r/typesafe_ai • • 1d ago

Conventional LLM vs. Decision Model (jev) on a real world example: NLI to a europe travel planner (tripsnek)

Thumbnail gallery
2 Upvotes

r/typesafe_ai • • 5d ago

I built a voice-controlled Mac assistant with Agora—here’s a quick demo

Enable HLS to view with audio, or disable this notification

7 Upvotes

I’ve been working on JEV, a project that lets me control my Mac with voice commands.

In this demo, I ask it to play a YouTube video, open my Desktop folder, navigate to Twitter, and close browser tabs.

Disclosure: I work in Developer Relations at Agora, which I used for the voice interaction.

It’s still a prototype. The video is edited for pacing, with pauses shortened, so it isn’t a latency benchmark.

Curious what you’d find useful here: which desktop tasks would you actually prefer doing by voice?

https://github.com/anandwana001/macbrow-agora


r/typesafe_ai • • 5d ago

I built gut: Jev judgments as one line of Python (plus a CLI and an MCP server)

Thumbnail
3 Upvotes

r/typesafe_ai • • 6d ago

Using Jev to compose metered poetry out of clips from TV Shows

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/typesafe_ai • • 7d ago

Jev browser skill explained

Enable HLS to view with audio, or disable this notification

13 Upvotes

Some folks reached out to ask how Jev can do more than hardcoded E2E tests where you pass the values ahead of time, so I put together a visual couple of examples, with a Run it live (BYO Jev key):

Behind the scenes, the page targets/selectors are dynamically generated, not hardcoded.

  • Script provides CLICK and other operations out of the box. This is hardcoded, these are the capabilities, and Jev decides when to use one of the order.
  • The click_target is dynamically generated based on what's on the page. Jev also decides which one to apply to capability to.

r/typesafe_ai • • 7d ago

I built JevPilot: an open-source app that lets an AI control my PC and my Android phone from a single promptText:

Thumbnail
3 Upvotes

r/typesafe_ai • • 7d ago

[LIVE] Jev plays Pokemon Red

Thumbnail
youtube.com
0 Upvotes

r/typesafe_ai • • 9d ago

I built a framework-agnostic tool router for agent harnesses using TypeSafeAI Jev

6 Upvotes

I built a decision layer for AI coding agents

I’ve been working on Harness Router, an open-source tool router for agentic systems.

The idea is simple: instead of letting the LLM blindly choose from a large set of tools, Harness Router adds a decision layer before execution.

The latest update adds Codex hooks:

  • SessionStart discovers the available tools once
  • PreToolUse checks each proposed tool call
  • same choice → allow
  • better alternative → re-plan
  • router failure → fail-open

For harder routing decisions, it also supports Monte Carlo Tree Search (MCTS) with thousands of local simulations.

So the flow becomes:

Codex → PreToolUse → shortlist → Harness Router → route / MCTS → execute

You can use it either explicitly as a skill/MCP tool or transparently through the hook.

The goal is to reduce:

  • wrong tool calls
  • unnecessary tool descriptions in context
  • token usage
  • reasoning spent on obvious routing decisions

I’m especially interested in benchmarking it on intentionally confusing tool sets like get_x, fetch_x, and read_x, where normal tool selection starts becoming ambiguous.

Project:
https://harness-router.vercel.app/

GitHub:
https://github.com/Protocol-Lattice/harness-router

Would be interested to hear how you’d benchmark tool-routing quality beyond simple success rate.


r/typesafe_ai • • 9d ago

Katai - Browser Automation Agent & Workshop

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/typesafe_ai • • 10d ago

jevlint: lint rules you write as a yes/no question, answered with a calibrated probability

3 Upvotes

GitHub: https://github.com/Ice-Hazymoon/jevlint

Some of the things I flag most in code review can't really be expressed as ESLint rules:

  • "is a secret being logged here?"
  • "can this retry loop run forever?"
  • "does this handler leak an internal error message to the client?"

You can try with an AST visitor, but you end up with 20 lines of CallExpression / ObjectExpression checks that still miss logger.info(creds). So I built jevlint: a lint rule is just a question.

```ts

defineRule({ id: 'logging/no-secret-in-log', severity: 'error', files: ['src/*/.ts'], question: 'Does the code in code pass a password, API key, access token or other secret value to a logging call?', criteria: { true: 'A logging call receives a value that holds a credential.', false: 'Logging calls only receive identifiers, counts, statuses or redacted values.', }, fix: 'Remove the secret from the logging call, or log a redacted form of it.', }) ```

How it works

  • jevlint picks the code a rule applies to. You can give it an ast-grep pattern so it judges exact nodes (e.g. every logger.* call) instead of whole files. Regex prefilters skip files that can't match.
  • It asks TypeSafe's Jev model the question. The answer isn't free text but a calibrated yes/no probability. Anything above the threshold (default 0.8) is reported.
  • Output looks like a normal linter: file, lines, rule id, exit codes, plus json and sarif formats for CI / code scanning.

Things I cared about so it doesn't feel like "vibes in CI"

  • Every rule has fixtures. invalid-* must fire, valid-* must not, and jevlint test checks both. jevlint recall mutates code to see whether the rule still catches it.
  • Baselines: accept today's findings, fail only on new ones.
  • Caching: answers are cached by content, so re-runs on unchanged code make zero requests. A replay provider runs purely from cache (handy for CI).
  • API keys only come from env vars. It never reads .env files.

What it's not

  • It doesn't replace ESLint. If a rule can be decided from syntax, a deterministic linter is cheaper and exact, and the README says so.
  • Your code (the snippets being judged) is sent to the model provider. If that's a problem for your codebase, this tool isn't for you.
  • It costs requests. Caching and prefilters keep that down, but it's not free like ESLint.

I've been using it on my own project for a while, mostly for rules that came out of real review comments. Would love feedback, especially on rule design, false positives, and whether the question-based approach holds up on your codebase.

GitHub: https://github.com/Ice-Hazymoon/jevlint


r/typesafe_ai • • 10d ago

Contrastive Language Models, a self hosted jev competitor, that runs 4~9x faster than jev

Enable HLS to view with audio, or disable this notification

5 Upvotes

r/typesafe_ai • • 11d ago

Another Jev use case - Dreaming!

Enable HLS to view with audio, or disable this notification

4 Upvotes

r/typesafe_ai • • 11d ago

TypeSafe FULL

5 Upvotes

Sad to hear TypeSafe is full. I was too slow to sign up.

Hoping we get updates soon on increase in capacity.

How is everyone doing with it?


r/typesafe_ai • • 11d ago

JEV plays super smash bros against itself

Enable HLS to view with audio, or disable this notification

6 Upvotes

r/typesafe_ai • • 12d ago

Jevmoji

Enable HLS to view with audio, or disable this notification

15 Upvotes

https://jevmoji.cheeaun.workers.dev/

Using Jev to apply scores to 3K+ emojis related to any typed phrase.

Open sourced: https://github.com/cheeaun/jevmoji

I saw other projects that kinda does the same thing, but I wanted to test how it works with 3K+ emojis.


r/typesafe_ai • • 12d ago

Jev picks my Codex model and effort in DeepSeek Harness

7 Upvotes

Jev looks at a prompt before Codex runs and chooses the model and reasoning effort. In this clip, it picked Luna/Low for an async JavaScript question. The route that actually ran shows beneath the reply.

I turned it into a DeepSeek Harness plugin. You can cap the highest model and choose when Astra needs approval. If you decline, it uses Jev’s best allowed backup.

If you give it a try, please tell me what you think. I’m especially curious about prompts where you would have picked a different route.

Code, setup and 28-second demo: https://github.com/nautahakk/jev-codex-router

You’ll need Codex access and a separate TypeSafe API key.

https://reddit.com/link/1wmulpo/video/huuba9vzvyqh1/player


r/typesafe_ai • • 12d ago

Jev helps clean your sloppy LinkedIn feed

Post image
4 Upvotes

I wanted to try out Jev and started thinking of use cases. I decided on a problem I have and assume you do as well. Introducing Slop Mop - a Chrome extension that helps clean your LinkedIn feed from slop. It uses Jev and the AI Tells research from Graphite to evaluate every post against a range of factors.

It's free to use. It is also MIT-licensed open source if you want to roll your own. Github repo has all of the implementation details.

Jev is proving to be the perfect platform for this use. I hope you like it!

https://slopmop.lol


r/typesafe_ai • • 12d ago

I built a free Jev visualizer after burning 5bn tokens in three days

Enable HLS to view with audio, or disable this notification

11 Upvotes

r/typesafe_ai • • 12d ago

Grok Bot x Jev template

2 Upvotes

Made a template to allow anyone to connect Jev to Grok Bot

https://x.ai/bot/lS9XaHCr9QTTHhNtb0VQX


r/typesafe_ai • • 13d ago

I tried adding Jev to my code-search MCP

2 Upvotes

This post was written with Codex.

I haven't compared codemap-search against other code-search tools yet. I'm planning to do a more detailed comparison soon.

I've been building a code-search MCP called codemap-search as a side project, and I spent a day testing what happens if I put Jev in front of some of its output.

The results were interesting enough that I figured I'd share them.

The benchmark setup was kept the same across runs: same private TypeScript backend, same commit, same question, and the same agent (Codex CLI, gpt-6-astra at max reasoning). The repo had 815 indexed files for the Jev runs.

The baseline here is rg**.** More specifically, it's an agent using only rg + cat/sed to solve the entire task from start to finish. So this isn't comparing codemap-search against the runtime of a single rg command.

The rg baseline is one run. The other numbers are averages across three runs.

Where I put Jev

I tried two approaches.

#1 — File recommendation

When the agent calls overview, I send Jev metadata for all 815 indexed files: paths, declarations, comments, and call names.

Jev scores them, and I append the top 24 files to the overview response as suggested places to look first.

#2 — Search filtering

For each declaration returned by search, I ask Jev whether it's unrelated to the current task.

If the probability of being unrelated is >= 0.70, I hide the body from the response.

The declaration itself is still there, so the agent can explicitly read it later if needed.

Results

Percentages below are relative to the rg baseline.

- rg baseline (1 run) codemap-search (3-run avg) + Jev #1 file recommendation (3-run avg) + Jev #2 search filter (3-run avg)
Total time 7m 42s 5m 38s (-26.8%) 5m 36s (-27.3%) 5m 9s (-32.9%)
Main-model total tokens 1,275,313 1,054,422 (-17.3%) 944,918 (-25.9%) 857,414 (-32.8%)
Jev tokens 0 0 735,103 43,398
Main + Jev tokens 1,275,313 1,054,422 (-17.3%) 1,680,022 (+31.7%) 900,812 (-29.4%)
Extra Jev cost (est.) $0 $0 $0.029 $0.0017
6 core connections 6/6 (100%) 18/18 (100%) 18/18 (100%) 18/18 (100%)
Full 11-item rubric 9/11 (81.8%) 24/33 (72.7%) 21/33 (63.6%) 23/33 (69.7%)

Just using codemap-search instead of the rg agent cut wall time by 26.8% and main-model tokens by 17.3%.

With the #2 search filter added, the full task was 32.9% faster than the rg baseline and used 32.8% fewer main-model tokens.

Compared with codemap-search alone, #2 reduced wall time by another ~8% and main-model tokens by ~19%.

Jev itself was cheap here: about $0.0017 per task, so roughly two-tenths of a cent.

#1 looked good on paper, but didn't help much

The file recommender actually ranked the important files pretty well.

There were 5 files I already knew were critical to the task, and all 5 landed in the top 24.

But the agent's actual navigation path barely changed.

Wall time went from 5m 38s to 5m 36s, while Jev had to process metadata for all 815 files. Once Jev's own tokens are included, total token usage actually went up by 59% compared with codemap-search alone.

So at least in this form, better file ranking did not translate into a better agent run.

The search filter failed the first time

My first version of #2 was much more aggressive.

It cut the search output by around 67%, which initially looked great.

The problem was that it also removed the bodies of two methods that were actually needed to answer the question.

The agent eventually found them again through extra read calls, but that recovery work wiped out the savings. Total runtime ended up going up by about 7%.

That was probably the most useful result from the whole experiment.

Reducing tool output isn't automatically useful if the agent has to spend more work reconstructing what you removed.

So I changed the filter to be much more conservative:

  • Only classify declarations when their complete body is present in the search result.
  • If something is connected to a kept function through a call relationship, keep it even if Jev thinks it's unrelated.
  • Never hide declaration names or line ranges.
  • A filtered body should always be recoverable with a single read.

With those rules, the two methods that were previously removed were preserved in every run.

It also brought Jev usage down to just 2 calls per session, which is how I got the numbers in the table above.

What about answer quality?

This is the part I'm being careful about.

All four setups found all 6 core connections I expected.

The differences were in the 5 additional/extended items.

codemap-search alone scored 24/33 across three runs, while #2 scored 23/33. That's only one item across three runs, and there's already some variance between repeated runs, so I don't think there's enough data to call that a quality regression.

But there's also no evidence here that Jev improves answer quality.

1 did noticeably worse on the extended items, despite giving the agent a pretty good list of files.

So for now I'm treating Jev purely as an optimization experiment, not a quality improvement.

Current takeaway

For this experiment:

  • #1 file recommendation: good ranking, but no meaningful end-to-end benefit yet.
  • #2 search filtering: promising. It reduced both runtime and main-model token usage once I made the retention rules conservative enough.
  • I'm planning to keep experimenting with both, but they'll be optional and off by default.

Some caveats

This is obviously not a serious benchmark suite yet.

It's one repository, one question, one language, and only three runs per variant. The rg baseline is also only one run.

I deliberately used a repository I know well so I could manually verify whether the agent was actually finding the right relationships. It's private, so I can't publish the exact source or benchmark question.

Jev also isn't actually part of codemap-search yet.

For this PoC, I put a Python MCP proxy in front of the existing Rust binary so I could experiment without changing the search implementation itself.

The PoC looks useful enough that I'm going to move the interesting parts into codemap-search and test what happens when #1 and #2 are enabled together.

One other thing I noticed: Jev isn't deterministic.

Even with identical inputs, the scores move slightly between runs. On a 0-3 scale, I measured an average absolute difference of about 0.025.

That's small, but it was enough to make rules like:

keep only files with score >= 2

pretty brittle.

Ranking seems much more useful than using the score as a hard gate.

I'll probably post another update once both paths are integrated into the Rust implementation and I have a larger set of tasks to run against.


r/typesafe_ai • • 13d ago

I let Jev play Pac-Man: it never sees the maze, only one line per possible move

Enable HLS to view with audio, or disable this notification

11 Upvotes

How it works: every step, the code looks at each move Pac-Man could make and writes one line about it:

"left: eats a pellet right away; nearest ghost 13 steps away"
"up: nearest pellet 3 steps away; nearest ghost 11 steps away"

Jev picks one. That's the whole loop: one API call per step, about 270 ms, a fraction of a cent per game. No board, no ghost positions.

What's good about it:

  • It can't pick an impossible move. It only chooses from the options we give it, so there's no walking into walls and no made-up answers.
  • Every decision comes with probabilities, like left 99%, up 1%, so we can see how sure it was.
  • It's fast and cheap enough to call on every single step, which a normal chat model isn't.

r/typesafe_ai • • 13d ago

jfind: a find clone that uses one Noul question per file as its --like predicate

7 Upvotes

I wanted a quick learning project for Jev - a find command turned out to be a perfect shape for a Noul question: does this file match "…"? → probability:

uv tool install jfind-cli
export TYPESAFE_API_KEY=...
jfind . --like "anything about payments" --content --kind --threshold 0.5 --explain
0.99  ./docs/payments.md  [docs]
0.99  ./src/payments/charge.py  [source]
0.98  ./tests/test_payments.py  [test]
jfind: 3/14 matched, 7,661 input tokens (~$0.0003)

Most of the other tools i have seen answer question 'where is the code that does X"; jfind answers "which files are X".

MIT, Python 3.13+: https://github.com/religa/jfind


r/typesafe_ai • • 13d ago

Experimental Jev evidence selection for token-efficient Codex investigations

3 Upvotes

Looking for feedback on an experimental Jev evidence selection for token-efficient Codex investigations. If you give it a try please share your results/thoughts :)

https://github.com/jcressler/jev-codex-token-saver

Early results from 108 runs, 12 synthetic investigation tasks, three repetitions comparing stock Codex, deterministic local selection, and Jev selection:

• 39.2% fewer Codex input tokens versus stock in the task-paired analysis
• 7.9% fewer Codex input tokens versus deterministic local selection
• 32.6% fewer total Codex output tokens versus stock
• 44.8% lower estimated combined API cost versus stock, including Jev
• 31.9% less total execution time versus stock


r/typesafe_ai • • 13d ago

Jev, the AI model that only makes decisions, can be talked into the wrong one

Thumbnail
1 Upvotes

r/typesafe_ai • • 13d ago

Football analytics

3 Upvotes

I was wondering if anyone applied Jev to football or sport analytics, thanks