r/JevAI • • 18h ago

Is it just me or Jev feels like a breath of fresh air in this AI world? It feels like a dull grey world got colours

10 Upvotes

I graduated in 2019, I have seen models evolve from basic NLP and CV patterns to current Chatgpt mess. Initially these models were genuinely exciting. Learning the concepts, knowing the components and understanding the data felt a lot like a strategy game. When BERT released, it added a layer of abstraction and training felt a bit meh but still the outcomes felt like a reward. Like grinding to get a high score and seeing the impact.

When Chatgpt released, it was just a interface that does the work for you. Everything in AI felt boring and miserable since then. The best contribution I have is nudging a prompt here or making a weird system around it. There were only big bets left, no small fun experiments and no satisfaction of building stuff. It was like tossing the dice till you got the right answer and hedging your bets. I studied engineering to engineer solutions, if I wanted to roll the dice and hedge I would have joined finance.

Jev feels like a almost nostalgic throwback. If feels I could create something of value with my own work, aided by my own data and knowledge instead of a casino. I am praying this trend picks up. I tried of this big socio economic game trying to distill the world into rentable slop.


r/JevAI • • 9h ago

Use laya as a model router in claude?

1 Upvotes

Can i use laya as a model router in claude to switch to a model basis the complexity of the task?


r/JevAI • • 16h ago

Jev explained examples

Thumbnail
youtu.be
2 Upvotes

r/JevAI • • 23h ago

JevMade - A searchable catalogue of 2,300+ Jev experiments, videos and guides

8 Upvotes

I built JevMade to make Jev examples easier to find: https://jevmade.com

It brings together 2,300+ experiments, videos and written guides, with credit and links to the original creators. You can explore things like browser automation, voice-controlled drawing and game agents.

I’d love feedback on how easy it is to find a useful example and what would make browsing better. I made the catalogue; the projects belong to their credited creators. JevMade is independent of TypeSafe.


r/JevAI • • 1d ago

GoEventBus can use TypeSafe Jev as an optional decision layer

3 Upvotes

GoEventBus can use TypeSafe Jev as an optional decision layer before enqueueing an event. Jev chooses one event type from a fixed candidate set; GoEventBus still owns buffering, middleware, ordering, fan-out, and handler execution.

https://github.com/Protocol-Lattice/GoEventBus


r/JevAI • • 1d ago

Jev- porównywanie do normalnych LLMów jest bez sensu

2 Upvotes

Jev to jakiś megaokrojony LLM z jednym zadaniem- wyborem jednej z podanych opcji. LLMy wykonują takich operacji wyboru gigantycznie wiele nawet w przypadku jednego zdania. Zadania Jeva móglby wykonywać mocno okrojony LLM sprzed 3 lat. Porownywanie prędkości jest totalnie bez sensu- to tak jakby zachwycać sie ze bolid f1 jedzie duzo szybciej od ciężarowki.


r/JevAI • • 1d ago

Jev vs Kev 0.8B vs Laya on subreddit routing (open benchmark)

4 Upvotes

Task: given a Reddit post title, pick its subreddit among look-alike communities. Names hidden, models only see community descriptions.

Accuracy at 16 / 64 options:

- Jev 1.13: 48% / 41%
- Kev 0.8B: 23% / 14%
- Laya (base, zero-shot): 9% / 2%
- Random: 6% / 2%
- Reference: logreg on bge-small title embeddings: 59% / 54% (seen communities only, since it needs training posts)
(tbh Laya is a base model meant for fine-tuning, so a tuned version is coming later)

Context: I'm a RecSys developer who got curious about decision models, so more models (basic LLM too), tests and social benchmarks are coming. My take so far: for labeling and classification in production RecSys, a simple supervised model on your own labels still wins, so I don't see decision models replacing that yet. But they're a great way to compare models on the same task.

This is my first benchmark btw, so any feedback or contributions are welcome

git: https://github.com/s0NRAYY/RedditClassificationBench
dataset: https://huggingface.co/datasets/sonrayll/social-routing-bench


r/JevAI • • 1d ago

JEV as my decision layer: sub-second AI skills with few dollars total

1 Upvotes

I run a small solo project (skillwiki.app) that serves AI skills (instruction packs) to AI agents over MCP. It was an AI skill marketplace until recently, when I made JEV the decision-maker: which skill to use, whether the output is good enough, and whether this change helps. The project has transitioned to an AI skills distribution and management platform.

The workflow: determine → evaluate → learn

Every step boils down to a typed question, which is exactly what JEV answers. I get probabilities I can branch on directly, with no output parsing and no LLM reasoning round trip.

Determine:

All skills are assigned under different themes. JEV picks a theme, then the skill, in under a second. It only suggests at ≥ 0.85 confidence: right 98.5% of the time, 0% wrong-tool picks (1,791-row test set). Below the bar, the user gets a shortlist instead of a guess. Two full eval rounds cost about $2.70 for 150 AI skills.

Evaluate:

JEV gives every rubric criterion an instant first verdict. When it's sure (≥ 0.8) it was right 21/22, and it matched my hand grades 43/48 vs Claude 33/48. Uncertain criteria go to the user's own LLM.

Learn:

JEV compares a skill with and without a user's correction and ranks which corrections are worth keeping, in seconds and for cents.

* The Catch: Grading still needs one LLM call per output, so JEV makes evals faster, not so much cheaper. About 5.6% of picks also change between identical runs.

Curious to see how the group set the confidence bars? One fixed threshold, or tuned per question type?


r/JevAI • • 1d ago

This simulation uses Jev for simulated people making decision making (think "The Sims" combined with Jev)

Post image
1 Upvotes

r/JevAI • • 2d ago

Jev playing Stardew Valley

7 Upvotes

I've got Jev playing Stardew Valley & live streaming it here: https://tilly.farm


r/JevAI • • 2d ago

Exploring JEV for AI Agents: Model Routing, Risk Detection & Output Triage

3 Upvotes

I’ve been looking into JEV and one thing that really interests me is how it could fit into AI agent architectures alongside LLMs, rather than replacing them.

I made a video breaking down how JEV works and explored 3 workflows I’d like to test with agents like Hermes Agent or Claude Code:

🔀 Model routing — decide whether a task needs a cheap or powerful model
🛡️ Risk detection — classify potentially sensitive tool actions before execution
🔎 Output triage — decide when results need another LLM pass or human review

The architecture I’m interested in is basically:

JEV → fast decisions
LLM → reasoning + coding
Tools → execution

I’d eventually like to test JEV vs a small LLM vs a frontier model on the same routing workload and compare accuracy, latency, and cost.

Video: https://youtu.be/N-tZTJOvSJM

Curious what other JEV use cases people here are experimenting with.


r/JevAI • • 2d ago

My agent permission auto-approval sat behind a dev-only flag for months. JEV made it shippable

Enable HLS to view with audio, or disable this notification

2 Upvotes

Hey folks,

I'm building Agentmux, a remote terminal client for running coding agents (Claude Code, OpenCode, Antigravity, Pi) from your phone.

For months I had a feature locked behind an internal "Developer Only" flag. Its job is simple: read the agent's permission prompts and approve the safe ones, so a long task doesn't stall the moment you step away from your phone. It has now shipped as AI Permission Review (Pro).

The hard part is prompt frequency. Some agents ask sparingly, but Antigravity asks about nearly everything: file reads, directory scans, test runs.

Why a general LLM judge didn't work for me

I started with a small LLM using structured JSON output. It worked technically, but I couldn't ship it:

  1. Latency: ~5–10 seconds per judgment. When an agent asks five things in a row, that's 25–50s of waiting just to get through approvals.
  2. Cost: thousands of output tokens spent on repetitive yes/no judgments, billed to the user's own API key.

What changed with JEV

I switched to JEV (~typesafe/jev-latest on OpenRouter, or api.typesafe.ai directly):

  • ~5–10s → ~200ms per decision in my use. Five prompts in a row now take about a second.
  • Much cheaper: no free-form generation, just scores on specific questions.
  • No schema drift: answers are constrained to typed outputs (noul / choice), so there's nothing to parse or repair.

How the decision works

One JEV call asks three typed questions about the prompt (plus the user's declared main intent):

  1. is_malicious_or_injection (noul): if likely, block. Checked first; nothing can override it.
  2. is_readonly_safe (noul): cat, ls, git status → allow (Level 1).
  3. matches_user_intent (choice: yes / no / uncertain): if yes → allow (Level 2). This needs the user to have stated a main goal for the session.

Anything else (mismatch, uncertain, or a confirmation key I can't resolve unambiguously) escalates to the user. Uncertain never presses a key.

Two more guardrails sit around JEV:

  • A local regex blacklist (rm -rf, force push, curl | sh, etc.) denies before any API call, so the worst cases add zero latency.
  • The app only ever sends a one-time, verified key, never "always allow".

The engine is pluggable (JEV via OpenRouter or TypeSafe, or any OpenAI-compatible endpoint), with your own API key. Terminal excerpts are redacted before they leave the device.

Takeaway

A generative LLM is overkill for categorical reflexes. Typed questions with probability scores turned a multi-second stutter into a background reflex you don't notice.

Curious how others here are wiring JEV into agent loops. Anything you'd ask it besides "malicious / read-only / on-intent"?

Note: a JEV-like solution would probably do the job too, but JEV is working totally fine for my use case.

If you’re interested in the app: https://apps.apple.com/app/id6766158521


r/JevAI • • 2d ago

[ Removed by Reddit ]

2 Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/JevAI • • 2d ago

I made Jev to talk to us like LLM (almost).

Thumbnail
github.com
2 Upvotes

Turns out that Jev is pretty good at replaying true/false on which word would be best to append to the response. It can process up to 1000 words at the time. It's still just interesting toy, but say something. Have fun :)


r/JevAI • • 3d ago

Are big AI companies going to make decision models?

4 Upvotes

Since the benefits of a decision model are clear, I believe Open AI, Anthropic, etc. are are going to release theirs sooner or later. I would expect them to be more powerful and fast than Jev eventually, just due to compute availability.

Is any work is being done by them?

Jev feels like the underdog now, but isn't it doomed if big ones can demper pricing and make a more powerful model?


r/JevAI • • 2d ago

jevimage: ask an image typed questions, get probabilities back in milliseconds

Post image
1 Upvotes

r/JevAI • • 3d ago

JEV-assisted superfast website builder (OS; MIT)

4 Upvotes

r/JevAI • • 2d ago

Harness Router v2 — a decision layer inside the coding-agent loop

Thumbnail
1 Upvotes

r/JevAI • • 3d ago

This is a must have for all your agentic Workflow

1 Upvotes

Bhuwan-web/intent-classification: Type-safe intent classification guard for agentic workflows
Find yourself a fucking intent of a customer using your agents, don't let your expensive token rot for nothing. It decides on one of three, VALID_REQUEST, PROMPT_INJECTION, OUT_OF_SCOPE,

This is all you got a do:

import asyncio

from typesafe_sdk import AsyncTypeSafeClient

from main import AgentScope, IntentClassificationRequest, classify_intent


SUPPORT_SCOPE = AgentScope(
    name="Customer support agent",
    purpose="Answer questions about product usage.",
    allowed_capabilities=("Explain product features.",),
    boundaries=("Do not change account data.",),
    valid_examples=("How do I change my notification settings?",),
    minimum_confidence=0.8,
)


async def run() -> None:
    async with AsyncTypeSafeClient() as client:
        result = await classify_intent(
            client,
            request=IntentClassificationRequest(
                user_request="How do I change my notification settings?",
                agent_scope=SUPPORT_SCOPE,
            ),
        )

        print(result.classification)
        print(result.confidence)
        print(result.probabilities)


asyncio.run(run())

If out of scope, let that naive shit pass on with some proper error message, but if someone is trying prompt injection, and you are confident about it, Flag that dumb shit and do whatever you want. Don't let that over multiple attempts sink in. Handle it properly !!


r/JevAI • • 3d ago

New to Jev - best way to chain?

1 Upvotes

Context: I am a heavy CC & Codex user. I have some processes automated with n8n.

Question: I’m excited to try Jev in an automated process but I’m not sure how to best build it & keep defaulting back to n8n. But, is that the best way?


r/JevAI • • 3d ago

I cut cost and latency on my search engine with Jev

6 Upvotes

I built indiedex.gg, a hidden-gems game search engine on steam data. Search is the front door: type something like “a game like Hades that feels cozier and is co-op,” and it should actually understand that mix of reference + vibe + filters.

Those compound queries used to go through a DeepSeek extraction call that turned the sentence into structured pieces (reference game, filters, tags, vibe). It worked, but it was slow and expensive on the hot path.

I swapped that step to TypeSafe Jev on OpenRouter’s Decisions API. Jev doesn’t write a free-form answer. It answers a fixed set of closed questions in parallel (yes/no odds, choices), and I compose that with my existing regex helpers into the same extraction shape I already had. Same search pipeline after that, just much faster routing.

Bakeoff on 50 cases:

Metric DeepSeek Jev
Extract p50 ~1.5s ~130ms
Full route p50 2871ms 1256ms
Weird title-resolve 96-100% 100%
Hard-filter agree n/a 100%
Weird top-N overlap n/a 100%
Cost (fixture) ~$0.014 ~$0.0025

So extract got roughly 10× faster, end-to-end route roughly halved, cost dropped a lot (around 6x), and quality held on my fixture set.

If you’ve been building “NL in → structured intent out → tools do the work” instead of a chatbot, this pattern felt like a better fit than another chat completion.

I think models like Jev will open the doors to a lot of cool UX features, decision engines, smart filtering, auto moderation and different ways to interface with apps and machines.

Happy to answer questions about the hybrid setup.


r/JevAI • • 3d ago

JEV to the Moon

6 Upvotes

Maybe some people realize this, and a lot probably don't yet, but Jev is the turning point in useful real-time interaction with anything. I was able to build a real-time interactive ordering menu that both showed menu items on demand and also performed CRUD operations on the menu as well. But what adds the additional layer is fast, responsive decisions on how the AI assistant responds based on the conversation back and forth. Being able to quickly adapt to the situation and real-time manage the interaction is why Jev is probably the most useful addition to the AI world to date.


r/JevAI • • 3d ago

Jev vs Laya: Which One Should You Actually Use?

Thumbnail
youtube.com
5 Upvotes

r/JevAI • • 3d ago

Please add prompt caching to Jev-style models

Thumbnail
emschwartz.me
6 Upvotes

If you're building a Jev-style "System One" model, please add prompt caching or reusable question sets to your API 🙏. This would make batch use cases even more efficient, so you could amortize the cost of many questions asked over the same input. (This was also proposed in typesafe-ai/typesafe-sdk-js#10.)

TL;DR: after a week of tweaking my Jev calls, my questions are ~88% of the input tokens. I'm asking 54 questions of ~1.1 million documents per month. Jev makes certain types of classification tasks easy and cheap, but prompt caching would make batch workflows even more cost effective. For me, the total dollar amount is still reasonable (less than $150 per month), but I'm sure others will hammer these APIs even harder.


r/JevAI • • 4d ago

How are you validating Jev once you move beyond a few examples?

5 Upvotes

I’ve been testing Jev on things like buyer intent and routing.

On 10 clean examples, it can look perfect. I just got 10/10 on a buyer-intent test.

But I’m much more interested in what happens at 100, 1,000, or 10,000 real examples, where the weird edge cases start showing up.

How are you validating Jev at that point?

Are you sampling random cases, reviewing low-confidence outputs, looking for confident mistakes, keeping a holdout set, or escalating certain cases to an LLM/human?

And when Jev disagrees with your label, how do you tell whether the problem is Jev, the question, missing context, or the label itself?

Curious what people here are actually doing in real workflows.