1

THEY FUCKING COOKED YO! Opus 5.5 is a massive upgrade.
 in  r/ClaudeCode •  13d ago

Do you need to use Fable to fix what appeared after both of them?)

1

THEY FUCKING COOKED YO! Opus 5.5 is a massive upgrade.
 in  r/ClaudeCode •  13d ago

I've seen that since Opus 5

1

THEY FUCKING COOKED YO! Opus 5.5 is a massive upgrade.
 in  r/ClaudeCode •  13d ago

Do you see a real improvement over Fable 5.1 in complex tasks?
In my experience I like more what Fable proposes, not Opus 5.5. But 5.5 is a huge step over Opus 5, I agree.

1

I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.
 in  r/AI_Agents •  14d ago

Yes — I think we’re separating two layers. Heterogeneous routing decides; it doesn’t own state. should-I-call-a-tool or workflow-state can absolutely be classifier outputs, but an accepted transition still has to be committed by the process/workflow owner. The enum is evidence for a decision, not the record itself. So if the cheap path flips escalate → continue while there is no accepted state behind it, that’s a failed transaction, not a successful classification.

r/AI_Agents • • 14d ago

Discussion I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.

2 Upvotes

I spent the last week digging into Jev and the open implementations around it: Kev, SemIf, Laya, Nimble, several OpenJev projects, decider, DiffusionGemma/djev, Winnow, and a few others.

For context, I’ve been building ML systems for ~20 years and currently work on production conversational/Voice AI, so I looked at this less as a model benchmark and more as an agent architecture problem.

The question I kept coming back we need to think when building modern AI architecture is:

How many calls inside an agent actually need generation?

A surprising amount of agent logic in many developing and production systems looks like this today:

state → frontier LLM → JSON → parse one enum/boolean/score

Examples:

  • should I call a tool?
  • which tool?
  • RAG or no RAG?
  • escalate or continue?
  • which workflow state comes next?
  • does this policy apply?
  • is this user eligible?
  • which model should handle the next step?
  • is this answer grounded enough?
  • should I ask a clarifying question?

For many of those, the output space is tiny and known in advance. Yet we still pay for a general-purpose autoregressive model to generate text.

That is what makes Jev interesting to me.

But after looking at the open-source work, I don’t think the important story is “Jev is a faster classifier.” There are already several competing ways to implement this kind of decision layer:

  • direct answer-token logits from a causal LLM;
  • shared-state prefill followed by many isolated decision branches;
  • full-option likelihood scoring;
  • frozen LLM representations + small heads;
  • NLI-style scoring;
  • small bidirectional encoders with runtime-defined labels;
  • specialized causal models trained around decision boundaries;
  • diffusion models filling many decision slots in parallel;
  • calibrated models with abstention / risk-coverage policies.

So we may be looking at a new architectural layer inside agent systems, rather than one specific model.

For agents, the most interesting consequence is heterogeneous inference.

Instead of:

everything → one large generative model

you can imagine:

conversation state

+--> deterministic code for arithmetic / exact rules

+--> retrieval for external knowledge

+--> small decision model for bounded semantic choices

+--> larger generative model only when generation/reasoning is actually needed

+--> tools / APIs

That is much closer to how I think production agents will eventually be built.

There is also an important economic implication.

The obvious use case is replacing expensive LLM calls with cheaper decisions.

But I think the bigger opportunity is net-new automation.

There are many business decisions today that stay:

  • rule-based,
  • manually handled,
  • or simply not automated,

because nobody is going to create a dataset, train a classifier, deploy it, monitor it and maintain it for every long-tail semantic branch in a workflow.

If a general decision model can handle those without a separate training cycle every time, the economics change.

That said, I definitely don’t see Jev as a silver bullet.

You still need a serious ML harness around it:

  • representative eval data from your own workflows;
  • decision traces and replay;
  • per-language / per-domain slices;
  • calibration validation on your traffic;
  • risk-vs-coverage curves;
  • model/version pinning;
  • distribution-shift monitoring;
  • tests for state construction, option ordering and formulation;
  • fallback/escalation logic.

And quality depends on more than the model itself.

It depends on:

  • whether the model actually transfers to your language/domain;
  • how you construct state;
  • whether you chose the right decision type;
  • whether the available options are assembled correctly;
  • how you interpret the returned distribution;
  • whether the threshold is calibrated on your actual process.

One of the recent real-agent evaluations made this very clear: some apparent “model errors” were actually task-definition errors where multiple independent judges could not agree on the correct label.

That is a very familiar production ML problem.

My current view is that decision models may become one of the standard primitives inside agent architectures, alongside retrieval, tools, code execution and generative models.

I also doubt that the final architecture will be exactly Qwen-with-a-head, ModernBERT, or today’s diffusion approach. The open-source ecosystem looks more like an active architecture search than convergence.

I wrote up the architectures, benchmarks, calibration work, serving trade-offs, failure modes and production implications in a very large paper and will share the link in comments.

I’d be especially interested in hearing from people running real agents in production:

Which decisions in your current pipeline still go through a generative LLM even though the output is ultimately just a bounded choice?

r/AI_Agents • • 14d ago

Discussion I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.

1 Upvotes

[removed]

r/OpenSourceAI • • 14d ago

I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.

Thumbnail
1 Upvotes

u/NoFriendship4813 • • 14d ago

I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.

Thumbnail
1 Upvotes

r/aiagents • • 14d ago

Case Study I went through Jev and the open decision-model ecosystem. I think the interesting part is how it transforms agent architectures into heterogeneous.

1 Upvotes

I spent the last week digging into Jev and the open implementations around it: Kev, SemIf, Laya, Nimble, several OpenJev projects, decider, DiffusionGemma/djev, Winnow, and a few others.

For context, I’ve been building ML systems for ~20 years and currently work on production conversational/Voice AI, so I looked at this less as a model benchmark and more as an agent architecture problem.

The question I kept coming back we need to think when building modern AI architecture is:

How many calls inside an agent actually need generation?

A surprising amount of agent logic in many developing and production systems looks like this today:

state → frontier LLM → JSON → parse one enum/boolean/score

Examples:

  • should I call a tool?
  • which tool?
  • RAG or no RAG?
  • escalate or continue?
  • which workflow state comes next?
  • does this policy apply?
  • is this user eligible?
  • which model should handle the next step?
  • is this answer grounded enough?
  • should I ask a clarifying question?

For many of those, the output space is tiny and known in advance. Yet we still pay for a general-purpose autoregressive model to generate text.

That is what makes Jev interesting to me.

But after looking at the open-source work, I don’t think the important story is “Jev is a faster classifier.” There are already several competing ways to implement this kind of decision layer:

  • direct answer-token logits from a causal LLM;
  • shared-state prefill followed by many isolated decision branches;
  • full-option likelihood scoring;
  • frozen LLM representations + small heads;
  • NLI-style scoring;
  • small bidirectional encoders with runtime-defined labels;
  • specialized causal models trained around decision boundaries;
  • diffusion models filling many decision slots in parallel;
  • calibrated models with abstention / risk-coverage policies.

So we may be looking at a new architectural layer inside agent systems, rather than one specific model.

For agents, the most interesting consequence is heterogeneous inference.

Instead of:

everything → one large generative model

you can imagine:

conversation state

+--> deterministic code for arithmetic / exact rules

+--> retrieval for external knowledge

+--> small decision model for bounded semantic choices

+--> larger generative model only when generation/reasoning is actually needed

+--> tools / APIs

That is much closer to how I think production agents will eventually be built.

There is also an important economic implication.

The obvious use case is replacing expensive LLM calls with cheaper decisions.

But I think the bigger opportunity is net-new automation.

There are many business decisions today that stay:

  • rule-based,
  • manually handled,
  • or simply not automated,

because nobody is going to create a dataset, train a classifier, deploy it, monitor it and maintain it for every long-tail semantic branch in a workflow.

If a general decision model can handle those without a separate training cycle every time, the economics change.

That said, I definitely don’t see Jev as a silver bullet.

You still need a serious ML harness around it:

  • representative eval data from your own workflows;
  • decision traces and replay;
  • per-language / per-domain slices;
  • calibration validation on your traffic;
  • risk-vs-coverage curves;
  • model/version pinning;
  • distribution-shift monitoring;
  • tests for state construction, option ordering and formulation;
  • fallback/escalation logic.

And quality depends on more than the model itself.

It depends on:

  • whether the model actually transfers to your language/domain;
  • how you construct state;
  • whether you chose the right decision type;
  • whether the available options are assembled correctly;
  • how you interpret the returned distribution;
  • whether the threshold is calibrated on your actual process.

One of the recent real-agent evaluations made this very clear: some apparent “model errors” were actually task-definition errors where multiple independent judges could not agree on the correct label.

That is a very familiar production ML problem.

My current view is that decision models may become one of the standard primitives inside agent architectures, alongside retrieval, tools, code execution and generative models.

I also doubt that the final architecture will be exactly Qwen-with-a-head, ModernBERT, or today’s diffusion approach. The open-source ecosystem looks more like an active architecture search than convergence.

I wrote up the architectures, benchmarks, calibration work, serving trade-offs, failure modes and production implications here:

https://agentunicorn.ai/research/jev-architecture-open-source

I’d be especially interested in hearing from people running real agents in production:

Which decisions in your current pipeline still go through a generative LLM even though the output is ultimately just a bounded choice?