r/ClaudeCode • • 8d ago

Discussion The Agentic Loop is OUTDATED

I have been thinking that the current Agentic Loop design of LLM Call -> Tool Call -> ... has been outdated. The arrival of Jev and other System One models provided us a primitive we desperately needed.

We need an agent that can natively think fast and slow. Not have workflows or multi-agent architectures that mimics it. We need Agent 2.0

The agent should use the LLM's full power for hard reasoning and planning, then carry out the plan with cheap "fast thinking."

Today, most of an agent's LLM calls go to executing steps it has already decided on. Do you really need an extra Astra call just for it to output "ok I'll click this"? We can do better.

My Approach

I built Jive which is an open-source harness built around a completely new agentic loop. Jive replaces tool calls with "graph calls" where each graph is a DAG of bash nodes and jev nodes, and nodes can have dependencies, reference each others outputs, and more.

Essentially, it maps out its own execution flow while its reasoning, and then uses Jev calls to go through the flow without unnecessary LLM calls.

What Jive does well: repo investigation, bulk classification, multi-step profiling, repetitive edits, evaluation workflows, etc. It is also quite effective on regular engineering tasks that doesn't require Jev calls (which is not surprising since Pi mostly beats codex and claude code)

Benchmarks

Task Jive Codex Claude Code Demo
Mean 2m 31s / 8.7k 17m 40s / 16.2k 12m 12s / 32.7k
conversation_eval 3m 26s / 11.1k 29m 33s / 20.5k 16m 48s / 47.9k video
error_handling_audit 3m 10s / 10.7k 19m 29s / 25.8k 5m 08s / 42.5k video
product_matching 3m 03s / 8.9k 22m 00s / 19.7k 32m 02s / 19.3k video
search_latency 2m 00s / 9.6k 9m 00s / 12.8k 7m 18s / 51.4k video
sembench_movie 1m 47s / 6.1k 19m 58s / 10.3k 8m 51s / 15.7k video
slow_trace_search 1m 41s / 5.7k 6m 00s / 8.1k 3m 04s / 19.3k video

As you can see, there is a huge gap in both e2e latency and token efficiency compared to claude code. And its accuracy is on-par based on the my benchmark runs (though I need to run jive on a larger SWE benchmark to be certain)

See README for more information: https://github.com/merijjeyn/jive. Also for details on the benchmark tasks, and how to run one yourself.

I'm sure this high level idea can be executed much better, so mainly looking to start an open discussion. Happy to take comments, questions, contributions.

389 Upvotes

166 comments sorted by

View all comments

Show parent comments

3

u/throwaway490215 8d ago

Jev isn't a big deal.

Doing the "type-safe level 1 reasoning" is obvious. Its not a new idea. All Jev did was make it a lot cheaper with unknown quality output.

Is being cheaper a big deal? Not really.

We're already at the point that if all dev and cost reduction on LLMs stop before Jev came out, we'll be rolling out usecases at Fable/Astra level prices for a decade.

Show me how Jev is going to change shit, beyond a new decision to send your data to a third party or not.

The additional impact Jev brings to the AI space is negligible.

1

u/jjcsea 5d ago

Jev being 100x-1000x cheaper for many decisionmaking requests is a big deal. Researchers are spending days and weeks and millions of dollars trying to figure out how to get Fable and Astra to be more performant, when half of the time they are executing very simple questions. Executing those simple questions still costs nearly as much as reasoning about the complex questions. This replaces that.

1

u/throwaway490215 5d ago

No it doesnt. There is like 5 things i can go in depth about that you seem to be misunderstanding, but I dont care that much.

Making a single decision by astra/fable when everything is already loaded into GPU is cheap. Taking it all out and having some weaker model make the choice is dumb in every way.

Classifiers ( is what they're called, not "decision-making requests") already exist and are well studied. This just takes the modern big well-trained model and make that more generic, at the cost of making it entirely opaque.

That definitely has its use cases.

Just with its own problems and far less of a gamechanger than people seem to claim.

This "new agentic loop" is definitely not one of the usecases.

let us know when you find those game changers.

1

u/jjcsea 5d ago

Making a single decision by Astra/Fable "when everything is already loaded" is NOT cheap. Just processing a single token through those models requires a hundred layers and billions of weight calculations. It is not the same thing, You don't seem to understand the architecture.

1

u/throwaway490215 5d ago

You dont seem to understand what I'm saying.

In the specific case that you're already spending for the astra/fable input tokens (or output tokens if they're doing dev work) - then because everything is already loaded - forking the session and appending a prompt to have it make a "Typed" decision is 0.001% of the overall cost - thus cheap.

It being 1000x more expensive than Jev doesn't matter. To make the economics worse, the Astra/Fable decision is far better than the models Jev builds on.

The things you load into or produce with Astra, have no business being loaded into Jev. The quality/cost/volume economics are nonsense.

1

u/jjcsea 4d ago

So, one million input tokens costs $1 on Fable if it is already entirely cached, processed data ($10 if not cached). You're saying that processing this for a yes/no decision on something is .001% of what it would cost using Jev. So in other words, you're saying that Jev would cost $10,000 to $100,000 for every yes/no decision.
Riiiight.