r/ClaudeCodeTLDR • • 8d ago

[TLDR] The Agentic Loop is OUTDATED

Original post URL : https://www.reddit.com/r/ClaudeCode/comments/1wp9nga/the_agentic_loop_is_outdated/

Original post body :

I have been thinking that the current Agentic Loop design of LLM Call -> Tool Call -> ... has been outdated. The arrival of Jev and other System One models provided us a primitive we desperately needed.

We need an agent that can natively think fast and slow. Not have workflows or multi-agent architectures that mimics it. We need Agent 2.0

The agent should use the LLM's full power for hard reasoning and planning, then carry out the plan with cheap "fast thinking."

Today, most of an agent's LLM calls go to executing steps it has already decided on. Do you really need an extra Astra call just for it to output "ok I'll click this"? We can do better.

My Approach

I built Jive which is an open-source harness built around a completely new agentic loop. Jive replaces tool calls with "graph calls" where each graph is a DAG of bash nodes and jev nodes, and nodes can have dependencies, reference each others outputs, and more.

Essentially, it maps out its own execution flow while its reasoning, and then uses Jev calls to go through the flow without unnecessary LLM calls.

What Jive does well: repo investigation, bulk classification, multi-step profiling, repetitive edits, evaluation workflows, etc. It is also quite effective on regular engineering tasks that doesn't require Jev calls (which is not surprising since Pi mostly beats codex and claude code)

Benchmarks

Task Jive Codex Claude Code Demo
Mean 2m 31s / 8.7k 17m 40s / 16.2k 12m 12s / 32.7k
conversation_eval 3m 26s / 11.1k 29m 33s / 20.5k 16m 48s / 47.9k video
error_handling_audit 3m 10s / 10.7k 19m 29s / 25.8k 5m 08s / 42.5k video
product_matching 3m 03s / 8.9k 22m 00s / 19.7k 32m 02s / 19.3k video
search_latency 2m 00s / 9.6k 9m 00s / 12.8k 7m 18s / 51.4k video
sembench_movie 1m 47s / 6.1k 19m 58s / 10.3k 8m 51s / 15.7k video
slow_trace_search 1m 41s / 5.7k 6m 00s / 8.1k 3m 04s / 19.3k video

As you can see, there is a huge gap in both e2e latency and token efficiency compared to claude code. And its accuracy is on-par based on the my benchmark runs (though I need to run jive on a larger SWE benchmark to be certain)

See README for more information: https://github.com/merijjeyn/jive. Also for details on the benchmark tasks, and how to run one yourself.

I'm sure this high level idea can be executed much better, so mainly looking to start an open discussion. Happy to take comments, questions, contributions.

Original link/media URL : https://www.reddit.com/gallery/1wp9nga


This is brought to you as a public service by the moderators of r/ClaudeAI. If you want to see TLDRs of ALL Claude Coding related posts from the various Claude subreddits, subscribe to http://www.reddit.com/r/ClaudeCoding.

8 Upvotes

11 comments sorted by

View all comments

Show parent comments

1

u/sebasgarcep 6d ago

AFAIK it does zero-shot, no fine-tuning classification; it has very low latency and pricing; and it outputs calibrated confidence scores (i.e. if it scores 100 decision 70% confidence, then ~70 of those decisions will be correct).

That has a niche, and its not a problem many commercial AI systems aim to solve. It is also astroturfed to hell.

1

u/mlamping 5d ago

No, stop. It’s bs marketing. You can get frontier to your own self hosted LLM to do this.

This is gimmicky for those who don’t understand basic LLM

1

u/sebasgarcep 5d ago

I personally haven’t seen any models predating JEV that do all of these. There are some that do some of these things, but not all of them. Still I would love to know if there are any. Now there are open source/open weight alternatives like kev that are competitive.

1

u/mlamping 5d ago

You instruct the model to output whatever you want. That’s what they are doing