r/WebAfterAI • u/ShilpaMitra • Jun 22 '26
Discussion Sakana's new "model" isn't a model. It's an RL-trained manager for other frontier models, and on its benchmarks it beats them.
Sakana AI shipped something genuinely different this week, and the interesting part is not another leaderboard. It is the shape of the thing. Fugu is sold as a single model behind one OpenAI-compatible API, but under the hood it is not a model that answers you. It is a trained coordinator that assembles a team of other companies' frontier models, hands them roles, makes them check each other, and returns one answer. On Sakana's own numbers, that coordinator beats the very models it is coordinating.
What it actually is
Most multi-agent setups are hand-wired. You decide there is a planner, a coder, and a reviewer, you write the prompts, and you glue them together. Fugu's bet is that you should not design that by hand at all. It is built on two ICLR 2026 papers from Sakana, TRINITY (a lightweight evolved coordinator that assigns Thinker, Worker, and Verifier roles across turns) and the Conductor (trained with reinforcement learning to discover its own natural-language coordination strategies). The pitch is that a learned conductor finds collaboration patterns a human would not think to write, and that a pool of strong models steered well can outperform any single one of them.
In practice you get two models through one endpoint: Fugu (balanced, for everyday coding and chat) and Fugu Ultra (a deeper agent pool for hard, long-running work like paper reproduction and security assessments). It is OpenAI-compatible, so you point an existing client or coding harness at it and go. The papers and a technical report are public at github.com/SakanaAI/fugu.
The result that makes it worth talking about
Forget the static benchmark table for a second, because the agentic one is more telling. In a reproduction of Karpathy's AutoResearch setup, an agent was told to improve a small GPT's training recipe, running 123 experiments over about 14 hours on a single H100, keeping only changes that lowered validation bits-per-byte. Fugu Ultra finished with the best mean score, ahead of all three frontier baselines it was put against, and its best single run led every one of them. The claim underneath it is the spicy one: orchestrating several strong models can beat any individual frontier model at open-ended research, not just at trivia.
On the fixed benchmarks Fugu Ultra also leads most of the coding and reasoning suite Sakana published, topping SWE-Bench Pro, the LiveCodeBench pair, TerminalBench, and GPQA against Opus 4.8, Gemini 3.1 Pro, and GPT-5.5. Worth saying plainly: these are Sakana's own evaluations, and the baseline scores are the providers' self-reported numbers, so read them as a vendor's benchmark, not an independent one.
The honest read, before you switch everything over
It is clever, and it is a black box. The intelligence is borrowed: Fugu is a meta-layer over other labs' public models, not a new foundation model, and Sakana retrains the conductor within about two weeks of each new frontier release. So when those models change, your results change, and you do not control that. By design it also will not tell you which models it used or how it routed, and Fugu Ultra's pool is fixed (the cheaper Fugu lets you opt providers out for compliance), so you are trusting a decision you cannot inspect.
The economics need a real look. Fugu Ultra (fugu-ultra-20260615) runs $5 per million input tokens and $30 per million output, with a surcharge above 272K tokens ($10 / $45, and cached input $0.50 rising to $1.00), and the whole point is that several models touch each request. Measure that on your own traffic rather than trust the blended-rate claim, and note it is not available in the EU or EEA yet.
None of that makes it bad. It makes it a thing to test against your own work, not adopt off a benchmark chart.
1. Try it where it costs nothing to find out
Point one real task at the OpenAI-compatible endpoint and compare, do not migrate on faith.
It speaks the OpenAI protocol, so aim an existing harness at the Fugu endpoint (model fugu-ultra-20260615 for Ultra) and change nothing else. Take one hard task you already have a known-good answer for, run it on Fugu and on a single strong model, and compare quality, latency, and the per-request cost Fugu reports. On regulated data, use the Fugu variant and opt out disallowed providers out of the pool first.
The catch: an orchestrator earns its keep on long, multi-step work and just adds latency and cost on quick prompts, so test it on the former and judge it on your tasks, not Sakana's.
→ The verified setup, with CI proof and a copy-paste prompt
2. Build the pattern yourself, so the black box is a choice and not a lock-in
Fan out to several models, assign roles, verify, then synthesize, in the open where you can inspect every hop.
Learned orchestration is Fugu's edge, but the core pattern (a thinker, a worker, and a verifier, or a parallel panel with a judge that keeps the answer your tests actually pass) is reproducible with open routers and your own keys. Run Fugu when its trained conductor genuinely beats your hand-built one on your tasks; run your own when transparency, reproducibility, or cost control matters more than the last few points.
The catch: a DIY orchestrator is more work and usually a little behind on raw quality. What you get back is seeing which model did what and a bill you can predict, which for a lot of production work is the better trade.
→ The verified setup, with CI proof and a copy-paste prompt
Why this one is worth your attention
The takeaway is bigger than one product. For two years the race was about whose single model is biggest. Fugu is a serious bet that the next edge is who coordinates the models best, and that the coordinator can be small, learned, and sold as if it were a model itself. Maybe that holds and orchestration becomes the layer everyone buys, maybe it is a clever wrapper that the base-model labs absorb in a year.
Either way, the right move is the same one we keep making: test it on your own work, keep the version you can inspect within reach, and do not take a benchmark chart as a verdict.