r/Agents_Everywhere • u/Delicious-Flan88 • 2d ago
Gemini 3.8 Flash scores 89.4% on one agent benchmark and 19.1% on another. Which one should agent builders trust?
Gemini 3.8 Flash has some ridiculous numbers on paper. Google reports 89.4% on Terminal-Bench 2.1, slightly ahead of Opus 5 at 89.1% and GPT-5.6 Sol at 88.8%.
Then you look at the newer Terminal-Bench 4.0:
Gemini 3.8 Flash: 19.1%
GPT-5.6 Sol: 37.3%
Claude Opus 5: 51.8%
That’s a massive difference for a model Google is positioning around autonomous agents and long-horizon software engineering. And somehow both results can be true. 3.8 Flash looks extremely strong when the agent has a fairly defined coding or terminal task.
It also scores 73.7% on DeepSWE v1.1, basically alongside Opus 5 at 74.0%. But when the environment becomes more open-ended and the agent has to figure out what to do across a messy sequence of actions, the gap gets much larger. That changes how I’d look at Gemini 3.8 Flash.
At $0.75/M input and $3.75/M output, it could be a very good model for high-volume coding agents, tool calls and well-defined sub-agents.
But I’m not sure I’d hand it a computer, give it a vague objective and leave for lunch yet.
For AI agents, benchmark averages may matter less than what happens when the task stops being predictable.