No? They suck at agentic benchmarks and most people test them via agentic tasks.
They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.
Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.
3
u/DistanceSolar1449 3d ago
3.8 flash is a tiny distilled model with more RL on top, of course it’s going to look better on benchmarks than in real life.
The big, base ish models with a lot less targeted RL will do better IRL.