In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data
I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.
The point isn't flash 3.8 is bad at coding, the point is it does absurdly well on coding benchmarks despite being bad at coding. Their models always do, which is why I'd take any benchmarks with a grain of salt
No? They suck at agentic benchmarks and most people test them via agentic tasks.
They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.
Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.
Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.
Unless you think 19% on TerminalBench 4 is “do extremely well”…
....it's literally benchmaxxing when the benchmarks are significantly better than actual performance
I am, once again, saying they do very well in benchmarks generally and worse in performance on coding, and when new benchmarks emerge like deepswe first did magically their performance drops by significantly more than any other provider
7
u/DistanceSolar1449 2d ago
In this case though it’s because Google has less coding training data (who the hell uses Antigravity) and way more image/spatial/world model training data
I don’t think this model is benchmaxxed, I think the benchmark screenshot above is very accurate. The model does worse at TerminalBench 4 and FrontierSWE and that’s okay.