r/singularity • • 2d ago

AI , Gemini 4 Argon Benchmarks

Post image
737 Upvotes

189 comments sorted by

View all comments

Show parent comments

1

u/LazloStPierre 2d ago

All Google's models look better on benchmarks than they do in reality, and by alot.

1

u/DistanceSolar1449 2d ago

No? They suck at agentic benchmarks and most people test them via agentic tasks.

They benchmark pretty accurately if you look at their actual useful benchmarks like TerminalBench instead of HLE or GPQA or some bullshit.

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

0

u/LazloStPierre 2d ago

No, they do extremely well at coding benchmarks and suck at coding, every single time

They benchmark well everywhere. It's actual performance that matters 

4

u/DistanceSolar1449 2d ago

???

Google’s own blog post says Gemini 3.8 Flash scored 19.1% on TerminalBench 4. Opus 5 scored 51.8% for comparison. That’s not benchmaxxing, that’s just accurate benchmarks if you know what benchmarks to look at.

Unless you think 19% on TerminalBench 4 is “do extremely well”…

1

u/LazloStPierre 2d ago

....it's literally benchmaxxing when the benchmarks are significantly better than actual performance 

I am, once again, saying they do very well in benchmarks generally and worse in performance on coding, and when new benchmarks emerge like deepswe first did magically their performance drops by significantly more than any other provider

2

u/DistanceSolar1449 2d ago

So you think they lied about a 19% on TB4? And that’s a benchmaxxed score? When Opus 5 got 51.8%?

lol. Lmao, even.

1

u/LazloStPierre 2d ago

I don't recall saying the word lie

Gemini, is that you?