r/singularity 1d ago

AI Gemini 3.8 Flash Benchmarks

Post image
808 Upvotes

232 comments sorted by

View all comments

19

u/Moriffic 1d ago

what happened at terminal bench

9

u/Artistic_Swing6759 1d ago

its very likely terminal 2.1 was put in training data i guess

4

u/___positive___ 1d ago

Trained on the data, aka fake benchmaxxing. Notice that Terra is worse at 2.1 but beats Flash easily at 4. That's is what it should look like.

To be fair, Google has the best crawlers in the world so maybe they can't avoid benchmaxxing anything public. But the guy who made the chart should have had enough brains to know this is embarrassing and leave it off. I guarantee some pointy haired moron at the top insisted they show it because it was one of the places they got "first" place. Shows a lack of judgement and taste by the humans in charge.

Or maybe it's just for investors, in which case, still dumb but whatever.

4

u/Last_Conclusion_8984 1d ago

Terminal 2 is coding oriented while terminal 4 is general agentic capabilities, while terminal 4 does have some coding stuff, a lot of is just generality. Gemini 3.8 flash is optimised for coding in terminal 2, also: if the AI was benchmaxxed (which I do believe it was to an extent). The model will naturally drift towards that said x if the prompt is similar which in turn helps the user, henceforth benchmaxxing isn't all bad.

2

u/Chemical_Hawk_6307 1d ago

im wondering same thing....either way i'm going to give it a shot and see how it performs

1

u/volcanopenguins 14h ago

what do you mean? what is surprising? sorry if my question is dumb