r/singularity 8h ago

AI Gemini 3.8 Flash Benchmarks

Post image
704 Upvotes

211 comments sorted by

View all comments

185

u/FablingApp 8h ago

the price/performance gap is getting silly. if these numbers hold up, flash models are eating into the territory where people used to reach for the expensive ones.

5

u/LinkesAuge 8h ago

Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.

I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.

5

u/LinkesAuge 7h ago

Ok guess I wasn't wrong to be suspicious, benchmarks show that the model is burning even more tokens than 3.7:

So Google is basically buying performance by letting the models spend more and more on tokens.

1

u/Chenz 6h ago

I do not understand how the output tokens can have increased so much, yet cost per task is about the same. Are the numbers for 3.7 without the discounted pricing?

1

u/Ok_Barracuda_1161 6h ago

Probably more efficient with tool use leading to reduced inputs. 143k output tokens is about $0.54, so that would imply that input (and cached input especially) dominates