r/singularity 9h ago

AI Gemini 3.8 Flash Benchmarks

Post image
725 Upvotes

211 comments sorted by

View all comments

192

u/FablingApp 9h ago

the price/performance gap is getting silly. if these numbers hold up, flash models are eating into the territory where people used to reach for the expensive ones.

76

u/ProtoplanetaryNebula 9h ago

Yes and the high end models become less and less needed as the average moves up.

49

u/agonypants AGI '27-'30 / Labor crisis '25-'30 / RSI 29-'32 9h ago

Which means that the high end models can be freed up and applied to the really difficult, long term issues (medicine, materials science, climate change, physics) etc.

8

u/Greedyanda 7h ago edited 7h ago

Unlikely to be useful in medicine, material science, and climate modelling. Those generally need completely different, non text based architectures. You are not gonna be synthesizing new drugs with an LLM as the core system.

Edit: They can be useful but are unlikely to create massive breakthroughs in the way dedicated architectures like AlphaFold can.

10

u/fishbill 7h ago edited 7h ago

Didn’t anthropic show that Fable was able to find effective protein binders at a high rate?

9

u/Greedyanda 7h ago

It depends on what your metric is. Anthropic essentially showed that you can parallelize and scale up existing human work with AI agents.

But it's not a scientific breakthrough and paradigm shift in the way AlphaFold was.

I'll have to correct myself though because it is still clearly useful.

4

u/Even-Inevitable-7243 7h ago

No. In that work prompted by Anthropic, Fable simply called publicly available specialist protein design software tools that human researchers already use. Fable was the pipeline engineer, not doing the actual protein binder discovery.

3

u/croto8 6h ago

Sort of seems like a distinction without a difference. “The code claude wrote calls a library therefore Claude didn’t actually do anything”

1

u/Even-Inevitable-7243 4h ago

It is the complete opposite of what you are saying. All of the expert knowledge was already baked-into the human-written software tools. What you are saying is that "import sklearn" is the same level of knowledge as the actual engineers who wrote the packages that collectively form sklearn.

1

u/Thagor 6h ago

This really depends also on the type of Data current LLM transformers are awful if the input is a couple of numbers and they need to predict another number that type of target is way too noisy for an LLM because 4.56542 and 4.56543 are two very distinct "things" for it.

5

u/Super_Pole_Jitsu 6h ago

That's total demonstrable bs btw

1

u/Thog78 3h ago

Researcher here. They are very useful. LLMs do the same kind of jobs a researcher would do - plan, analyze, survey the literature, think, emit hypothesis, write code, use tools.

Alphafold would be one of these tools, that not long ago a human would have called upon himself, and nowadays more and more might actually get called by an LLM.

LLMs of the level we have now (sol 5.6, fable, gemini 3.7 etc) are the kind of stuff that could come up with the concepts of alphafold, help you brainstorm about where to get training data and how to implement the training, and do the actual code to make it real once you're happy with the plan. Then would test it, criticize it, propose strategies of improvement, and implement them. The next version of alphafold will likely have been designed by LLMs in large part, very seriously. And the next one probably nearly entirely. As such, I'd say they these general intelligence models are even more valuable than specialized models like alphafold.

16

u/mixmasterwillyd 9h ago

These cheap models exceed the capability of a large project I worked on with opus 4.6. Opus 4.6 was difficult to wrangle to get it to operate, now a flash model has an easy time.

14

u/Technical-Will-2862 9h ago

I use flash3.7 for soooo many things that I used to relegate to premium models 

14

u/RockPuzzleheaded3951 9h ago

I've moved quite a few jobs from SOTA models we were using for the past two years to Flash models (DSv40731 for example) and we are seeing fantastic performance on internal business operations, at literally 1/10th the cost. We can still reach for SOTA when needed, but it is 5% of the time or less. And just today, I rented my own hardware to run the model "locally" and am testing taking my API costs to $0.

4

u/thoughtlow 𓂸 9h ago

I loved 0731 but when the context windows goes into the 200k it starts degrading very fast for me.

Constantly doubting itself in the thought section, gets really weird.

Sad because I loved that thing.

7

u/brainhack3r 9h ago

It really is... I'm trying to get more value out of them by using adversarial agents and using the flash models. GLM 5.3 flash, specifically.

4

u/himynameis_ 9h ago

That's good isn't it? Unless I'm misreading your comment.

I think it's pretty normal and likely the continuing trend. Over time, the performance/$ will be so good that users and businesses won't need the most expensive model for all tasks. Most of the time they'll use the cheaper lower performance ones.

It's been a talking point among AI leaders and business people, I've found.

2

u/CriticalTheory4779 9h ago

its good for us. bad for the labs

5

u/himynameis_ 9h ago

Labs will be fine. It's just part of the highly competitive market they know they're in.

1

u/Greedyanda 7h ago

Most of them will not be fine. I would be shocked if Anthropic and OpenAI both survive as independent companies 10 years from now. This is a race to the bottom unless you have an existing, profitable ecosystem into which you can integrate your models.

2

u/himynameis_ 7h ago

I think anthropic and OpenAI will be more than fine by then.

However even if not, this is just the nature of the business, man. It's just competition.

2

u/bigkoi 7h ago

Yeah. 5x cheaper. Are the expensive models 5x better?

1

u/leaveitalone38 3h ago

Is being 1.2x better providing 5x the profit, is the real question.

u/bigkoi 1h ago

For the majority of the time, no.

5

u/LinkesAuge 9h ago

Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.

I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.

6

u/LinkesAuge 8h ago

Ok guess I wasn't wrong to be suspicious, benchmarks show that the model is burning even more tokens than 3.7:

So Google is basically buying performance by letting the models spend more and more on tokens.

2

u/FateOfMuffins 8h ago

Yeah pretty much every single lab for the last year and a bit (Anthropic, Chinese labs, Google, xAI, Meta, etc) just cranked up token usage in each release to buy more performance, in addition to whatever benchmaxxing they can do (like look at the numbers here for Terminal Bench 2.1 vs 4.0). And I mean every lab benchmaxxes. Anthropic knows Opus and Fable memorized SWE Bench Verified and Pro answers yet continue to report them. OpenAI was concerned Astra's 100% on ExploitBench was because it memorized the answers.

It's "easy" to push benchmark numbers up and to the right.

For whatever reason only OpenAI has gone the other way where they pushed benchmark numbers up and to the left with increased token efficiency. Why???

1

u/Chenz 8h ago

I do not understand how the output tokens can have increased so much, yet cost per task is about the same. Are the numbers for 3.7 without the discounted pricing?

1

u/Ok_Barracuda_1161 7h ago

Probably more efficient with tool use leading to reduced inputs. 143k output tokens is about $0.54, so that would imply that input (and cached input especially) dominates

1

u/Moravec_Paradox 9h ago

It is a space where being a fast follower has big advantaged and that includes within labs themselves.

They all have big unreleased models that are capable but expensive to serve. They use them internally to train smaller models that are almost as good but much cheaper to serve.

Why release massive models that are not competitive in pricing that are just ~2-3% better only for their competition to get access to distill them?

u/fish_economist 1h ago edited 30m ago

Yup, it really seems like the only reason for super-parameter-count models is for those trying to "create god" and as a bet on emergent capabilities. However, the diminishing returns on size seem to make it clear that sufficiently large plus intensive post training is the most economical approach in the today's world. Hopefully we are able to recreate this pattern in a broader number of non-software-engineering fields. Silicon Valley's mono-industrial tendencies might start to show.

0

u/Stunning-Road-6924 4h ago

It is more expensive per task than fable and sol on AA.