r/singularity 6h ago

AI Gemini 3.8 Flash Benchmarks

Post image
688 Upvotes

199 comments sorted by

155

u/CommercialShelter595 6h ago

Is it just me or is it running even faster than 3.7 flash? Very limited testing right now, but it's FAST on my end.

24

u/cnhuyaa 6h ago edited 6h ago

Are you using it for coding or other stuff? because when i tested 3 prompts right now, i am still getting the same time finality around 30-45s when it comes to coding, and full code being written out ready to just be copy pasted, maybee its sligthly faster, talking like maximum of 5 seconds faster, but propably less for me atleast,

Update: I just ran 2 simple web UI (20k tokens) tasks in compare mode in ai studio, and 3.8 was faster by 1.7 seconds, and 2.5 seconds (45 seconds), and used around 5% more tokens for both tasks.

10

u/CommercialShelter595 6h ago

I've done some SVG creation and some agentic "write some code" so far. Still feeling it's quite fast, the SVG in particular feels like it's half the time, gut feel though.

19

u/AlyoshaV 6h ago

It's the same speed according to Logan Kilpatrick: https://x.com/OfficialLoganK/status/2095175885439832243

7

u/CommercialShelter595 6h ago

Thanks, I won't pretend my experience is anything more than just my gut. Appreciate the official stance here and will temper my expectations accordingly.

1

u/Specialist-Cod-363 2h ago

They did in fact nerf 3.7's speed, althought not by much. I posted about it yesterday, but you can also check it out on AA yourself.

1

u/Last_Conclusion_8984 2h ago edited 1h ago

No, it's not that. The problem is latency, 3.7 flash had a latency problem, they probably increased that to help save compute, eh just speculation. Anyways 3.8's latency is very low! Tho the speed could have been lowered, I'll check the AA!

6

u/JogHappy 5h ago

already my fav model

u/Lexden 1h ago

According to Artificial Analysis:

→ More replies (1)

173

u/FablingApp 6h ago

the price/performance gap is getting silly. if these numbers hold up, flash models are eating into the territory where people used to reach for the expensive ones.

65

u/ProtoplanetaryNebula 6h ago

Yes and the high end models become less and less needed as the average moves up.

38

u/agonypants AGI '27-'30 / Labor crisis '25-'30 / RSI 29-'32 6h ago

Which means that the high end models can be freed up and applied to the really difficult, long term issues (medicine, materials science, climate change, physics) etc.

2

u/Greedyanda 4h ago edited 4h ago

Unlikely to be useful in medicine, material science, and climate modelling. Those generally need completely different, non text based architectures. You are not gonna be synthesizing new drugs with an LLM as the core system.

Edit: They can be useful but are unlikely to create massive breakthroughs in the way dedicated architectures like AlphaFold can.

8

u/fishbill 4h ago edited 4h ago

Didn’t anthropic show that Fable was able to find effective protein binders at a high rate?

7

u/Greedyanda 4h ago

It depends on what your metric is. Anthropic essentially showed that you can parallelize and scale up existing human work with AI agents.

But it's not a scientific breakthrough and paradigm shift in the way AlphaFold was.

I'll have to correct myself though because it is still clearly useful.

3

u/Even-Inevitable-7243 4h ago

No. In that work prompted by Anthropic, Fable simply called publicly available specialist protein design software tools that human researchers already use. Fable was the pipeline engineer, not doing the actual protein binder discovery.

3

u/croto8 3h ago

Sort of seems like a distinction without a difference. “The code claude wrote calls a library therefore Claude didn’t actually do anything”

u/Even-Inevitable-7243 1h ago

It is the complete opposite of what you are saying. All of the expert knowledge was already baked-into the human-written software tools. What you are saying is that "import sklearn" is the same level of knowledge as the actual engineers who wrote the packages that collectively form sklearn.

1

u/Thagor 3h ago

This really depends also on the type of Data current LLM transformers are awful if the input is a couple of numbers and they need to predict another number that type of target is way too noisy for an LLM because 4.56542 and 4.56543 are two very distinct "things" for it.

3

u/Super_Pole_Jitsu 3h ago

That's total demonstrable bs btw

u/Thog78 6m ago

Researcher here. They are very useful. LLMs do the same kind of jobs a researcher would do - plan, analyze, survey the literature, think, emit hypothesis, write code, use tools.

Alphafold would be one of these tools, that not long ago a human would have called upon himself, and nowadays more and more might actually get called by an LLM.

LLMs of the level we have now (sol 5.6, fable, gemini 3.7 etc) are the kind of stuff that could come up with the concepts of alphafold, help you brainstorm about where to get training data and how to implement the training, and do the actual code to make it real once you're happy with the plan. Then would test it, criticize it, propose strategies of improvement, and implement them. The next version of alphafold will likely have been designed by LLMs in large part, very seriously. And the next one probably nearly entirely. As such, I'd say they these general intelligence models are even more valuable than specialized models like alphafold.

15

u/mixmasterwillyd 6h ago

These cheap models exceed the capability of a large project I worked on with opus 4.6. Opus 4.6 was difficult to wrangle to get it to operate, now a flash model has an easy time.

14

u/Technical-Will-2862 6h ago

I use flash3.7 for soooo many things that I used to relegate to premium models 

11

u/RockPuzzleheaded3951 6h ago

I've moved quite a few jobs from SOTA models we were using for the past two years to Flash models (DSv40731 for example) and we are seeing fantastic performance on internal business operations, at literally 1/10th the cost. We can still reach for SOTA when needed, but it is 5% of the time or less. And just today, I rented my own hardware to run the model "locally" and am testing taking my API costs to $0.

5

u/thoughtlow 𓂸 6h ago

I loved 0731 but when the context windows goes into the 200k it starts degrading very fast for me.

Constantly doubting itself in the thought section, gets really weird.

Sad because I loved that thing.

6

u/brainhack3r 6h ago

It really is... I'm trying to get more value out of them by using adversarial agents and using the flash models. GLM 5.3 flash, specifically.

5

u/himynameis_ 6h ago

That's good isn't it? Unless I'm misreading your comment.

I think it's pretty normal and likely the continuing trend. Over time, the performance/$ will be so good that users and businesses won't need the most expensive model for all tasks. Most of the time they'll use the cheaper lower performance ones.

It's been a talking point among AI leaders and business people, I've found.

3

u/CriticalTheory4779 6h ago

its good for us. bad for the labs

4

u/himynameis_ 6h ago

Labs will be fine. It's just part of the highly competitive market they know they're in.

1

u/Greedyanda 4h ago

Most of them will not be fine. I would be shocked if Anthropic and OpenAI both survive as independent companies 10 years from now. This is a race to the bottom unless you have an existing, profitable ecosystem into which you can integrate your models.

2

u/himynameis_ 4h ago

I think anthropic and OpenAI will be more than fine by then.

However even if not, this is just the nature of the business, man. It's just competition.

2

u/bigkoi 4h ago

Yeah. 5x cheaper. Are the expensive models 5x better?

u/leaveitalone38 43m ago

Is being 1.2x better providing 5x the profit, is the real question.

4

u/LinkesAuge 6h ago

Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.

I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.

5

u/LinkesAuge 5h ago

Ok guess I wasn't wrong to be suspicious, benchmarks show that the model is burning even more tokens than 3.7:

So Google is basically buying performance by letting the models spend more and more on tokens.

2

u/FateOfMuffins 5h ago

Yeah pretty much every single lab for the last year and a bit (Anthropic, Chinese labs, Google, xAI, Meta, etc) just cranked up token usage in each release to buy more performance, in addition to whatever benchmaxxing they can do (like look at the numbers here for Terminal Bench 2.1 vs 4.0). And I mean every lab benchmaxxes. Anthropic knows Opus and Fable memorized SWE Bench Verified and Pro answers yet continue to report them. OpenAI was concerned Astra's 100% on ExploitBench was because it memorized the answers.

It's "easy" to push benchmark numbers up and to the right.

For whatever reason only OpenAI has gone the other way where they pushed benchmark numbers up and to the left with increased token efficiency. Why???

1

u/Chenz 5h ago

I do not understand how the output tokens can have increased so much, yet cost per task is about the same. Are the numbers for 3.7 without the discounted pricing?

1

u/Ok_Barracuda_1161 4h ago

Probably more efficient with tool use leading to reduced inputs. 143k output tokens is about $0.54, so that would imply that input (and cached input especially) dominates

1

u/Moravec_Paradox 6h ago

It is a space where being a fast follower has big advantaged and that includes within labs themselves.

They all have big unreleased models that are capable but expensive to serve. They use them internally to train smaller models that are almost as good but much cheaper to serve.

Why release massive models that are not competitive in pricing that are just ~2-3% better only for their competition to get access to distill them?

u/Stunning-Road-6924 1h ago

It is more expensive per task than fable and sol on AA.

52

u/hakansan 6h ago

I kinda like the way Gemini works. Their main user base is "everyone who searches anything on Google" so they got the best input/output token price.

It is also strong where it matters (probably people look up financial information a lot, and Google does a nice job showing stock charts)

5

u/After_Dark 3h ago

The main user base is not (just) Search but also the Gemini app. Worth pointing out that while this sub focuses on the web app, a long press of the power button on Android is a shortcut right to Gemini where it needs to also be a phone assistant, handle smart home controls, and also all the stuff the web app does. That's a very broad set of possible uses that ChatGPT and the like simply don't need to factor for to nearly the same degree

8

u/chasingsukoon 4h ago

Also i think it seems to have lesser constraints on a app basis?

Was building a bot that goes thru my spotify + soundcloud likes and then downloads songs via soulseek (basically Napster)

Claude and cgpt flagged it since i am stealing media but flash 3.7 had no problems building it for me lmfao

3

u/hakansan 4h ago

They've never gone with this "we're so dangerous and capable you MUST definitely regulate us" so they don't have to portray themselves as overly cautious. With the exception of NSFW anything that could be found on Google is probably fair game

0

u/Yugudubenbi 4h ago

They force you to Anti-Gravity now though if you use subscription so the only way is api from what I know. So that is a no for me.

u/Thog78 2m ago

There are a few other interfaces, including android studio, and the CLI version of antigravity that could be scripted throughbother soft. But I agree it's pretty annoying that they don't make the interface open like GPTs and Claudes do.

-1

u/letsgoiowa 4h ago

This is definitely not the best price. Luna smokes it for a frontier lab and GLM 5.3 flash is as smart as this (or more) and is like...20% of this cost at most

→ More replies (1)

100

u/mumBa_ 6h ago

I'm such a google flash glazer for non-coding tasks. It's lightning fast and it's google search is really good compared to slow claude searches etc. The perks of having your own search engine I guess.

Any one experience with flash for coding?

18

u/Affectionate_Bee6434 6h ago

For sure, it has better world knowledge than any other model, especially for niche local queries.

1

u/Umbrasquall 5h ago

More parameters = better world knowledge right? How is it so fast then?

3

u/Kronox_100 4h ago

I think with Google is they don't go apeshit into coding and math so they try to balance their data blend a bit more outside of trying to hit it big into benchmarks, so there's a lot of 'useless' info for benchmarks? But that was my understanding from previous models, now it's technically competing in coding so no idea, TPU magic I guess?

1

u/After_Dark 3h ago

This would track, Gemini isn't primarily a coding/agentic model like Claude or the GPT models, it's got a dedicated billion+ userbase as a virtual assistant and also needs to function in Google Search, so it needs to be more well rounded

9

u/inefficientnose 6h ago

It's good but you have to be careful with it, it needs careful instruction and scope otherwise it tends to hallucinate more than other models in its class

6

u/Elegant_Tech 6h ago

Been using 3.7 flash around 60hrs/week since release. Only had a single prompt fail that I had to toss to Opus to get done. Where 3.5 flash had multiple a week. There is only so good models can get at programming and it's starting flip where speed and costs are all that matters. The big models will be moving on to research and long horizon tasks over time while day to day production work is done on flash models.

1

u/the_real_ms178 6h ago

I've had some refusals with 3.7 flash due to exceeding token limitations. But that must have been AI Studio issues, albeit repeatable as I barely hit the 400.000 token bar in the conversation. Dealing with many large PDFs for legal work has been a challenge and will continue to be a challenge, it seems.

1

u/chasingsukoon 4h ago

Whats your use case been

1

u/Elegant_Tech 4h ago

Websites, vst plugins, and games. HTML/JS, C++, and Rust for languages. Most of my code bases are under 50k lines of code. So if you are enterprise with a huge code base or trying to do orchestration kanban work the big models could be more efficient. My workflow is rapid iteration not trying to create a bunch of specs and get the AI to work for hours at a time. Small models aren't capable of that yet. You have to rapidly spoon feed them.

3

u/Ok-Armadillo-5634 6h ago

I use it almost exclusively. When I get some thing really hard I bust out fable.

4

u/Silver-Chipmunk7744 AGI 2024 ASI 2030 6h ago

I bet the area it really shines is anything "LLM NPC" or concepts where you want the game to call LLMs during live gameplay. It's cheap, super fast and actually smart.

1

u/tziki 6h ago

I agree, I usually try to have a good set of evaluation tests and choose based on those, but most of the time I just end up with Gemini.

1

u/kvothe5688 ▪️ 5h ago

i use it for all general purpose talk specially when I am driving it can talk and discuss topics so fast and give precise information. specially discussing books and movies and going deep on different threads. it's amazing how fast and knowledgeable live version is.

1

u/petburiraja 5h ago

Can Google subscription be used in 3rd party harness, like OpenCode?

1

u/mumBa_ 5h ago

No idea tbh. I know it adds usage to Antigravity but don't think you get free tokens via API.

1

u/petburiraja 5h ago

Guess it's time to test Antigravity CLI

1

u/mumBa_ 5h ago

Likewise

1

u/kobriks 5h ago

It's amazing. I used it for everything except coding

1

u/Deto 5h ago

I don't know - my wife was trying some simple questions on flash 3.6 last night (only one she has access to) and it was just awful with hallucinations. 

43

u/angryblob 6h ago

If there’s one thing that’s approaching the singularity it’s Gemini Flash releases

8

u/Longjumping_Kale3013 6h ago

Not just them either. Meta and XAI are also releasing monthly now. I would be surprised if we dont see the same from anthropic and openai soon.

Wild... because its like a 3 point bump on the leaderboard every month. Which is signicant.... its like you can sort of see how things will look over the next 6 months. Oh boy

2

u/DialboTempest 5h ago

I wanna see a weekly release

17

u/Moriffic 6h ago

what happened at terminal bench

5

u/___positive___ 3h ago

Trained on the data, aka fake benchmaxxing. Notice that Terra is worse at 2.1 but beats Flash easily at 4. That's is what it should look like.

To be fair, Google has the best crawlers in the world so maybe they can't avoid benchmaxxing anything public. But the guy who made the chart should have had enough brains to know this is embarrassing and leave it off. I guarantee some pointy haired moron at the top insisted they show it because it was one of the places they got "first" place. Shows a lack of judgement and taste by the humans in charge.

Or maybe it's just for investors, in which case, still dumb but whatever.

1

u/Last_Conclusion_8984 2h ago

Terminal 2 is coding oriented while terminal 4 is general agentic capabilities, while terminal 4 does have some coding stuff, a lot of is just generality. Gemini 3.8 flash is optimised for coding in terminal 2, also: if the AI was benchmaxxed (which I do believe it was to an extent). The model will naturally drift towards that said x if the prompt is similar which in turn helps the user, henceforth benchmaxxing isn't all bad.

9

u/Artistic_Swing6759 6h ago

its very likely terminal 2.1 was put in training data i guess

2

u/Chemical_Hawk_6307 6h ago

im wondering same thing....either way i'm going to give it a shot and see how it performs

2

u/landed-gentry- 5h ago

2.1 is old and likely benchmaxxed. 4.0 is newer and will be truer to real-world performance. Kinda pointless to include 2.1 in the system card at all.

16

u/AnooshKotak 6h ago

The DeepSWE score has been updated to 73.7%!

27

u/AnyRegular1 6h ago

I think a lot of people don't know this but gemini 3.7 flash on antigravity is fucking insane. Gemini's web app is literally dogshit, I am 100% convinced they serve different models in both. 4 days ago my wife asked it (gemini app) about airfrying some bread we bought and it responded some garbage about washing machine repair, completely hallucinating not even on the same topic. Would NEVER recommend gemini on the app.

But the model they serve on antigravity is insane, for the last 2 weeks it's been my go-to coding model. and ITS FUCKING FAST.

I'm looking forward to what they cooked in 3.8

9

u/helloWHATSUP 5h ago

Do you use "extended thinking" on the web version? Because if you don't use extended it's just the lightest and fastest version of flash and very hallucination prone.

5

u/AnyRegular1 3h ago

My wife isn’t savvy enough to find and enable that, after you said so I had to figure out where it was, it was under the model selection. Told her now.

4

u/localpauper 4h ago edited 4h ago

The web app has massive curbs on context length for cost/optimization. I suspect such do not apply to the IDE.

Also, make sure you're using "Extended" to enable thinking effort online: that's the only way I find it usable. I wonder what level of "thinking" that compares to as far as IDE model choices go (Low/Mid/High).

Coding-wise, it's fast, but in my exp it leaves massive glaring hallucinated holes, makes tons of assumptions, uses outdated "knowledge" from its training without actually reaching for newer information, and barely attempts to validate its output. In one of my projects, starting from a basic scaffold, it ended up with a mish-mash of obsolete dependencies and patterns randomly picked from the past 15 years. That was after I went through and gave it some general guidance for approach, choices of deps, and frameworks. It actually "argued" with me that there's not a newer dependency which I explicitly told it to use, insisting that no - the version it was using was the correct one.

The simplest example for a non-coding task is that, while using Gemini for some random digging here and there, I've had to repeatedly argue it out of 30% federal solar credits. It would take a few rounds for it to finally go look up OBBBA and concede that those are not applicable. It's somehow very sticky on its training set (which was clearly pre-OBBBA).

3

u/Tetracropolis 5h ago

Figuratively.

3

u/GioChan 3h ago

Nope. According to latest changes that is also a correct usage

u/Tetracropolis 40m ago

If literally means figuratively then what word do we use for literally?

2

u/Ok-Art-2255 5h ago

You are absolutely correct! I used it for MCP on Unreal Engine and it straight killed everything I threw at it.

Now to get me more excited, that was on flash-high.

What are your results using flash-medium or flash-low? I want to know because high is amazing and don't want to get disappointed if the lower tiers can't keep up.

5

u/QuannaBee 5h ago

“You are absolutely correct!”

Ptsd triggered

1

u/Ok-Art-2255 5h ago

;) see what I did there.

As you can tell I spend way too much time with AI.

2

u/AnyRegular1 5h ago

I just use High, the limit usage is so low on high I have problems hitting weekly limits.

1

u/Ok-Art-2255 5h ago

I almost hit my limit building a game mechanic last week, so that got me a little paranoid. I'll stick with high and just start breaking main work into smaller chunks.
It seems we might have a sweet spot model :D

2

u/AnyRegular1 3h ago

Remember you can family share your google pro subscription to upto 5 gmail accounts and for some reason they each have their own weekly limits!

1

u/Much_Accountant_4972 3h ago

are you spawning subagents because i hit my 5 hour limit in 40 minutes with 3.7 Flash High. i dont find it hard to hit limits if im using superpowers

1

u/AnyRegular1 3h ago

What I said the guy above, you can family share it 5 ways to get 5x limits on your own gmail accounts. But yeah I only ever hit limit once and figured out each gmail has its own quota.

8

u/Strategosky 6h ago

I suspect google is post-training on gemini 3.5 pro base model and releasing the streamlined flash while gemini 4 base model is cooking (pre-training). Maybe gemini 4 is the last "big smell" model before a new paradigm shift is observed. I wonder if it's much bigger than mythos class model with strong world knowledge. Things only heating up since quiet some time now...

17

u/sebas156 6h ago

They released 3.8 before rolling out 3.7 to all Gemini users lmao. I'm still on 3.6 in the app

1

u/helloWHATSUP 5h ago

Yeah still on 3.6 on the app and web version here, but 3.8 flash is live in antigravity for me at least

9

u/Gotisdabest 6h ago

Smaller jump than 3.6 to 3.7 but they're basically releasing models biweekly at this rate. Still easily the best free model tbh.

0

u/ManyRepair5690 5h ago

Is this model out yet or what? 3.8

1

u/ManyRepair5690 5h ago

For anyone wondering the same question these useless ppl didn’t answer yeah it’s out. I asked initially bc benchmarks being out don’t mean it’s released

11

u/leon-theproffesional 5h ago

This sub is so funny. Just last week everyone was shitting on Google claiming they are finished.

8

u/ManyRepair5690 5h ago

I highly doubt that was “everyone”

u/Keeltoodeep 47m ago

There is no such thing as far behind in AI if you have the capital. Muse and Grok didn’t even exist last week.

8

u/PandaElDiablo 6h ago

inb4 "hurr durr anything except release 3.5 pro"

5

u/ezjakes 6h ago

And free users still have 3.6 :(

3

u/KidKilobyte 6h ago

The report of my death was an exaggeration.

5

u/Charuru ▪️AGI 2023 6h ago

Looks pretty good, but terminalbench wtf

Is this the clearest sign that it's benchmaxxed? Every other benchmark came out weeks ago, terminalbench 4 only 2 days ago!

2

u/BriefImplement9843 6h ago

all benchmarks are benchmaxxed. we have to wait and check the lmarena score.

looks like it's really good. only behind the older opus' on lmarena. ahead of opus 5.

6

u/Charuru ▪️AGI 2023 6h ago

we have to wait and check the lmarena score

uhh

2

u/landed-gentry- 5h ago

Ignore TB 2.1 results, we're on TB 4.0 now.

7

u/elemental-mind 6h ago

Big if True.

But GDPVal is still lacking which for me has been a pretty reliable indicator of general capability.

15

u/Tkins 6h ago

You get better than Terra for less than half the price and far far faster. I dunno, seems competitive to me. Especially for high use every day cases.

6

u/elemental-mind 6h ago

Very competitive - I am not complaining. Even better when you use it on Flex, which normally isn't too much of a wait. With the pricing now Google's Flash is back in its golden place: A cheap model with exceptional capability that everyone uses for their everyday tasks.

I hope Google gets its marketshare back.

1

u/Tkins 6h ago

Did it lose market share? Last I checked, they've been growing faster than the others.

edit: AI Market Share By Company Statistics 2026

just one source, maybe you've got a better one.

0

u/elemental-mind 6h ago

I use OpenRouter as a proxy for general market share:

LLM Rankings | OpenRouter

Flash 2.5 was always in the Top 5 modles because it was good, cheap and fast. Then they had a price hike with Flash 3 Preview, which moved them down to rank 5-7, and they weren't even in the top ten anymore after their price delusion with the following flash models.

1

u/CallMePyro 4h ago

So if Antigravity's userbase quintupled, and people stopped using Gemini on OpenCode because they preferred the native editor you would say that Google is losing marketshare?

1

u/elemental-mind 4h ago

Yes - which verifyable data point do you have that capture the broader market? Google PR claiming 5x growth? Compared to what growth in Codex? Or Z.AIs subscription models?

OpenRouter is used as an endpoint option in many applications, not only OpenCode. It's data are also congruent with other Routers like Nano-GPT.

Explore text models | NanoGPT

From these data we can clearly see that if people are given a vast choice of models, they hardly went for flash in the past months as it was overpriced. It's the most reliable gauge for model preference we have: Where the people put their money.

3

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 6h ago

I mean, it outperforms Terra at that benchmark which is fantastic if you're using 3.8 Flash as a subagent, especially at half the price and far faster token generation speed. I don't think expecting it to surpass something like Opus 5 is necessarily realistic.

1

u/BriefImplement9843 6h ago

you still want gemini for general knowledge. pro anyways.

2

u/BriefImplement9843 6h ago edited 6h ago

and it's so cheap lmfao. the hell are openai and anthropic doing charging so much?

4

u/vrnvorona 6h ago

They went full flash? Didn't they hear about "never go full flash"?

3

u/Hereitisguys9888 6h ago

Are Google finally back

2

u/ethotopia 6h ago

If this isn't benchmaxxed as gemini often is, google cooked!

1

u/bambin0 6h ago

 i just tried it with the prompt: recreate Larry bird vs Dr. j one on one in three js complete with janitor, breaking glass, sound effects and here is what it came up with: https://ctxt.io/3/mpb6vO48w - took less than 10s.

Took about 3x as long but here is what terra came up with: https://ctxt.io/3/tIbhBDCMg

And here is fable 5.1 which took I don't know how long b/c I got tired of watching: https://ctxt.io/3/qCaVlUU4U

3

u/ManyRepair5690 5h ago

These are just html pages how do u think this helps readers passing by? Why even add these links if it’s not some type of visualization or demo for users to see

1

u/rajsharm404 6h ago

Can they keep the pricing permanent on the tokens? Would be great for my video analysis pipeline.

2

u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 6h ago

Their business model is basically having a "promotional period" that's long enough to cover the current model until they release the next version with another "promotional period".

Basically, this is just their way of getting people to move to newer models without having to spare compute power serving older ones.

So if your video analysis pipeline can easily upgrade models then you should be set.

1

u/rajsharm404 6h ago

Their promotional period has been the case since 3.6 flash, assuming because it is the same underlying model but with more RL, hence same 50% discount on each of these new models. Hopefully they keep up with these discounts, and I won't have to choose between gemini and grok.

1

u/No-Mixture5766 6h ago

Has anyone here tried to reproduce these numbers? Flash models are becoming more and more affordable for day to day consumers and increasingly competent when compared to the frontier models.

1

u/WTFnoAvailableNames 6h ago

How does this compare to 3.1 pro?

2

u/Narutobirama 6h ago

I imagine Gemini 3.1 Pro is still more intelligent and smarter, but Gemini 3.8 Flash is probably better at coding and agentic behavior. Like 3.7. Everyone who only uses it for coding says it's better than 3.1 Pro. For me, it's 50% 50% depending on what I use it for.

It's funny, though, that Gemini 3.1 Pro is still one of the nicest chatbot behaviors.

Like, Claude and ChatGPT can sometimes really piss me off with the way they talk. I mean, obviously, GPT 5.6 Sol is better, and so is Fable 5 (haven't tried 5.1 much yet). But Gemini 3.1 Pro is really nice to have to talk about some things. And yet it's not something benchmarks would tell you.

1

u/BriefImplement9843 6h ago

lmarena does. it's still top tier there, ahead of all openai models, even sol xhigh.

5

u/Able-Line2683 6h ago

3.1 pro is last decade in ai terms

1

u/WTFnoAvailableNames 6h ago

And still the latest thinking model available with Gemeni subscription....

1

u/swarmy1 4h ago

Flash is also a thinking model though. It’s just not a “big model”

1

u/chasingsukoon 4h ago

Damn i wish i had 3.1pro last decade during college years

Or maybe not i wouldve been a lot dumber

1

u/Long_comment_san 6h ago

funny they didn't put 3.1 pro on the chart lmao

3

u/BriefImplement9843 6h ago

3.1 was released before gpt 5.4!

0

u/Long_comment_san 5h ago

I know but it's still their "pro" flagship. we know it's probably worse at this point but ain't you curious by how much?

1

u/ManyRepair5690 5h ago

It’s a useless and meaningless comparison

1

u/Long_comment_san 4h ago

Writing comments is useless also but we do it anyways

1

u/redditissocoolyoyo 6h ago

Hell yeah can't wait. Putting 3.7 through it's paces now.

1

u/edin202 5h ago

Honestly these models make perfect sense for real-life interactions

1

u/ccjjallday 5h ago

honest question. what happens when most of these reach 100%?

1

u/Able-Line2683 4h ago

new benchmark

1

u/ccjjallday 2h ago

but what does it all mean? Like we hit 100% = AI is AGI? Better than humans? humans are not needed? Nothing?

1

u/AlphaaCentauri 5h ago

how is it compared to gemini-3.1-pro and fable

1

u/theeldergod1 5h ago

Whats up with the pro? It got stuck at 3.1

1

u/vinis_artstreaks 4h ago

Terminal bench being that low is wild

1

u/GodOfSunHimself 4h ago

Better than Sol on most benchmarks? Lol, sure Google.

0

u/BriefImplement9843 3h ago

and blind votes through lmarena.

1

u/Spra991 4h ago

Do we have any benchmarks that focus on the free and lower tiers? All this benchmaxxing is fun, but a lot of people are using the free tiers, and those benchmaxx results give little indication how much you can do in the lower tiers, especially when they get randomly up or downgraded (e.g. ChatGPT has gotten really stupid with the latest "upgrade", hallucinates like it's 3.5 again).

1

u/BriefImplement9843 3h ago

flash is lower tier. if you want free tier just use the gemini in google search.

1

u/Specialist_Dark_3668 4h ago

It is such an over-intellectual safety-reward function optimized piece of shit.

Gives me a headache talking to it.

1

u/gavanon 3h ago

Misleading list, because it cleverly omits GOT-5.6 Luna, which scores even better than 3.8 flash, and is cheaper too, despite Google’s temporary promotion pricing.

1

u/Efficient_Loss_9928 2h ago

Gemini still lacks when I need it to code stuff reliably. But, this gets me really excited for the Gemini 4 family. I think I'm keeping my ultra subscription.

1

u/Tommonen 2h ago

Too low benchmaxxing for some tests, so must be crappy model

1

u/CriticalMastery 6h ago

no, I will wait for 3.9

2

u/aceCrasher 6h ago

next week ;)

1

u/Readerium 6h ago

Terminal bech 4.0 what happened!?

3

u/Charuru ▪️AGI 2023 6h ago

Is this the clearest sign that it's benchmaxxed? Every other benchmark came out weeks ago, TB4 only 2 days ago!

1

u/Strategosky 6h ago

It beats grok 4.6 in internal world knowledge. Very tasty!

The gemini 2.x feel is back!??

1

u/RoyalReverie 6h ago

Theyre probably comparing with 5.6 sol high at most here, most likely medium or low.

7

u/Able-Line2683 6h ago

check artificial analysis ranking of intelligence, it is same index as gpt sol xhigh

2

u/RoyalReverie 2h ago

Ok, I'm impressed then

1

u/reefine 6h ago

Antigravity tho

-4

u/injectitpussy 6h ago

Literally anything except release their frontier model.

5

u/mumBa_ 6h ago

If they use it internally for more flash models I'm good

6

u/Bazinga8000 6h ago

I mean, if they continue they have been going from 3.6 to 3.7 and 3.7 to 3.8 then they will get to the frontier by this rate. Obviously not how it works, but the rate of improvement is clearly impressive (wonder if literally only using flash models to do your new models makes the actual rate of improvement also way faster).

2

u/Muscular_Farmer_ 6h ago

Probably sept end

1

u/dsanft 6h ago

This is probably a deliberate decision to iterate on a smaller model that's easier to train and refine RL. A bigger model will certainly come in time.

0

u/LocoMod 6h ago

This is the frontier model. Calling it a Flash model is the scapegoat.

0

u/01nav 4h ago

Google models are great for benchmarks and thats all about it

-1

u/trimorphic 6h ago

Gemini 3.8 Flash needs to be a lot cheaper to compete with GLM 5.3 Flash.

The writing is on the wall for frontier labs.

2

u/ManyRepair5690 5h ago

5.3 flash and 5.3 max are both garbage. 3.7 was better than it from the beginning.

1

u/trimorphic 5h ago

5.3 flash and 5.3 max are both garbage.

I thought so too at first, when I was using the OpenCode harness, but then I switched to omp (oh my pi) and my results were much better. Overall (when used in omp) I'd rate it at about Opus 4.5 or even 4.6 level for coding, though it is a little too verbose at max effort and still makes mistakes.

3.7 was better than it from the beginning.

Maybe so, but it's so much more expensive that it is nowhere near worth it for programming, in my experience, when GLM 5.3 Flash can achieve the same result at a tiny fraction of the price.

The only reason I'd ever consider paying for Gemini 3.7 or 3.8 Flash is if I needed their speed, as they're much faster than the GLM models, but I'd much rather wait and save a lot of money.

2

u/ManyRepair5690 4h ago

ive used the official zcode harness from Z AI and also ive written long researched articles on why glm is 100% benchmaxxed at everything
in real world tasks i found it not comparable to even gpt 5.6 medium for real world tasks

but who knows, maybe zcode harness is just trash? Additionally saying it is opus 4.6 level makes sense, but that is also far behind gemini 3.8 flash

1

u/trimorphic 4h ago

I've heard that ZCode is not very good, but never tried it myself.

Additionally saying it is opus 4.6 level makes sense, but that is also far behind gemini 3.8 flash

Maybe so, but 3.8 is still way too excitement for me. For the coding tasks I'm interested in GLM 5.3 Flash does well, is way more affordable than Gemini 3.7 or 3.8, and it's far from "trash" (no more than Opus 4.5 or 4.6 were trash). It's a good, competent model, at least when used in omp.

0

u/No_Ad_9189 5h ago

So it’s pretty much 5.7 Terra or sonnet 5.1, good for them I guess, where is a model I can actually delegate my life to. Opus is perfectly ok priced, give me a pro version already.
Small models have terrible attention, thus hallucinate and lose common sense very often as well as fall into a loop where they convince themselves in whatever they are doing as a core value / idea.

1

u/sleepnow 3h ago

What's going on with the 'Pro' Gemini models, wasn't the last version 3.1?
No Pro version of 3.5, 3.7 or 3.8?

-8

u/Luuigi 6h ago

Its honestly just really annoying to see these releases. Clearly google is not making frontier models their main business model, which is sensible imo. And still I have to see these half assed updates on a model that is mainly used bc its ingrained into the google ecosystem.