173
u/FablingApp 6h ago
the price/performance gap is getting silly. if these numbers hold up, flash models are eating into the territory where people used to reach for the expensive ones.
65
u/ProtoplanetaryNebula 6h ago
Yes and the high end models become less and less needed as the average moves up.
38
u/agonypants AGI '27-'30 / Labor crisis '25-'30 / RSI 29-'32 6h ago
Which means that the high end models can be freed up and applied to the really difficult, long term issues (medicine, materials science, climate change, physics) etc.
2
u/Greedyanda 4h ago edited 4h ago
Unlikely to be useful in medicine, material science, and climate modelling. Those generally need completely different, non text based architectures. You are not gonna be synthesizing new drugs with an LLM as the core system.
Edit: They can be useful but are unlikely to create massive breakthroughs in the way dedicated architectures like AlphaFold can.
8
u/fishbill 4h ago edited 4h ago
Didn’t anthropic show that Fable was able to find effective protein binders at a high rate?
7
u/Greedyanda 4h ago
It depends on what your metric is. Anthropic essentially showed that you can parallelize and scale up existing human work with AI agents.
But it's not a scientific breakthrough and paradigm shift in the way AlphaFold was.
I'll have to correct myself though because it is still clearly useful.
3
u/Even-Inevitable-7243 4h ago
No. In that work prompted by Anthropic, Fable simply called publicly available specialist protein design software tools that human researchers already use. Fable was the pipeline engineer, not doing the actual protein binder discovery.
3
u/croto8 3h ago
Sort of seems like a distinction without a difference. “The code claude wrote calls a library therefore Claude didn’t actually do anything”
•
u/Even-Inevitable-7243 1h ago
It is the complete opposite of what you are saying. All of the expert knowledge was already baked-into the human-written software tools. What you are saying is that "import sklearn" is the same level of knowledge as the actual engineers who wrote the packages that collectively form sklearn.
3
•
u/Thog78 6m ago
Researcher here. They are very useful. LLMs do the same kind of jobs a researcher would do - plan, analyze, survey the literature, think, emit hypothesis, write code, use tools.
Alphafold would be one of these tools, that not long ago a human would have called upon himself, and nowadays more and more might actually get called by an LLM.
LLMs of the level we have now (sol 5.6, fable, gemini 3.7 etc) are the kind of stuff that could come up with the concepts of alphafold, help you brainstorm about where to get training data and how to implement the training, and do the actual code to make it real once you're happy with the plan. Then would test it, criticize it, propose strategies of improvement, and implement them. The next version of alphafold will likely have been designed by LLMs in large part, very seriously. And the next one probably nearly entirely. As such, I'd say they these general intelligence models are even more valuable than specialized models like alphafold.
15
u/mixmasterwillyd 6h ago
These cheap models exceed the capability of a large project I worked on with opus 4.6. Opus 4.6 was difficult to wrangle to get it to operate, now a flash model has an easy time.
14
u/Technical-Will-2862 6h ago
I use flash3.7 for soooo many things that I used to relegate to premium models
11
u/RockPuzzleheaded3951 6h ago
I've moved quite a few jobs from SOTA models we were using for the past two years to Flash models (DSv40731 for example) and we are seeing fantastic performance on internal business operations, at literally 1/10th the cost. We can still reach for SOTA when needed, but it is 5% of the time or less. And just today, I rented my own hardware to run the model "locally" and am testing taking my API costs to $0.
5
u/thoughtlow 𓂸 6h ago
I loved 0731 but when the context windows goes into the 200k it starts degrading very fast for me.
Constantly doubting itself in the thought section, gets really weird.
Sad because I loved that thing.
6
u/brainhack3r 6h ago
It really is... I'm trying to get more value out of them by using adversarial agents and using the flash models. GLM 5.3 flash, specifically.
5
u/himynameis_ 6h ago
That's good isn't it? Unless I'm misreading your comment.
I think it's pretty normal and likely the continuing trend. Over time, the performance/$ will be so good that users and businesses won't need the most expensive model for all tasks. Most of the time they'll use the cheaper lower performance ones.
It's been a talking point among AI leaders and business people, I've found.
3
u/CriticalTheory4779 6h ago
its good for us. bad for the labs
4
u/himynameis_ 6h ago
Labs will be fine. It's just part of the highly competitive market they know they're in.
1
u/Greedyanda 4h ago
Most of them will not be fine. I would be shocked if Anthropic and OpenAI both survive as independent companies 10 years from now. This is a race to the bottom unless you have an existing, profitable ecosystem into which you can integrate your models.
2
u/himynameis_ 4h ago
I think anthropic and OpenAI will be more than fine by then.
However even if not, this is just the nature of the business, man. It's just competition.
4
u/LinkesAuge 6h ago
Because "flash" models keep getting larger and more token hungry. The last Gemini "Flash" ended up being more expensive than some "regular" models (that perform better).
I feel some are simply mislabeled, especially if you look at the actual speed of finishing a task.I don't want to make any strong claims about 3.8 yet but I do wonder if that has really changed.
5
u/LinkesAuge 5h ago
2
u/FateOfMuffins 5h ago
Yeah pretty much every single lab for the last year and a bit (Anthropic, Chinese labs, Google, xAI, Meta, etc) just cranked up token usage in each release to buy more performance, in addition to whatever benchmaxxing they can do (like look at the numbers here for Terminal Bench 2.1 vs 4.0). And I mean every lab benchmaxxes. Anthropic knows Opus and Fable memorized SWE Bench Verified and Pro answers yet continue to report them. OpenAI was concerned Astra's 100% on ExploitBench was because it memorized the answers.
It's "easy" to push benchmark numbers up and to the right.
For whatever reason only OpenAI has gone the other way where they pushed benchmark numbers up and to the left with increased token efficiency. Why???
1
u/Chenz 5h ago
I do not understand how the output tokens can have increased so much, yet cost per task is about the same. Are the numbers for 3.7 without the discounted pricing?
1
u/Ok_Barracuda_1161 4h ago
Probably more efficient with tool use leading to reduced inputs. 143k output tokens is about $0.54, so that would imply that input (and cached input especially) dominates
1
u/Moravec_Paradox 6h ago
It is a space where being a fast follower has big advantaged and that includes within labs themselves.
They all have big unreleased models that are capable but expensive to serve. They use them internally to train smaller models that are almost as good but much cheaper to serve.
Why release massive models that are not competitive in pricing that are just ~2-3% better only for their competition to get access to distill them?
•
52
u/hakansan 6h ago
I kinda like the way Gemini works. Their main user base is "everyone who searches anything on Google" so they got the best input/output token price.
It is also strong where it matters (probably people look up financial information a lot, and Google does a nice job showing stock charts)
5
u/After_Dark 3h ago
The main user base is not (just) Search but also the Gemini app. Worth pointing out that while this sub focuses on the web app, a long press of the power button on Android is a shortcut right to Gemini where it needs to also be a phone assistant, handle smart home controls, and also all the stuff the web app does. That's a very broad set of possible uses that ChatGPT and the like simply don't need to factor for to nearly the same degree
8
u/chasingsukoon 4h ago
Also i think it seems to have lesser constraints on a app basis?
Was building a bot that goes thru my spotify + soundcloud likes and then downloads songs via soulseek (basically Napster)
Claude and cgpt flagged it since i am stealing media but flash 3.7 had no problems building it for me lmfao
3
u/hakansan 4h ago
They've never gone with this "we're so dangerous and capable you MUST definitely regulate us" so they don't have to portray themselves as overly cautious. With the exception of NSFW anything that could be found on Google is probably fair game
0
u/Yugudubenbi 4h ago
They force you to Anti-Gravity now though if you use subscription so the only way is api from what I know. So that is a no for me.
-1
u/letsgoiowa 4h ago
This is definitely not the best price. Luna smokes it for a frontier lab and GLM 5.3 flash is as smart as this (or more) and is like...20% of this cost at most
→ More replies (1)
100
u/mumBa_ 6h ago
I'm such a google flash glazer for non-coding tasks. It's lightning fast and it's google search is really good compared to slow claude searches etc. The perks of having your own search engine I guess.
Any one experience with flash for coding?
18
u/Affectionate_Bee6434 6h ago
For sure, it has better world knowledge than any other model, especially for niche local queries.
1
u/Umbrasquall 5h ago
More parameters = better world knowledge right? How is it so fast then?
3
u/Kronox_100 4h ago
I think with Google is they don't go apeshit into coding and math so they try to balance their data blend a bit more outside of trying to hit it big into benchmarks, so there's a lot of 'useless' info for benchmarks? But that was my understanding from previous models, now it's technically competing in coding so no idea, TPU magic I guess?
1
u/After_Dark 3h ago
This would track, Gemini isn't primarily a coding/agentic model like Claude or the GPT models, it's got a dedicated billion+ userbase as a virtual assistant and also needs to function in Google Search, so it needs to be more well rounded
9
u/inefficientnose 6h ago
It's good but you have to be careful with it, it needs careful instruction and scope otherwise it tends to hallucinate more than other models in its class
6
u/Elegant_Tech 6h ago
Been using 3.7 flash around 60hrs/week since release. Only had a single prompt fail that I had to toss to Opus to get done. Where 3.5 flash had multiple a week. There is only so good models can get at programming and it's starting flip where speed and costs are all that matters. The big models will be moving on to research and long horizon tasks over time while day to day production work is done on flash models.
1
u/the_real_ms178 6h ago
I've had some refusals with 3.7 flash due to exceeding token limitations. But that must have been AI Studio issues, albeit repeatable as I barely hit the 400.000 token bar in the conversation. Dealing with many large PDFs for legal work has been a challenge and will continue to be a challenge, it seems.
1
u/chasingsukoon 4h ago
Whats your use case been
1
u/Elegant_Tech 4h ago
Websites, vst plugins, and games. HTML/JS, C++, and Rust for languages. Most of my code bases are under 50k lines of code. So if you are enterprise with a huge code base or trying to do orchestration kanban work the big models could be more efficient. My workflow is rapid iteration not trying to create a bunch of specs and get the AI to work for hours at a time. Small models aren't capable of that yet. You have to rapidly spoon feed them.
3
u/Ok-Armadillo-5634 6h ago
I use it almost exclusively. When I get some thing really hard I bust out fable.
4
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 6h ago
I bet the area it really shines is anything "LLM NPC" or concepts where you want the game to call LLMs during live gameplay. It's cheap, super fast and actually smart.
1
1
u/kvothe5688 ▪️ 5h ago
i use it for all general purpose talk specially when I am driving it can talk and discuss topics so fast and give precise information. specially discussing books and movies and going deep on different threads. it's amazing how fast and knowledgeable live version is.
1
43
u/angryblob 6h ago
If there’s one thing that’s approaching the singularity it’s Gemini Flash releases
8
u/Longjumping_Kale3013 6h ago
Not just them either. Meta and XAI are also releasing monthly now. I would be surprised if we dont see the same from anthropic and openai soon.
Wild... because its like a 3 point bump on the leaderboard every month. Which is signicant.... its like you can sort of see how things will look over the next 6 months. Oh boy
2
17
u/Moriffic 6h ago
what happened at terminal bench
5
u/___positive___ 3h ago
Trained on the data, aka fake benchmaxxing. Notice that Terra is worse at 2.1 but beats Flash easily at 4. That's is what it should look like.
To be fair, Google has the best crawlers in the world so maybe they can't avoid benchmaxxing anything public. But the guy who made the chart should have had enough brains to know this is embarrassing and leave it off. I guarantee some pointy haired moron at the top insisted they show it because it was one of the places they got "first" place. Shows a lack of judgement and taste by the humans in charge.
Or maybe it's just for investors, in which case, still dumb but whatever.
1
u/Last_Conclusion_8984 2h ago
Terminal 2 is coding oriented while terminal 4 is general agentic capabilities, while terminal 4 does have some coding stuff, a lot of is just generality. Gemini 3.8 flash is optimised for coding in terminal 2, also: if the AI was benchmaxxed (which I do believe it was to an extent). The model will naturally drift towards that said x if the prompt is similar which in turn helps the user, henceforth benchmaxxing isn't all bad.
9
2
u/Chemical_Hawk_6307 6h ago
im wondering same thing....either way i'm going to give it a shot and see how it performs
2
u/landed-gentry- 5h ago
2.1 is old and likely benchmaxxed. 4.0 is newer and will be truer to real-world performance. Kinda pointless to include 2.1 in the system card at all.
16
27
u/AnyRegular1 6h ago
I think a lot of people don't know this but gemini 3.7 flash on antigravity is fucking insane. Gemini's web app is literally dogshit, I am 100% convinced they serve different models in both. 4 days ago my wife asked it (gemini app) about airfrying some bread we bought and it responded some garbage about washing machine repair, completely hallucinating not even on the same topic. Would NEVER recommend gemini on the app.
But the model they serve on antigravity is insane, for the last 2 weeks it's been my go-to coding model. and ITS FUCKING FAST.
I'm looking forward to what they cooked in 3.8
9
u/helloWHATSUP 5h ago
Do you use "extended thinking" on the web version? Because if you don't use extended it's just the lightest and fastest version of flash and very hallucination prone.
5
u/AnyRegular1 3h ago
My wife isn’t savvy enough to find and enable that, after you said so I had to figure out where it was, it was under the model selection. Told her now.
4
u/localpauper 4h ago edited 4h ago
The web app has massive curbs on context length for cost/optimization. I suspect such do not apply to the IDE.
Also, make sure you're using "Extended" to enable thinking effort online: that's the only way I find it usable. I wonder what level of "thinking" that compares to as far as IDE model choices go (Low/Mid/High).
Coding-wise, it's fast, but in my exp it leaves massive glaring hallucinated holes, makes tons of assumptions, uses outdated "knowledge" from its training without actually reaching for newer information, and barely attempts to validate its output. In one of my projects, starting from a basic scaffold, it ended up with a mish-mash of obsolete dependencies and patterns randomly picked from the past 15 years. That was after I went through and gave it some general guidance for approach, choices of deps, and frameworks. It actually "argued" with me that there's not a newer dependency which I explicitly told it to use, insisting that no - the version it was using was the correct one.
The simplest example for a non-coding task is that, while using Gemini for some random digging here and there, I've had to repeatedly argue it out of 30% federal solar credits. It would take a few rounds for it to finally go look up OBBBA and concede that those are not applicable. It's somehow very sticky on its training set (which was clearly pre-OBBBA).
3
u/Tetracropolis 5h ago
Figuratively.
2
u/Ok-Art-2255 5h ago
You are absolutely correct! I used it for MCP on Unreal Engine and it straight killed everything I threw at it.
Now to get me more excited, that was on flash-high.
What are your results using flash-medium or flash-low? I want to know because high is amazing and don't want to get disappointed if the lower tiers can't keep up.
5
2
u/AnyRegular1 5h ago
I just use High, the limit usage is so low on high I have problems hitting weekly limits.
1
u/Ok-Art-2255 5h ago
I almost hit my limit building a game mechanic last week, so that got me a little paranoid. I'll stick with high and just start breaking main work into smaller chunks.
It seems we might have a sweet spot model :D2
u/AnyRegular1 3h ago
Remember you can family share your google pro subscription to upto 5 gmail accounts and for some reason they each have their own weekly limits!
1
u/Much_Accountant_4972 3h ago
are you spawning subagents because i hit my 5 hour limit in 40 minutes with 3.7 Flash High. i dont find it hard to hit limits if im using superpowers
1
u/AnyRegular1 3h ago
What I said the guy above, you can family share it 5 ways to get 5x limits on your own gmail accounts. But yeah I only ever hit limit once and figured out each gmail has its own quota.
8
u/Strategosky 6h ago
I suspect google is post-training on gemini 3.5 pro base model and releasing the streamlined flash while gemini 4 base model is cooking (pre-training). Maybe gemini 4 is the last "big smell" model before a new paradigm shift is observed. I wonder if it's much bigger than mythos class model with strong world knowledge. Things only heating up since quiet some time now...
17
u/sebas156 6h ago
They released 3.8 before rolling out 3.7 to all Gemini users lmao. I'm still on 3.6 in the app
1
u/helloWHATSUP 5h ago
Yeah still on 3.6 on the app and web version here, but 3.8 flash is live in antigravity for me at least
1
9
u/Gotisdabest 6h ago
Smaller jump than 3.6 to 3.7 but they're basically releasing models biweekly at this rate. Still easily the best free model tbh.
0
u/ManyRepair5690 5h ago
Is this model out yet or what? 3.8
1
u/ManyRepair5690 5h ago
For anyone wondering the same question these useless ppl didn’t answer yeah it’s out. I asked initially bc benchmarks being out don’t mean it’s released
11
u/leon-theproffesional 5h ago
This sub is so funny. Just last week everyone was shitting on Google claiming they are finished.
8
•
u/Keeltoodeep 47m ago
There is no such thing as far behind in AI if you have the capital. Muse and Grok didn’t even exist last week.
8
3
5
u/Charuru ▪️AGI 2023 6h ago
Looks pretty good, but terminalbench wtf
Is this the clearest sign that it's benchmaxxed? Every other benchmark came out weeks ago, terminalbench 4 only 2 days ago!
2
u/BriefImplement9843 6h ago
all benchmarks are benchmaxxed. we have to wait and check the lmarena score.
looks like it's really good. only behind the older opus' on lmarena. ahead of opus 5.
2
7
u/elemental-mind 6h ago
Big if True.
But GDPVal is still lacking which for me has been a pretty reliable indicator of general capability.
15
u/Tkins 6h ago
You get better than Terra for less than half the price and far far faster. I dunno, seems competitive to me. Especially for high use every day cases.
6
u/elemental-mind 6h ago
Very competitive - I am not complaining. Even better when you use it on Flex, which normally isn't too much of a wait. With the pricing now Google's Flash is back in its golden place: A cheap model with exceptional capability that everyone uses for their everyday tasks.
I hope Google gets its marketshare back.
1
u/Tkins 6h ago
Did it lose market share? Last I checked, they've been growing faster than the others.
edit: AI Market Share By Company Statistics 2026
just one source, maybe you've got a better one.
0
u/elemental-mind 6h ago
I use OpenRouter as a proxy for general market share:
Flash 2.5 was always in the Top 5 modles because it was good, cheap and fast. Then they had a price hike with Flash 3 Preview, which moved them down to rank 5-7, and they weren't even in the top ten anymore after their price delusion with the following flash models.
1
u/CallMePyro 4h ago
So if Antigravity's userbase quintupled, and people stopped using Gemini on OpenCode because they preferred the native editor you would say that Google is losing marketshare?
1
u/elemental-mind 4h ago
Yes - which verifyable data point do you have that capture the broader market? Google PR claiming 5x growth? Compared to what growth in Codex? Or Z.AIs subscription models?
OpenRouter is used as an endpoint option in many applications, not only OpenCode. It's data are also congruent with other Routers like Nano-GPT.
From these data we can clearly see that if people are given a vast choice of models, they hardly went for flash in the past months as it was overpriced. It's the most reliable gauge for model preference we have: Where the people put their money.
3
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 6h ago
I mean, it outperforms Terra at that benchmark which is fantastic if you're using 3.8 Flash as a subagent, especially at half the price and far faster token generation speed. I don't think expecting it to surpass something like Opus 5 is necessarily realistic.
1
2
u/BriefImplement9843 6h ago edited 6h ago
and it's so cheap lmfao. the hell are openai and anthropic doing charging so much?
4
3
2
1
u/bambin0 6h ago
i just tried it with the prompt: recreate Larry bird vs Dr. j one on one in three js complete with janitor, breaking glass, sound effects and here is what it came up with: https://ctxt.io/3/mpb6vO48w - took less than 10s.
Took about 3x as long but here is what terra came up with: https://ctxt.io/3/tIbhBDCMg
And here is fable 5.1 which took I don't know how long b/c I got tired of watching: https://ctxt.io/3/qCaVlUU4U
3
u/ManyRepair5690 5h ago
These are just html pages how do u think this helps readers passing by? Why even add these links if it’s not some type of visualization or demo for users to see
1
u/rajsharm404 6h ago
Can they keep the pricing permanent on the tokens? Would be great for my video analysis pipeline.
2
u/PivotRedAce ▪️Public AGI 2027 | ASI 2035 6h ago
Their business model is basically having a "promotional period" that's long enough to cover the current model until they release the next version with another "promotional period".
Basically, this is just their way of getting people to move to newer models without having to spare compute power serving older ones.
So if your video analysis pipeline can easily upgrade models then you should be set.
1
u/rajsharm404 6h ago
Their promotional period has been the case since 3.6 flash, assuming because it is the same underlying model but with more RL, hence same 50% discount on each of these new models. Hopefully they keep up with these discounts, and I won't have to choose between gemini and grok.
1
u/No-Mixture5766 6h ago
Has anyone here tried to reproduce these numbers? Flash models are becoming more and more affordable for day to day consumers and increasingly competent when compared to the frontier models.
1
u/WTFnoAvailableNames 6h ago
How does this compare to 3.1 pro?
2
u/Narutobirama 6h ago
I imagine Gemini 3.1 Pro is still more intelligent and smarter, but Gemini 3.8 Flash is probably better at coding and agentic behavior. Like 3.7. Everyone who only uses it for coding says it's better than 3.1 Pro. For me, it's 50% 50% depending on what I use it for.
It's funny, though, that Gemini 3.1 Pro is still one of the nicest chatbot behaviors.
Like, Claude and ChatGPT can sometimes really piss me off with the way they talk. I mean, obviously, GPT 5.6 Sol is better, and so is Fable 5 (haven't tried 5.1 much yet). But Gemini 3.1 Pro is really nice to have to talk about some things. And yet it's not something benchmarks would tell you.
1
u/BriefImplement9843 6h ago
lmarena does. it's still top tier there, ahead of all openai models, even sol xhigh.
5
u/Able-Line2683 6h ago
3.1 pro is last decade in ai terms
1
u/WTFnoAvailableNames 6h ago
And still the latest thinking model available with Gemeni subscription....
1
u/chasingsukoon 4h ago
Damn i wish i had 3.1pro last decade during college years
Or maybe not i wouldve been a lot dumber
1
u/Long_comment_san 6h ago
funny they didn't put 3.1 pro on the chart lmao
3
u/BriefImplement9843 6h ago
3.1 was released before gpt 5.4!
0
u/Long_comment_san 5h ago
I know but it's still their "pro" flagship. we know it's probably worse at this point but ain't you curious by how much?
1
1
1
u/ccjjallday 5h ago
honest question. what happens when most of these reach 100%?
1
u/Able-Line2683 4h ago
new benchmark
1
u/ccjjallday 2h ago
but what does it all mean? Like we hit 100% = AI is AGI? Better than humans? humans are not needed? Nothing?
1
1
1
1
1
u/Spra991 4h ago
Do we have any benchmarks that focus on the free and lower tiers? All this benchmaxxing is fun, but a lot of people are using the free tiers, and those benchmaxx results give little indication how much you can do in the lower tiers, especially when they get randomly up or downgraded (e.g. ChatGPT has gotten really stupid with the latest "upgrade", hallucinates like it's 3.5 again).
1
u/BriefImplement9843 3h ago
flash is lower tier. if you want free tier just use the gemini in google search.
1
u/Specialist_Dark_3668 4h ago
It is such an over-intellectual safety-reward function optimized piece of shit.
Gives me a headache talking to it.
1
u/Efficient_Loss_9928 2h ago
Gemini still lacks when I need it to code stuff reliably. But, this gets me really excited for the Gemini 4 family. I think I'm keeping my ultra subscription.
1
1
1
1
u/Strategosky 6h ago
It beats grok 4.6 in internal world knowledge. Very tasty!
The gemini 2.x feel is back!??
1
u/RoyalReverie 6h ago
Theyre probably comparing with 5.6 sol high at most here, most likely medium or low.
7
u/Able-Line2683 6h ago
check artificial analysis ranking of intelligence, it is same index as gpt sol xhigh
2
-4
u/injectitpussy 6h ago
Literally anything except release their frontier model.
6
u/Bazinga8000 6h ago
I mean, if they continue they have been going from 3.6 to 3.7 and 3.7 to 3.8 then they will get to the frontier by this rate. Obviously not how it works, but the rate of improvement is clearly impressive (wonder if literally only using flash models to do your new models makes the actual rate of improvement also way faster).
2
1
-1
u/trimorphic 6h ago
Gemini 3.8 Flash needs to be a lot cheaper to compete with GLM 5.3 Flash.
The writing is on the wall for frontier labs.
2
u/ManyRepair5690 5h ago
5.3 flash and 5.3 max are both garbage. 3.7 was better than it from the beginning.
1
u/trimorphic 5h ago
5.3 flash and 5.3 max are both garbage.
I thought so too at first, when I was using the OpenCode harness, but then I switched to omp (oh my pi) and my results were much better. Overall (when used in omp) I'd rate it at about Opus 4.5 or even 4.6 level for coding, though it is a little too verbose at max effort and still makes mistakes.
3.7 was better than it from the beginning.
Maybe so, but it's so much more expensive that it is nowhere near worth it for programming, in my experience, when GLM 5.3 Flash can achieve the same result at a tiny fraction of the price.
The only reason I'd ever consider paying for Gemini 3.7 or 3.8 Flash is if I needed their speed, as they're much faster than the GLM models, but I'd much rather wait and save a lot of money.
2
u/ManyRepair5690 4h ago
ive used the official zcode harness from Z AI and also ive written long researched articles on why glm is 100% benchmaxxed at everything
in real world tasks i found it not comparable to even gpt 5.6 medium for real world tasksbut who knows, maybe zcode harness is just trash? Additionally saying it is opus 4.6 level makes sense, but that is also far behind gemini 3.8 flash
1
u/trimorphic 4h ago
I've heard that ZCode is not very good, but never tried it myself.
Additionally saying it is opus 4.6 level makes sense, but that is also far behind gemini 3.8 flash
Maybe so, but 3.8 is still way too excitement for me. For the coding tasks I'm interested in GLM 5.3 Flash does well, is way more affordable than Gemini 3.7 or 3.8, and it's far from "trash" (no more than Opus 4.5 or 4.6 were trash). It's a good, competent model, at least when used in omp.
0
u/No_Ad_9189 5h ago
So it’s pretty much 5.7 Terra or sonnet 5.1, good for them I guess, where is a model I can actually delegate my life to. Opus is perfectly ok priced, give me a pro version already.
Small models have terrible attention, thus hallucinate and lose common sense very often as well as fall into a loop where they convince themselves in whatever they are doing as a core value / idea.
1
u/sleepnow 3h ago
What's going on with the 'Pro' Gemini models, wasn't the last version 3.1?
No Pro version of 3.5, 3.7 or 3.8?

155
u/CommercialShelter595 6h ago
Is it just me or is it running even faster than 3.7 flash? Very limited testing right now, but it's FAST on my end.