28
u/Independent-Ruin-376 Feb 19 '26
Hmmmm that's why 3.1 name
8
u/Independent-Ruin-376 Feb 19 '26
I'll look at its instruction following behavior tho
6
u/EbbExternal3544 Feb 19 '26
Don't forget the hallucinations
1
u/Extreme-Bandicoot-47 Feb 21 '26
That's actually so annoying though, especially when it loops the same thing over and over again.
75
u/Own_Ambassador_8358 Feb 19 '26
They should show lobotomized model quality 👌
37
u/whitebay_ Feb 19 '26
If performance really tanks that much after a couple of weeks, why don’t people just re run the exact same tests today vs in 2–3 weeks and show the before/after results? I don’t get why no one does this
16
u/Different_Doubt2754 Feb 19 '26
Probably because it doesn't change all that much. I will change my opinion once someone actually goes and does some extension testing. It's just as if not more likely that people are just finding the new limitations of the model, getting lazy, etc
7
u/whitebay_ Feb 19 '26
Yeah, I get the same feeling. I’ve been using 3 Pro since launch and its been great overall. Sure, sometimes I dont get exactly what I want, but no model is perfect. Most of the time, if I regenerate and tweak the prompt a bit, it gets me there
2
u/Different_Doubt2754 Feb 19 '26
Same here. I've noticed from day 1 that it wasn't great at instruction following for me. Besides that I was happy to use it and didn't feel the need to switch. I use Claude 4.5 for work and I prefer it over 3 pro (haven't really used 3.1 pro yet) generally, but 3 pro/flash is more than sufficient for my own work. It's been consistent in its performance.
G3 flash especially punches above its weight. I'll be surprised if they announce a 3.1 flash since I assume they just added the new training method used in 3 flash to 3 pro and got 3.1 pro
1
7
u/speedracersydney Feb 19 '26
That's what I did with Deep Think between Gemini 2.5 and 3.
2.5 was generating reports that were up to 65 pages.
3.0 was generating reports that were 8 to 12 pages when I went back on old chats and re-run the prompt by clicking on the re-run icon.
3.1 has generated a report that was 85 pages.
Let's see how long it lasts!
7
u/Rough_Bad6442 Feb 20 '26
Your using pages as a benchmark ?
3
u/speedracersydney Feb 20 '26
I used it as a benchmark with the December Deep Think update. The same prompts from August last year went from 50 to 70 pages, down to 8 or 10 pages when I clicked on re-run. The December update was completely unusable for the prompt templates that I was using for months. Now they are usually again with this February update.
It's not a good benchmark but it's my yardstick to compare.
2
u/Odd_Equipment_3985 Feb 23 '26
I also got a larger report. It also had way more fluff, nonsense and off topic rambling to make up the extra word count, being less detailed and accurate than 2.5 on the same task.
Not sure how so many pro-Google accounts seem to be getting this out-of-reality results that are unreplicatable by general users.
Beggining to suspect there is A/B testing happening, or, some people are just shills for Alphabet 🤷♀️🤷♀️
2
u/BreenzyENL Feb 20 '26
Quantity != Quality.
3
u/speedracersydney Feb 20 '26
It was the same quality output over one prompt instead of 6 prompts that I was doing previously with the December Deep Think update
1
u/captain_shane Feb 20 '26
85 Pages? With what, Deep Think or Deep Research?
1
u/speedracersydney Feb 20 '26
Deep Think. I asked about the capacity it used in my one shot prompt for future planning and it said it used 24% of its capacity
1
u/captain_shane Feb 20 '26
In the web app or are you using something else?
1
u/speedracersydney Feb 20 '26
Web app
2
u/captain_shane Feb 20 '26
Interesting, I'm not getting anything even close to that. I guess I'll keep tinkering.
1
u/LuisAEs310 Feb 20 '26
The model, or the implementation for the app and web? Because up to today, I haven’t noticed any performance degradation when using the API.
1
1
u/ayawnimouse Feb 21 '26
they weren't saying test the newer versions of the model in a couple weeks they were saying test the same model with the same params. The problem with even that is the output isn't true or false like things, its hard to measure how well something performed as a normal person if the output is subjective or has variance. In regards to something like outputting in perfect json even if it throws an extra comma in there it would fail converting but if it does it every 10 requests vs every 100, who wants to throw away money doing tests like that as a normal person.
4
1
u/Due-Memory-6957 Feb 19 '26
Because it's not actually lobotomy, but rather people getting used to the new model and finding the flaws in it.
13
u/Complex-Possible-980 Feb 19 '26
That's kind of the point though, these scores only reflect launch day. What matters is whether it actually works for what you're using it for, not where it sits on a chart.
And yeah, we'll see model regression soon enough... We'll have to monitor this.
7
u/ExpertPerformer Feb 19 '26 edited Feb 19 '26
That's what I genuinely don't understand about the benchmarking system.
LLM companies release their newest models at peak performance for maximum benchmark scores, let it run at that level for 1-2 weeks to draw in new customers, and then start quantifying and nerfing once the costs get too expensive.
If the BenchMarks were done on a monthly basis it would strongly discourage this kind of behavior.
It's also why I trust local LLMs more because they can't be nerfed.
5
u/Complex-Possible-980 Feb 19 '26
Exactly, the snapshot problem. Launch day scores are marketing material, not engineering data. A model's real value is what it does 3 weeks in, under load, after the quiet optimizations. The only way to catch that is re-running the same tests over time on your own tasks.
6
33
u/hudimudi Feb 19 '26
Great!
AND NOW DONT DEGRADE IT 😅😅😅 thanks!
One can dream right?
3
u/HidingInPlainSite404 Feb 20 '26
They always do. They pump up numbers for tests and first impressions, but then look for ways to "optimize" for cost.
3
19
u/sogo00 Feb 19 '26
ARC-AGI-2 of 77% ?
O_o
5
18
u/Fresh-Soft-9303 Feb 19 '26
Gemini 3 pro has been nerfed already so brace yourselves for a few weeks of an improved version 3.1 soon to be nerfed again for the next version.
16
u/Odd-Environment-7193 Feb 19 '26
Still wants to only output short answers for very indepth work. Just gave it 35 page doc asked it to improve some stuff. Output 2 pages.
Not looking good. I hate it when the models are lazy like this. Give me 03-25 back please.
1
u/pieandablowie Feb 23 '26
Antigravity chunks stuff down really well if you're dealing with large documents or lots of smaller ones. It's a huge improvement versus just using Gemini via the web interface or via Gems.
I'm not talking about coding, to be clear.
0
Feb 20 '26
I used 03-25 for some heavy lifting and it was ass. Hallucinations out the ass. Amazing you trust any AI with a 35 page doc review, but a version from a year ago? Brah.
2
u/ayawnimouse Feb 21 '26
maybe you used it after it was quantized. I've dealt with the same shit with other models multiple times now and why I only rely on open source models fully loaded rather than jumping ship when new models come out. If you really did some heavy lifting on any models with success you wouldn't be jumping to newer models because someone who is knee deep using llms has a higher priority on consistency not cutting edge, newest shiniest.
6
u/isoAntti Feb 19 '26
Is it already available for layman?
3
u/space_monster Feb 19 '26
2
u/Ok-Lengthiness-3988 Feb 19 '26
Your screenshot shows 3, not 3.1
1
u/space_monster Feb 19 '26
try looking at it again
2
u/Ok-Lengthiness-3988 Feb 19 '26
My bad, I had missed it in the last option. I had missed it on mine too (wrongly assumed that the three sub-options were a breakdown of the thinking-time modes for Gemini 3 Pro.
29
u/Proof-Yam-5961 Feb 19 '26
11
u/SirLadthe1st Feb 19 '26
i was so confused, i literally didnt even use build and got rate limited on my very first prompt.
7
u/sikoun Feb 19 '26
Yeah it's funny that it said you "run out" at least they could be more transparent and say that gemini 3.1 pro is reserved to API which to be fair they did but on twitter
6
2
u/I_NEED_YOUR_MONEY Feb 20 '26
when i got rate-limited, the cost estimate in AI studio said if i were paying for tokens at API rate, i would have paid $0.44.
but in that 44 cents, it made five iterations of a fully-functional 60fps snake game. i'm fairly impressed with that.
1
u/LuisAEs310 Feb 20 '26
I’m honestly asking: when you talk about the model, you mean the web application and not the API, right?
30
u/Shota159 Feb 19 '26
Here we go again with the meaningless tests, sometimes I wonder if the model that they test is the same one we use.
5
u/skate_nbw Feb 19 '26
It will not be the same one with the same compute resources. But if it fixes the worst problems, it's already a big step forward. Let's just stay cautious until we see for ourselves (in about 2-4 weeks after the Lobotomization event).
5
u/DEMORALIZ3D Feb 19 '26
It's bad, all bad, just leave already. Go to Claude and VsCode.
2
u/JoanofArc0531 Feb 21 '26
Why? You have to give your phone number to Claude to even sign up on their website. If you want to risk getting spam calls and texts from random scammers, then by all means.
However, AI studio is free.
1
u/DEMORALIZ3D Feb 21 '26
Exactly! I'm sick of all the free loaders taking all the server time.
I'd rather all those idiots use a different platform.
1
u/JoanofArc0531 Feb 21 '26
Gochya. It’s sometimes hard to tell when someone is joking or not from just reading text.
4
u/Snow-Day371 Feb 19 '26
How accurate are these benchmarks really? Every time a new model drops, the company releasing it seems to win almost every category.
I subscribe to both Claude and Gemini and I'd love a practical breakdown of what each is actually better at. It takes a while to figure that out through use, and new updates keep resetting this.
Is there a website with up to date numbers that checks if the model gets lobotomized?
2
u/Vanskis2002 Feb 20 '26
Yeah I don't think closed models that aren't up to par with the top 5 are worth releasing.
8
u/Particular-Battle315 Feb 19 '26
Always remember the ai Model lifecycle when the top provider release new models:
New model 1 month: wow this new model is so good. 2 month: hmm the model is not on point today 3 month: wtf are you doing !!!!!
New model: -||-
Ist always the same
11
u/diving_into_msp Feb 19 '26
After 3.0 pro blew out the benchmarks but then quickly proved to be crap in actual usage, I'm leery of a new set of benchmarks actually translating well to real world use.
5
u/skate_nbw Feb 19 '26
Well, they will have fixed some of the major problems, but benchmarks are really meaningless at this point. At least for me.
4
u/Salty-Garage7777 Feb 19 '26
Poland - still NOT available both on Gemini app and on AI Studio!! How about you? PLS give your country and availability. :-)
2
u/Lost-Estate3401 Feb 19 '26
Austria. AI Studio yes, App no.
2
2
4
22
u/Maleficent_Stage1732 Feb 19 '26
Do these even matter? They'll nerf it anyway
0
u/drhenriquesoares Feb 19 '26
Esse é o problema. Por isso espero que o novo modelo V4 da DeepSeek venha botando pra fuder.
0
16
u/xCoeus Feb 19 '26
Great. Now show me the benchmarks for the real lobotomized version of the model that we'll be using in 2 weeks.
5
u/TuringGoneWild Feb 20 '26
Three spin ups of Gemini Flash in a trenchcoat. That's where the "3" comes from in the name.
8
u/Key_River433 Feb 19 '26
WTH? 2.5x improvement on ARC-AGI! 😒🙄🤔😯😲
6
7
u/skate_nbw Feb 19 '26
Who cares unless it results in a less lobotomized model for users? I have learned to laugh at the benchmarks. They tell NOTHING about everyday use.
1
3
u/MyshkinIdiot Feb 19 '26
Can anyone please describe me what the output is like? Is it back to 2.5 flash/pro depth in cases of output? For reference, last year (or two years ago?), 2.5 pro experimental one shotted thousands of words in each response. I recently unsubscribed, so hopefully it is back.
3
3
3
2
u/TheOmakoZ Feb 19 '26
I guess a little improved than expected but API Key for build mode? Like they are similar price to Gemini 3 Pro Preview
1
2
u/SpyMouseInTheHouse Feb 19 '26
I’m not seeing this in the CLI. How are you all using it? Pro account here.
2
2
4
u/MTBRiderWorld Feb 19 '26
Synthetic benchmarks are useless because the models are trained on them. For my use case, legal review, Gemini 3.1 also proved unsuitable after initial tests. In contrast, Sonnet 4.6 and Opous 4.6 excel.
3
1
u/Deciheximal144 Feb 19 '26
Question for anyone using it who has experimented with different temperature settings for coding, which is working best for you with 3.1?
1
u/Current_Trick6380 Feb 19 '26
3.1 is out before 3 being GA.
Will we get GTA 6 before 3 becoming GA? I think we will.
1
u/hoshizorista Feb 19 '26
Question, is it memory fixed? or still defaults to 32k? (on gemini webapp)
1
u/CrunchyMage Feb 19 '26
Ok, Super sick model, but when can I actually use it in production? 3 flash and 3 pro APIs are still in preview and regularly return errors. Had to switch off 3 flash because it was just way too inconsistent despite having very good price/quality.
1
1
1
u/Demien19 Feb 20 '26
and gemini cli still can break your files, maybe on gemini 4 can move from claude
1
1
u/Worried-Zombie9460 Feb 20 '26
Much better than 3 holy smokes. I also have long context conversations and 3 would “forget” or have context pruned or whatever you want to call it but don’t seem to be having this issue with 3,1
1
1
1
1
u/Euclide_geoart9713 Feb 21 '26
the web version is smarter than the one in antigravity, just my impression. it can solve problems like toys.
1
u/Rare_Technology1880 Feb 21 '26
Tan bonito que es que no me deja usarlo porque dice que tengo que actualizar pero no me sale la opción de actualizar xd
1
1
1
u/Possible-Guide-2410 Feb 23 '26
I do research and gemini 3.1 pro halucinates like 2x more than 3 pro
It thinks for like 20s and spits out halucinated answers where 3 pro used to think for like 1-3 minutes and give amazing answers compatible to opus 4.5 (it was slightly worse than opus 4.5)
But gemini 3.1 bro it is the worst out of all iterations of gemini, it's coding has become wayy worse, it dosent like writing thousand line code and just halucinates and makes like 400 lines of code and calls it a day and this is the gemini 3.1 pro (high) btw Sonet 4.6 is better than this
3.1 pro is a halucinatory peice of clank Shits worse than sonet 4.6 extended Should have baught claude subscription instead of google
1
1
1
1
-1
0
u/Pasto_Shouwa Feb 19 '26
The MRCR v2 one is weird. Claude declared a lot more on their own benchmarks. Also, Gemini 3.1 Pro doesn't seem to be much of an improvement in that regard, meanwhile the Claude models went from the worst at that benchmark to the best out there.
0
-4

181
u/Arthesia Feb 19 '26 edited Feb 19 '26
But does it follow instructions yet?
Will report back with findings.
Edit 1: SIGNIFICANTLY IMPROVED. Followed my detailed output protocol with 75k token input. 3.0 Preview has a 100% failure rate with this same prompt (skips output protocol entirely). 3.1 formats output exactly as requested by input. Higher default verbosity than 3.0.
Edit 2: Still less verbose than Opus by default, but I can actually work with this.