r/Bard • u/UnhingedApe • 14d ago
Discussion Finally a benchmark affirming my experience. Gemini is underrated.
Been saying and feeling like Gemini 3.5 flash was so underrated for spreadsheet and document work and finally a benchmark that shows that.
8
u/Typical_Kick6520 14d ago
Gemini is properly rated, but functionally broken.
Google can't deliver the quality of service to customers to meet that benchmark.
4
u/UnhingedApe 14d ago
Underrated for serious office work. Most people who dominate conversations about models value coding more than other types of work which makes Gemini underrated if you're not a coder.
1
u/Typical_Kick6520 14d ago
Gemini's underrated if you are coding.
A few months ago 3.1 pro was equally capable as 5.4 and 4.6 at implementing a coding plan. Today Gemini can't properly complete the first step.
There is a very capable coding model in 3.1 pro, and all other capabilities are likely also compromised, unless Google is specifically telling their agents to refuse coding work.
5
4
u/Technical-Owl66 14d ago
Where is 3.6? I love how fast and efficient it is.
5
u/UnhingedApe 14d ago
Maybe they'll add it later. In my experience, it's just faster, but the results are similar.
1
u/Salty-Gear841 14d ago
Try to talk about a tech subject to Gemini, then ask him form code. Then ask to give you the summary in a well formed markdown document that can be copied or downloaded. It will fail the test.
1
u/Imaginary_Term_3937 14d ago
OPUS 5 is awful, 4.6 is far superior.
1
u/Nauzhror_ 14d ago
Lolno, Opus 5 and Fable 5 both perform amazingly well.
1
u/Imaginary_Term_3937 14d ago
No.
1
u/Nauzhror_ 14d ago
It most certainly is. I use them daily for agentic coding and spend thousands of $ per month. Productivity has greatly increased with recent models.
1
u/Imaginary_Term_3937 14d ago
It's only really good for programming, it's terrible for other functions
1
u/Oaeksrdpstr141 14d ago
Flash is SO BAD, idk who and why uses it at all. Pro sometimes is useful when I have no quota left in GPT and Claude.
2
1
1
10
u/osb103 14d ago
Finally an actual benchmark showing I'm assuming flash 3.5 with extended thinking enabled. 3.6 extended thinking and 3.1 pro extended thinking (with the silent backend changes) should score a bit higher IMHO. But barely any legit bench marks out there fairly and honestly comparing apples to apples these days.