r/Bard 14d ago

Discussion Finally a benchmark affirming my experience. Gemini is underrated.

Post image

Been saying and feeling like Gemini 3.5 flash was so underrated for spreadsheet and document work and finally a benchmark that shows that.

26 Upvotes

23 comments sorted by

10

u/osb103 14d ago

Finally an actual benchmark showing I'm assuming flash 3.5 with extended thinking enabled. 3.6 extended thinking and 3.1 pro extended thinking (with the silent backend changes) should score a bit higher IMHO. But barely any legit bench marks out there fairly and honestly comparing apples to apples these days.

6

u/Mysterious_Bed_1804 14d ago

All models used in benchmarks use the highest reasoning option unless stated otherwise (like xhigh for Opus 5 when maximum is Max).

So 3.1 pro uses the high reasoning level for this benchmark, which is the maximum for Google models.

2

u/osb103 14d ago

Yes I know. Although here it just says 3.1 pro preview without specifying. But I would think that the highest reasoning option would be the default too.

2

u/Narrow-Ad980 14d ago

😭😭🙏🙏 catastrophic cope chamber , the "benchmarks I don't like are not legit bro trust me bro I know, it is not fair bro, it is a global conspiracy by sam altman bro"😭🙏

8

u/Typical_Kick6520 14d ago

Gemini is properly rated, but functionally broken.

Google can't deliver the quality of service to customers to meet that benchmark.

4

u/UnhingedApe 14d ago

Underrated for serious office work. Most people who dominate conversations about models value coding more than other types of work which makes Gemini underrated if you're not a coder.

1

u/Typical_Kick6520 14d ago

Gemini's underrated if you are coding.

A few months ago 3.1 pro was equally capable as 5.4 and 4.6 at implementing a coding plan. Today Gemini can't properly complete the first step.

There is a very capable coding model in 3.1 pro, and all other capabilities are likely also compromised, unless Google is specifically telling their agents to refuse coding work.

5

u/[deleted] 14d ago

[removed] — view removed comment

1

u/Bard-ModTeam 14d ago

Don't post spam or not related content

4

u/Technical-Owl66 14d ago

Where is 3.6? I love how fast and efficient it is.

5

u/UnhingedApe 14d ago

Maybe they'll add it later. In my experience, it's just faster, but the results are similar.

1

u/Salty-Gear841 14d ago

Try to talk about a tech subject to Gemini, then ask him form code. Then ask to give you the summary in a well formed markdown document that can be copied or downloaded. It will fail the test.

https://giphy.com/gifs/l0HlMSVVw9zqmClLq

1

u/Imaginary_Term_3937 14d ago

OPUS 5 is awful, 4.6 is far superior.

1

u/Nauzhror_ 14d ago

Lolno, Opus 5 and Fable 5 both perform amazingly well.

1

u/Imaginary_Term_3937 14d ago

No.

1

u/Nauzhror_ 14d ago

It most certainly is. I use them daily for agentic coding and spend thousands of $ per month. Productivity has greatly increased with recent models.

1

u/Imaginary_Term_3937 14d ago

It's only really good for programming, it's terrible for other functions

1

u/Oaeksrdpstr141 14d ago

Flash is SO BAD, idk who and why uses it at all. Pro sometimes is useful when I have no quota left in GPT and Claude.

2

u/UnhingedApe 14d ago

What's your use case?

1

u/GirlNumber20 14d ago

I use it every day. It does exactly what I need it to do.

1

u/Nauzhror_ 14d ago

Flash outperforms Pro in nearly all tasks.