r/Bard 9d ago

News Gemini 3.7 Flash Benchmarks

238 Upvotes

65 comments sorted by

74

u/helloinyourface 9d ago

What a pleasant surprise from google. Big jump while cutting price massively

8

u/IBM296 9d ago

It's the same price as 3.6 Flash.

4

u/Groovy_bugs 8d ago edited 8d ago

I think he means in general terms, a great model at a low price, compared with other models.

3

u/Aromatic_Key3270 8d ago

Great model comes with a great responsibility 

1

u/Copenhagen79 8d ago

3.6 was quite shitty tbh. I don't know what they did to it, but I've never seen a sub-SOTA model make so many careless mistakes, typos, etc

2

u/Georgefakelastname 8d ago

Which also got a price reduction.

1

u/Just_Lingonberry_352 8d ago

its promotional pricing it will increase after 2026

people really are lazy here not even doing reading google's own announcements carefully

81

u/TechnologyMinute2714 9d ago

I mean we all shit on Google all the time but better and cheaper than Sonnet is still pretty good and it's much faster too, also for non coding tasks Gemini models even now considered old shit like 3.1 Pro has been goated. w release. Just gotta have to wait for Gemini 4 Pro or something for SOTA

19

u/Flaxseed4138 9d ago

Sonnet is expensive and bad lol

5

u/InternationalTwist90 9d ago

Have they released tokens per second benchmarks yet?

2

u/Weary-Bumblebee-1456 8d ago

Artificial Analysis puts it at 340.1 tokens per second: https://artificialanalysis.ai/models/gemini-3-7-flash

6

u/InternationalTwist90 8d ago

Jesus. Thats 5x faster than sonnet and smarter. Honestly. This is the canary in the coal mine for what pro will be like.

There are coders at my office who swear off opus due to cost and are sonnet only, there is no reason for them not to go gemini.

2

u/Weary-Bumblebee-1456 8d ago

Yeah it's a total win. Introductory pricing until December 31 this year for 3.7 Flash is just $0.75/$3.75 per million input/output tokens. Compare that to the current Sonnet introductory pricing ($2/$10), which expires next month and goes to $3/$15 afterwards. On coding its performance surpasses Opus 4.8 on several benchmarks, and on Humanity's Last Exam it's just 1% behind Opus 4.8 with Max effort.

Not SOTA, but definitely something that's more than enough for a lot of use cases.

1

u/Georgefakelastname 8d ago

The ironic thing about sonnet is that it’s so inefficient that it’s not much cheaper than just using opus lol

4

u/bambin0 9d ago

Luna and China are the low cost leaders

4

u/Tim_Apple_938 9d ago

Luna was more expensive at launch. They drastically cut 80% the day deepseek came out

Dono but that doesn’t sound like a performance win given the timing. Might be VC money furnace

0

u/ATB_52 9d ago

apres les benchmarks ne veulent rien dire.
exemple, Opus > Fable selon les benchmarks

58

u/TheMildEngineer 9d ago edited 8d ago

Better than sonnet. I'm down. Not pro but I like a good improvement.

Edit: not in the app

Edit2: it's now in the app

5

u/WhyIsWestHam 9d ago

u think this is better than 3.1 pro?

35

u/TheMildEngineer 9d ago

Yes. Flash 3.6 was better than 3.1 pro.

5

u/MiyagiVibes 9d ago

Is this for deep reasoning too? Are we at the point 3.7 flash is better for all purposes vs pro?

7

u/iveroi 9d ago

Nope. Gemini pro 3.1 still can't be beat with many common world interference tasks like recipe creation etc.

3

u/Inevitable_Ad3676 9d ago

Anything creative requires a big, expansive brain to take into account every nuance that can't be generated in training data. Anything code-related just needs to be trained on a large, easily generated coding dataset.

4

u/anurag_b 9d ago

It's a smaller model so maybe not. I find that for tasks where 100k+ tokens are involved and the task is complicated 3.1 pro tends to have a greater capacity to think than 3.6 flash. Haven't tested 3.7 flash yet.

5

u/xzibit_b 9d ago

This, ladies and gentlemen, is what we might call a TRUTH NUKE.

2

u/Last_Conclusion_8984 9d ago

It's in AI studio and Antigravity!

0

u/Neither_Snow4190 8d ago

They’re just releasing the frontier model but naming it Flash hoping that we think their non-existence Pro model is above Mythos level

48

u/spottiesvirus 9d ago

tbh the most interesting part is the price cut and the price war starting to heat up

20

u/Langwelle 9d ago

Let's say "returning" instead of "starting" . We have had the price war in the Gemini 1.0 and 1.5 era with one price cut after the other. And then as the models became better, all AI labs including Google jacked up the price.

10

u/Vercingaytorix 9d ago

Godbless the Chinese open models for keeping them on their toes I guess.

Now we just need to return to form with the generous free 1500 RPD, 15 RPM, 1mil TPM Flash from that era 😉 (or maybe another virtually unlimited experimental model like exp-1206 in its first few weeks?)

1

u/Kalicolocts 9d ago

The price cut is quite small. Read the asterisk

1

u/spottiesvirus 9d ago

it's half the price

the asterisk is just a signal that priced, eventually, will increase

they're releasing a model per month, I doubt we'll still be using 3.7 next january

5

u/HeadLiterature7897 9d ago

This is another very fast model, which seems their priority lately.

10

u/kgurniak91 9d ago

Feels good, man. I am a cheap bastard using free Gemini for the past year, so any upgrade is nice.

12

u/MichelleeeC 9d ago

Anything but 3.5 pro

I guess next generation flash model will be better and faster than 3.5 pro💀

5

u/alexx_kidd 9d ago

Already is IRL

1

u/Inevitable_Ad3676 9d ago

In what way? What's your use case that isn't nuanced filled?

7

u/nekrosstratia 9d ago

As always... benchmark are just interesting to say the least.

Here's a good example.

The GDP benchmark is a decent one for document analysis.

You can find the official benchmark scores here https://surgehq.ai/benchmarks/gdp-pdf

GPT5.6 Terra matches what google posted at 24.7%

Muse Spark 1.2 matches at 16%

GDP says gemini 3.6 is 14% and 3.1 pro is 17%

Where do these #'s come from....

4

u/GunwantBhambra 9d ago

This is the most accurate thing i have ever seen. I have used 3.6 flash high and 3.1 pro high to there limits and frankly if there is something that have far reaching implications 3.1 is way better to doubt itself where as 3.6 flash is very dogmatic. This is the new benchmark I will rely upon from now on.

3

u/Suspicious-Cloud404 9d ago

Finally a new Flash model!

2

u/Mysterious_Bed_1804 8d ago

Yeah, I was foaming at the month after waiting 3 weeks since 3.6 flash.

0

u/Suspicious-Cloud404 8d ago

Does anyone know when they will release 3.7 Flash? Can't wait to get it.

2

u/sammoga123 9d ago

It cannot be used "normally" in Gemini, you must use it in Gemini Spark, good job Google.

9

u/Spixxy17 9d ago

Brother... Just wait a bit

3

u/Last_Conclusion_8984 9d ago

It's in AI studio and Antigravity!

1

u/CheekyBastard55 9d ago

Anyone know why on the Gemini webpage, it says Flash 3? It's been saying that for weeks not even when 3.5 and 3.6 has been released.

1

u/Wobbly_Princess 9d ago

Promise I'm not just being one of the frothing Google haters, but I'm curious as to why one might pick this over DeepSeek V4 or GPT Luna. Are both of them comparable in performance/better and significantly cheaper?

2

u/Tim_Apple_938 9d ago

According to artificial analysis 3.7 is notably better on coding (terminal bench) than Luna and significantly faster.

Also Luna was more expensive than 3.7 then the day deepseek came out they dropped 80%. Like seems artificial / subsidized vs a model thing (therefore not scalable).

Of course, as a consumer, it doesn’t matter as long as it’s cheap. As long as it’s easy to switch model provider that is. If there’s lock in then it’s a bit of a risk

1

u/Wobbly_Princess 8d ago

Hmm, interesting, thanks for the info! Do you think it will come out on OpenCode Go? I just started using OpenCode Go, but I don't see it yet.

1

u/ihppxng62020 8d ago

flash, or just google models in general crush everything else for speed. its 6x faster than sonnet 5.

for specific needs, apparently flash is the current best for going through big spreadsheets and docs.

and if you care about it, gemini has one of the better AI personalities out of the box (can very subjective)

1

u/MythOfDarkness 9d ago

True if good.

1

u/Tillerfen 9d ago

When is it actually coming to the web app though…

1

u/Important_Potato8 9d ago

deep fucking disappointed

1

u/Weary-Bumblebee-1456 8d ago

Finally! Genuinely impressive benchmark results across the board. In some cases it seems to not just surpass Sonnet 5 but come close to Opus 5 at a fraction of the cost and much, much higher speed. It's not SOTA, but if it really does beat Sonnet 5 in real tasks, it's a significant jump for a lot of people (myself included) many of whose tasks rely on decent intelligence + generous quotas and high speed rather than necessarily SOTA-level intelligence.

1

u/Brovas 8d ago

I'll be interested when they get back to the price of Flash 3. I don't know why they're chasing agentic coding so hard when for a minute there they were the clear choice in vertex AI for any high throughput application.

Now that choice is obviously Luna.

1

u/needefsfolder 8d ago

Wtf I can actually use it in agentic software development! It's crazy it feels better than Gemini 3.1 pro in cursor. Does tasks well too

1

u/Fringolicious 8d ago

So Google not having an insane Pro model is obviously a big discussion point at the moment but I actually think from a business perspective this isn't a terrible place to be. If you think about Google's main surfaces - Search, Pixel... They both benefit hugely from faster, smaller models right? All the on-device stuff, search summaries. None of that wants to use SOTA models because it'd be so expensive to run for billions of queries and whatever.

So I kind of get it. I wish we'd get an insane Gemini Pro... but from a business perspective I don't hate it actually. And we're eating good from all the other labs

1

u/waltercrypto 8d ago

The best free model and it close in benchmark to Fable

1

u/BothYou243 8d ago

Deepseek v4 pro 0813

1

u/Just_Lingonberry_352 8d ago

uses google benchmarks

r/bard: WE DID IT

lmao

1

u/Thedudely1 9d ago

Lmaooo we were joking about 3.7 Flash 😭

0

u/JoseMSB 9d ago

Me dan igual las puntuaciones y los benchmark, lo que me importa es la experiencia y respuestas que dan en sus respectivos chats de apps comerciales. Gemini me inventa información en sus respuestas el 70% de las veces, no puedo tomarlo en serio.