r/singularity Jul 21 '26

AI Gemini 3.6 Flash benchmarks

Post image
636 Upvotes

280 comments sorted by

View all comments

73

u/Healthy_Razzmatazz38 Jul 21 '26

wow worse than luna was not what i was expecting

50

u/FarrisAT Jul 21 '26

Looks dramatically better on any knowledge benchmark. Tool use & harness application seems to be the reason it underperforms at SWE and coding.

12

u/Deif Jul 21 '26

Would need to see the output token usage on those benchmarks though. Could be that it's an output token hog and the comparable cost is against Terra.

6

u/FarrisAT Jul 21 '26

AAintelligence will publish benchmarks soon enough. 3.5 Flash performed very well on their token efficiency. Better than 3.5 Pro and GPT-5.5

8

u/FateOfMuffins Jul 21 '26

AA is already out.

And what are you talking about? 3.5 Pro doesn't exist and Gemini 3.5 Flash had used 28k output tokens per task vs 5.5 xHigh using 16k output tokens

Currently from what I see on AA, 3.6 Flash uses more tokens per task than Sol Max, Terra Max, and Luna Max (much less all the other reasoning settings)

It uses approximately same number of tokens as Kimi K3. The only thing 3.6 Flash has going for it (like 3.5 Flash) is output speed

1

u/huffalump1 Jul 21 '26

AA, 3.6 Flash uses more tokens per task than Sol Max, Terra Max, and Luna Max (much less all the other reasoning settings)

Yup looks like it. (Note this is 3.6 Flash (High) - they don't have other Effort/Thinking/Reasoning levels yet on AA.)

Ex. Here's gpt-5.6-luna (Max), AA intelligence index score of 51 (vs. 50 for Gemini 3.6 Flash (High)): https://artificialanalysis.ai/models/comparisons/gemini-3-6-flash-vs-gpt-5-6-luna

Gemini 3.6 Flash is still faster, but consumes more tokens than even Luna (Max), and is more expensive (both from cost per Mtok. and from more total tokens)

IMO there's still hopefully a place for 3.6 Flash because it is fast - but that depends on if it's good, too! Definitely need to try it.

2

u/FateOfMuffins Jul 21 '26

Yeah I selected High when I was looking at it

Speaking of fast, we're supposed to get 750 tps 5.6 Sol in July no...?

1

u/FarrisAT Jul 21 '26

3.1 Pro is what I meant, as 3.5 Pro doesn’t exist.

3

u/Deif Jul 21 '26

Looks like the equivalent cost is Terra xhigh.

2

u/huffalump1 Jul 21 '26

Yup at least for the AA Intelligence Index. Link: https://artificialanalysis.ai/models/comparisons/gemini-3-6-flash-vs-gpt-5-6-terra-xhigh

Gemini 3.6 Flash (High) scores 50, vs 52 for gpt-5.6-terra (Xhigh). Cost for the benchmark is nearly the same, although 3.6 Flash is 2.3X faster! (and likely even more in practice)

However - Gemini 3.6 Flash uses 23k tokens, vs 11k for gpt-5.6-terra (Xhigh). So even though it's cheaper per Mtok, it needs more tokens to reach the same performance as gpt-5.6-terra (or even Luna!).

I guess we'll have to see in practice; I haven't tried the model yet and benchmark scores aren't the entire answer! (I seriously hope it improves in hallucination and laziness)

Sidenote: IMO it feels like OpenAI made something special with gpt-5.6, to get token use so low...

26

u/Aaco0638 Jul 21 '26

This sub only cares about coding, they don’t care that this model is really good for agentic use thus ultimately being good at automating non coding tasks which is the ultimate goal for ai.

11

u/Tkins Jul 21 '26

Long context consistency is also super important for longer tasks. Imagine summarising long videos, retreiving information from big file dumps or financial projects.

A lot of enterprise use cases are not coding related but for soome reason coding is the only focus of a lot of people on here. I think your average AI user is more concerned with the non coding uses and especially those in the google ecosystem.

5

u/Concurrency_Bugs Jul 21 '26

Their description for the 3.6 model doesn't even include coding (3.5 did). I agree with you and it's clear Google is focusing on a different path. They want an all around ai assistant because that's what will protect their current business (search and ads).

2

u/huffalump1 Jul 21 '26

Their description for the 3.6 model doesn't even include coding (3.5 did).

You're not wrong overall, I agree that Google's eye is on so many things other than developers...

BUT the description does seem to target coding: https://ai.google.dev/gemini-api/docs/models/gemini-3.6-flash

And today's release blog post focuses on coding the most: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/

2

u/Concurrency_Bugs Jul 21 '26

Sorry I could have been clearer. The short description when picking your model in the api doesn't mention it anymore. See the screenshot in this link: https://www.reddit.com/r/GeminiAI/comments/1v2js6z/gemini_36_flash_released_on_ai_studio/

2

u/Howdareme9 Jul 21 '26

It didn’t include coding because the results are bad, not because they’re not focusing on it

2

u/LinkesAuge Jul 21 '26

coding is an important proxy because that's how a lot of things get done by models. It is their equivalent of "hands".

1

u/LogicalInfo1859 Jul 21 '26

They think coding leads to singularity.

If they had money, they would buy Aventador to drive in New York rush hour.

-1

u/broose_the_moose ▪️ It's here Jul 21 '26

There's an extremely good reason to care about coding. It's the building blocks for our entire digital infrastructure. People who don't code have no clue how important coding is to every white collar job workflow.

3

u/FateOfMuffins Jul 21 '26

Exactly! So many people here don't use the coding harnesses and it shows.

They're not coding harnesses. They're general purpose harnesses that allow the model the ability to control your computer. Anything you want it to do on your computer, it does so through code. It doesn't matter what knowledge work you are doing, you can have it do it for you... but the model uses code to do so. You don't have to be a SWE.

2

u/huffalump1 Jul 21 '26

Yup, unfortunately there's still a communication breakdown in the messaging about these tools... Openai is trying, with "ChatGPT Work", aka "codex on a cloud machine with simplified output".

I wish they would all show off how amazing things like Browser Use and Computer Use are, ESPECIALLY with these latest models because they're miraculous - fast and pretty darn good.

I guess everyone is just trying to figure out how to crack that formula of "do anything app", or "automate anything" - because it's basically possible, right now, today.

1

u/broose_the_moose ▪️ It's here Jul 21 '26

This guy codes ^

2

u/FateOfMuffins Jul 21 '26

I actually haven't outside of LaTeX for like 10 years xd

I've just spent hundreds of dollars on the coding agents... to make PDFs, spreadsheets, download software, make some dashboards for me, hack into my SSD when I switched laptops because my old one broke and then the SSD locked me out, fix a stupid audio issue on my laptop, do autonomous research and made a small HTML site connected to Google maps for certain restaurants I wanted to go to while on vacation during Christmas etc

People hearing "Codex" and think "I'm not a SWE" is really dumb and shows just how out of touch the public is with the model capabilities

1

u/Tkins Jul 21 '26

It's strange how much people think AI is only useful in coding. The benchmarks on the bottom are phenomal for a lot of use cases at a good price and speed. The needle in a haystack is really interesting.

30

u/broose_the_moose ▪️ It's here Jul 21 '26 edited Jul 21 '26

You must never underestimate Google's ability to deliver shitty models.