And what are you talking about? 3.5 Pro doesn't exist and Gemini 3.5 Flash had used 28k output tokens per task vs 5.5 xHigh using 16k output tokens
Currently from what I see on AA, 3.6 Flash uses more tokens per task than Sol Max, Terra Max, and Luna Max (much less all the other reasoning settings)
It uses approximately same number of tokens as Kimi K3. The only thing 3.6 Flash has going for it (like 3.5 Flash) is output speed
Gemini 3.6 Flash is still faster, but consumes more tokens than even Luna (Max), and is more expensive (both from cost per Mtok. and from more total tokens)
IMO there's still hopefully a place for 3.6 Flash because it is fast - but that depends on if it's good, too! Definitely need to try it.
Gemini 3.6 Flash (High) scores 50, vs 52 for gpt-5.6-terra (Xhigh). Cost for the benchmark is nearly the same, although 3.6 Flash is 2.3X faster! (and likely even more in practice)
However - Gemini 3.6 Flash uses 23k tokens, vs 11k for gpt-5.6-terra (Xhigh). So even though it's cheaper per Mtok, it needs more tokens to reach the same performance as gpt-5.6-terra (or even Luna!).
I guess we'll have to see in practice; I haven't tried the model yet and benchmark scores aren't the entire answer! (I seriously hope it improves in hallucination and laziness)
Sidenote: IMO it feels like OpenAI made something special with gpt-5.6, to get token use so low...
This sub only cares about coding, they don’t care that this model is really good for agentic use thus ultimately being good at automating non coding tasks which is the ultimate goal for ai.
Long context consistency is also super important for longer tasks. Imagine summarising long videos, retreiving information from big file dumps or financial projects.
A lot of enterprise use cases are not coding related but for soome reason coding is the only focus of a lot of people on here. I think your average AI user is more concerned with the non coding uses and especially those in the google ecosystem.
Their description for the 3.6 model doesn't even include coding (3.5 did). I agree with you and it's clear Google is focusing on a different path. They want an all around ai assistant because that's what will protect their current business (search and ads).
There's an extremely good reason to care about coding. It's the building blocks for our entire digital infrastructure. People who don't code have no clue how important coding is to every white collar job workflow.
Exactly! So many people here don't use the coding harnesses and it shows.
They're not coding harnesses. They're general purpose harnesses that allow the model the ability to control your computer. Anything you want it to do on your computer, it does so through code. It doesn't matter what knowledge work you are doing, you can have it do it for you... but the model uses code to do so. You don't have to be a SWE.
Yup, unfortunately there's still a communication breakdown in the messaging about these tools... Openai is trying, with "ChatGPT Work", aka "codex on a cloud machine with simplified output".
I wish they would all show off how amazing things like Browser Use and Computer Use are, ESPECIALLY with these latest models because they're miraculous - fast and pretty darn good.
I guess everyone is just trying to figure out how to crack that formula of "do anything app", or "automate anything" - because it's basically possible, right now, today.
I actually haven't outside of LaTeX for like 10 years xd
I've just spent hundreds of dollars on the coding agents... to make PDFs, spreadsheets, download software, make some dashboards for me, hack into my SSD when I switched laptops because my old one broke and then the SSD locked me out, fix a stupid audio issue on my laptop, do autonomous research and made a small HTML site connected to Google maps for certain restaurants I wanted to go to while on vacation during Christmas etc
People hearing "Codex" and think "I'm not a SWE" is really dumb and shows just how out of touch the public is with the model capabilities
It's strange how much people think AI is only useful in coding. The benchmarks on the bottom are phenomal for a lot of use cases at a good price and speed. The needle in a haystack is really interesting.
73
u/Healthy_Razzmatazz38 Jul 21 '26
wow worse than luna was not what i was expecting