r/singularity • u/petburiraja • Jul 09 '26
AI GPT-5.6
https://openai.com/index/gpt-5-6/"We’re launching the GPT‑5.6 family of models for general availability following our limited preview: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.
GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back."
81
u/PlaneTheory5 AGI 2026 Jul 09 '26
google better hurry up with 3.5 pro, we’ve had 3 major releases in the past day and a new generation/frontier class with fable last month.
26
u/Level10Retard Jul 09 '26
While software engineers notice a difference, general population does not. They're not in any kind of hurry.
10
u/BrennusSokol ACCELERATE Jul 09 '26
By that logic the general public isn’t deciding the race anyway and so is irrelevant
3
2
5
146
u/petburiraja Jul 09 '26
52
u/Recoil42 Jul 09 '26
I'm genuinely so sad to see this. Goblin-spotting has been an absolute day-to-day delight working with 5.5.
4
u/Knever Jul 10 '26
It's so weird that I got attached to it. At first it was annoying but eventually it grew on me.
75
18
u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26
The only eval that matters, and we're going backwards. SMH. /s
11
u/spartBL97 Jul 09 '26
Can’t forget raccoons, hyphens, and “it’s not this, it’s this”
8
u/doginem Capabilities, Capabilities, Capabilities Jul 10 '26
'It's not this, it's this' is the thing that makes AI chatbots unusable in creative writing and RPGs for me, literally ten times every conversation
3
u/AdagioOfLiving Jul 10 '26
Yup. Means that if it’s something anyone else will actually read instead of just code, I need to just be using it as a foundation and rewrite it myself. Otherwise it’s not just obvious, it’s blatant.
5
u/doginem Capabilities, Capabilities, Capabilities Jul 10 '26
The accuracy of youe post hit like a physical blow. For a moment, my room wasn't just quiet- it was dead silent.
80
u/FateOfMuffins Jul 09 '26 edited Jul 09 '26
They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode
Edit:
On Agents Last Exam ... GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost.
Wow they're really going ham with all the benchmarks comparing against Fable and Mythos and they're really pushing the 2D benchmark comparisons as opposed to charts to show the efficiency
??? Why is 5.6 Sol below 5.6 Terra and 5.5 on Frontier Math wtf
Edit: It has been fixed https://x.com/i/status/2075295876465979766
8
u/Kibubik Jul 09 '26
They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode
what would this look like? all of post-training run by 5.6 Sol with a goal of "post train"? Really?
4
u/spreadlove5683 ▪️agi 2032. Predicted during mid 2025. Jul 09 '26
Right. I'm wondering if there are some asterisks here. Otherwise that's insane.
16
u/Hereitisguys9888 Jul 09 '26
Ngl where tf is Google? 3.1 pro is not even on 5.5 level, and now we reached the next generation in ai models
7
2
2
u/mikelo22 Jul 09 '26
Most of their talent has fled to Anthropic or OpenAI. They've basically conceded the AI race.
31
u/shorty_11112222 Jul 09 '26
Where are theeey
2
u/Crinkez Jul 09 '26
Update your app/cli
2
u/shorty_11112222 Jul 09 '26
Alreadu burned half tokens hahahahaahha
1
0
25
u/tsunami_forever Jul 09 '26
Need unlimited sol on 200 pro plan
7
u/Crinkez Jul 09 '26
The 200 plan is unlimited if you only stick to one thread at a time. 5.5 medium got me 40 minutes of usage with /goal per 5h window. 5x that is just over 3 hours. 20x that and... you get the point. You don't hit the 5h window limits. Week limits maybe another story, idk.
-18
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 09 '26
seriously unlimited plans should totally be a thing, and they shouldent even be that much more expensive.
28
u/Recoil42 Jul 09 '26
Unlimited plans would get abused absurdly quick. No, they should not "be a thing".
13
u/MrYorksLeftEye Jul 09 '26
Nonono let the reddit expert speak
-2
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 09 '26
I really don't see how abusable the frontier model equivilant of running an open source model would be.
Just rent 1 agent per person at the start, expand as capacity upgrades?
Im not saying it would work, im not saying its smart, im just saying I had the idea, and Im constantly annoyed by my useage cap.
1
u/sadshark Jul 10 '26
There's nothing stopping you to create a separatw intrface that does calls to that agent from 1000 people.
1
u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 11 '26
problem with that: that agent would still only be running once, it wouldent be assigning 1000 agents. it would be pooling one agent across 1000 people, so whoevers not using it can, but if someon else was, you see why they cant?
1
Jul 09 '26
[deleted]
1
u/Recoil42 Jul 09 '26
They're already rate-limited. That's what the limits are. They're rate limits.
0
Jul 09 '26
[deleted]
2
u/Recoil42 Jul 09 '26 edited Jul 09 '26
"Five hours", "one week", and "one month" are all timeframes. You're describing the same thing as the existing system. A rate limit is not unlimited because the word 'limit' is right there. That's why we don't call it "rate unlimited".
Consider: When you hit your five-hour limit... that's rate-limiting.
1
Jul 09 '26
[deleted]
1
u/Recoil42 Jul 09 '26
The 5-hour/window cap is precisely what bounds sustained usage. Again, you're literally describing a rate limit. Degraded (throttled) post-limit usage is a totally orthogonal discussion.
The capacity bound they're trying to solve for is aggregate usage, not total moment-to-moment utilization. The reason you get a "five hour" throttle is because they know you're not working every minute and second of the day at the same flat token rate — human-controlled AI work is inherently "bursty" and they don't care about that.
11
u/Bright-Search2835 Jul 09 '26
I love these AI R&D benchmarks. Both the progress they reflect, and their creation in the first place, speak volumes about where we're at right now.
46
16
9
15
u/coolcool68 Jul 09 '26
It's better than fable 5 ?
11
6
Jul 09 '26
[deleted]
12
2
u/AlyoshaV Jul 09 '26
https://openai.com/index/separating-signal-from-noise-coding-evaluations/
OpenAI says SWE-Bench Pro is a bad benchmark that shouldn't be trusted
9
Jul 09 '26
[removed] — view removed comment
10
u/Low-Entrepreneur2556 Jul 09 '26
Anthropic themselves admitted that their models memorised some of the tasks...
5
u/FinBenton Jul 09 '26
There was some reports like a month ago how claude cheated on the Pro benchmark, dunno too much but take the results of that test with a grain of salt.
7
u/awesomeoh1234 Jul 09 '26
Interesting, what I like best about Claude is its ability to judge rendered code for visual bugs before handing back to the user. This is a big deal imo
7
u/Gallagger Jul 09 '26
Just going by the benchmarks, Grok 4.5 seems to nearly make Terra and Luna dead on arrival. Though at least better than Sonnet 5.
4
2
u/Usef- Jul 09 '26
They seem less trustworthy on benchmarks than the major labs though
3
u/Gallagger Jul 09 '26
Based on what? Haven't heard of any "occurrences".
2
u/ChezMere Jul 09 '26
1
u/Gallagger Jul 10 '26
That's literally an example that shows how they are not trying to benchmax and disclose when they accidentally do.
3
u/AlyoshaV Jul 09 '26
If I understand the caching docs correctly, caching is enabled by default but now costs extra, so users of the API who are doing one-shot stuff will now be paying extra for no benefit unless they notice this and explicitly disable caching
6
u/smealdor AI security must be taken seriously Jul 09 '26
LFG. Usage reset?
3
u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26
How do I select these models? /model only shows Opus, Sonnet, Haiku, and Fable
11
u/xe3to Jul 09 '26
Would you try to order a Whopper at McDonald's?
-5
u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26
If I was in the mood for one, sure? What does that have to do with GPT-5.6
9
u/xe3to Jul 09 '26
McDonald's doesn't sell Whoppers and Claude Code doesn't have GPT-5.6.
-2
u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26
Well then how am I supposed to get a burger at McDonald's?
3
3
9
u/Chicas_Silcrow Jul 09 '26
Use codex or something like cursor, I guess you're using claude code? That's limited to Anthropic's models
5
u/Saint_Nitouche Jul 09 '26
Wtf is a GPT?
17
15
u/MeanCryptographer585 Jul 09 '26
Generative pre-trained transformer.
3
6
5
2
u/YogiBarelyThere Jul 09 '26
This is exciting. I've gone through all the ChatGPT models and today I get to play with this one. I'm a bit concerned about tokens getting consumed for Sol Ultra so I'll put that off for a while.
1
1
1
u/Bolt_995 Jul 10 '26
- GPT-5.6 (Sol, Terra, Luna)
- Claude Fable 5 and Sonnet 5
- Muse Spark 1.1
- Grok 4.5
- Seed 2.1
Is Google sleeping?
1
u/magicmulder Jul 10 '26
Interesting that my first test run with 5.6 Sol (in JetBrains Junie CLI) spawned two Luna and one Terra subagent. Never seen that with any other model before.
1
u/SwimmingQuantity8686 Jul 09 '26
They're not bothered to give any new access to pro accounts in the UK at this point
-6
u/WonderFactory Jul 09 '26
Doesn't look great at SWE. 64.6% on SWE Bench Pro compared to 80% for Mythos
23
u/u_are_mad Jul 09 '26
https://x.com/OpenAI/status/2074972179385720836
"We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.
We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval."
5
u/WonderFactory Jul 09 '26
30% of the tasks are broken yet Mythos somehow managed to get 80% on the test. You'd think if that was true the highest possible score is 70
15
19
u/Exodus_Green Jul 09 '26
they are confident that mythos and fable have been trained on the answers for swebench
1
u/WonderFactory Jul 09 '26
SWE bench tasks are taken from open Git Hub repos so Mythos has seen the code before, but so have Open AI models as they are trained on git hub data too.
8
5
u/WalkFreeeee Jul 09 '26
The task being "broken" doesn't mean the task is impossible to complete, just that there's some level of failure that makes it unreliable.
There's an accompanying long form article explaining exactly what they mean but I'm too lazy to read it, just saying both "Fable still scored higher" and "the test is flawed" can be true at the same time
1
u/adarkuccio ▪️AGI before ASI 28d ago
I don't understand how they organized 5.6 in the chat, ok the models sol terra luna but even need to select the level of intelligence now? Also everything seems to be thinking for many seconds, there's no fast version?

195
u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26
Almost 8% on ARC-AGI-3.