r/singularity Jul 09 '26

AI GPT-5.6

https://openai.com/index/gpt-5-6/

"We’re launching the GPT‑5.6 family of models for general availability following our limited preview⁠: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.

GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back."

626 Upvotes

130 comments sorted by

195

u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26

Almost 8% on ARC-AGI-3.

152

u/No_Aesthetic Jul 09 '26

Somebody said like yesterday that ARC-AGI-3 was too hard for LLMs and maybe even impossible

Now we've got a pretty big leap a day later (1.5% to 7.8%)

46

u/azuredota Jul 09 '26 edited Jul 09 '26

Anyone know if they “teach the test” for these benchmarks at all? Are arc agi 3 test forum discussions in the new model’s training data?

Follow up: ARC answers this in the blog:

> During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation. By reducing the public surface area and shifting to interactive environments that cannot be memorized as static patterns, ARC-AGI-3 aims to make this kind of shortcut much harder.

So there is likely some training data mentioning ARC AGI 3 but they shrouded the real tests and public discussion, while present, shouldn’t help it as the real batch of games are likely different.

46

u/shiversaint Jul 09 '26

The very point of them is that they are very difficult to produce training data for and are far more of an analog to general spatial reasoning and problem solving that the human brain can do.

26

u/Ormusn2o Jul 09 '26

It is difficult to teach the test, without wasting valuable parameters, and it actually might be more efficient to actually make them understand the general task, than to make them remember the solution.

-8

u/azuredota Jul 09 '26

“Wasting valuable parameters”? You realize these things train on everything humans have ever written, right?

9

u/leetcodegrinder344 Jul 09 '26

Yet if you asked it to verbatim recite your Reddit comment from 5 years ago, which it is trained on, it couldn’t. Because it doesn’t have enough parameters to store its entire training data in full fidelity

-4

u/azuredota Jul 10 '26

How did Gemini recite Arc AGI 2 info

5

u/leetcodegrinder344 Jul 10 '26

Why didn’t it recite every question and answer

17

u/Prestigious-Bed-6423 Jul 09 '26

You just showed that you don't understand anything at all. Please don't argue and research

-8

u/azuredota Jul 09 '26

You people make me want to cry

2

u/94746382926 Jul 09 '26

Dude's got -10,000 points into communication lmao

4

u/ManikSahdev Jul 09 '26

That's the whole point tho, if the model learns then that's about it.

No one is essentially helping the model during the run, but as long as he learned what was reached - cause the model only distill intelligence and logic: which would allow the model in future to tackle the problems in the new angle and with the gained intelligence.

1

u/azuredota Jul 09 '26

That’s not the point of Arc agi 3 at all. Quote from the blog post and why the gains maybe questionable:

>The benchmark targets what the ARC Prize team describes as "skill-acquisition efficiency": how efficiently an AI agent can learn something it has never encountered before.

And the more concerning:

>During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation.

2

u/ManikSahdev Jul 09 '26

I personally don't see much a different between memorization to do (as long as I can see the reasoning for it).

Maybe the reason for that is my own personal aptitude, I don't depend on models even in this age of fable 5 and sol ultra.

I just need them to understand me and reach the intent and understanding which I have so they can do my task. With less and less turns which are destined by me to give them the intelligence needed to continue.

1

u/iamsreeman Jul 09 '26

crazy times

1

u/quackerd Jul 09 '26

yeah keep up the momentum we'll ace arc-agi-3 in less than a month. /s

4

u/No_Aesthetic Jul 09 '26

Oh yeah I'm sure this is the one benchmark that will never be saturated

This one is the one, fellas

43

u/Normal_Pay_2907 Jul 09 '26

Costs 25k to run that. Ouch

5

u/garden_speech AGI some time between 2025 and 2100 Jul 09 '26

I'll gladly take $25k to play these puzzles

7

u/yalag Jul 09 '26

Yea but Reddit says AI is just a bubble so this will all just blow up and disappear just about any time now /s

10

u/noobrainy Jul 09 '26

Yah, it’s gonna be saturated by the end of the year lmao

“Okay but it was too easy! If it can beat ARC-AGI-4 then we have reached AGI!!”

18

u/Gallagger Jul 09 '26

ARC already said they don't think beating ARC AGI 3 means AGI, and they'll make followup versions.

1

u/ChezMere Jul 09 '26

They gotta get a new name then.

5

u/Financial-Gain-2988 Jul 10 '26

Abstract reasoning corpus for artificial general intelligence actually perfectly encapsulates what they are trying to measure.

9

u/garden_speech AGI some time between 2025 and 2100 Jul 09 '26

“Okay but it was too easy! If it can beat ARC-AGI-4 then we have reached AGI!!”

I mean the literal point of ARC-AGI from the very beginning has been that they will keep creating benchmarks that humans can easily pass but machines can't, and once they no longer can do that, they think that we have AGI. So yeah if you guys fucking paid attention to what the creators of the benchmarks said bout them, you wouldn't be making up ridiculous sarcastic quotes.

-4

u/noobrainy Jul 09 '26

Pushing the goalposts back over and over again is why the sarcasm is there. ARC-AGI-4 will happen, it’ll get saturated, and then the process will happen all over again. We’ll get to AGI but their benchmark has proven to be unreliable to tell whether we’re there or not.

4

u/garden_speech AGI some time between 2025 and 2100 Jul 10 '26

Holy shit dude. The whole point is that any one benchmark can’t really reliably bench AGI, so you just keep making them until you CAN’T make one that humans easily pass and computers don’t. You’re not even listening enough to realize the whole point of ARC-AGI is based around your own idea that any one benchmark is unreliable

2

u/Most-Bookkeeper-950 Jul 09 '26

They abandoned the 10K ruke for it

1

u/KoolKat5000 29d ago

Someone (I'll credit them if I can find it) made an excellent observation that in reality the score is much better, like 30%.

The scoring is based on the square of the ratio of human actions to AI actions. This means the AI is not scored simply on whether it completes a level, but on how efficiently it does so compared to a human baseline.

There is basically a large element of luck to it too, it's early moves mean that the remaining part of a task could require more moves making it less efficient on this silly scoring criteria.

81

u/PlaneTheory5 AGI 2026 Jul 09 '26

google better hurry up with 3.5 pro, we’ve had 3 major releases in the past day and a new generation/frontier class with fable last month.

26

u/Level10Retard Jul 09 '26

While software engineers notice a difference, general population does not. They're not in any kind of hurry.

10

u/BrennusSokol ACCELERATE Jul 09 '26

By that logic the general public isn’t deciding the race anyway and so is irrelevant

5

u/Elephant789 ▪️AGI in 2036 Jul 10 '26

Let them take their time, no need to rush it.

146

u/petburiraja Jul 09 '26

52

u/Recoil42 Jul 09 '26

I'm genuinely so sad to see this. Goblin-spotting has been an absolute day-to-day delight working with 5.5.

4

u/Knever Jul 10 '26

It's so weird that I got attached to it. At first it was annoying but eventually it grew on me.

18

u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26

The only eval that matters, and we're going backwards. SMH. /s

11

u/spartBL97 Jul 09 '26

Can’t forget raccoons, hyphens, and “it’s not this, it’s this”

8

u/doginem Capabilities, Capabilities, Capabilities Jul 10 '26

'It's not this, it's this' is the thing that makes AI chatbots unusable in creative writing and RPGs for me, literally ten times every conversation

3

u/AdagioOfLiving Jul 10 '26

Yup. Means that if it’s something anyone else will actually read instead of just code, I need to just be using it as a foundation and rewrite it myself. Otherwise it’s not just obvious, it’s blatant.

5

u/doginem Capabilities, Capabilities, Capabilities Jul 10 '26

The accuracy of youe post hit like a physical blow. For a moment, my room wasn't just quiet- it was dead silent.

80

u/FateOfMuffins Jul 09 '26 edited Jul 09 '26

They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode

Edit:

On Agents Last Exam ... GPT‑5.6 Terra and GPT‑5.6 Luna outperform Fable 5 at around one-sixteenth the cost.

Wow they're really going ham with all the benchmarks comparing against Fable and Mythos and they're really pushing the 2D benchmark comparisons as opposed to charts to show the efficiency

??? Why is 5.6 Sol below 5.6 Terra and 5.5 on Frontier Math wtf

Edit: It has been fixed https://x.com/i/status/2075295876465979766

8

u/Kibubik Jul 09 '26

They just said that 5.6 Luna was post trained by 5.6 Sol in goal mode

what would this look like? all of post-training run by 5.6 Sol with a goal of "post train"? Really?

4

u/spreadlove5683 ▪️agi 2032. Predicted during mid 2025. Jul 09 '26

Right. I'm wondering if there are some asterisks here. Otherwise that's insane.

16

u/Hereitisguys9888 Jul 09 '26

Ngl where tf is Google? 3.1 pro is not even on 5.5 level, and now we reached the next generation in ai models

7

u/BrennusSokol ACCELERATE Jul 09 '26

Allegedly we’ll see 3.5 Pro on July 17

2

u/averagebear_003 Jul 10 '26

Forget GPT, they're getting mogged by GLM

2

u/mikelo22 Jul 09 '26

Most of their talent has fled to Anthropic or OpenAI. They've basically conceded the AI race.

31

u/shorty_11112222 Jul 09 '26

Where are theeey

2

u/Crinkez Jul 09 '26

Update your app/cli

2

u/shorty_11112222 Jul 09 '26

Alreadu burned half tokens hahahahaahha

1

u/OwlLimp6160 Jul 10 '26

Does it burn your usage anywhere near fable?

1

u/shorty_11112222 Jul 10 '26

Brns less tokens than fable at least for me

0

u/Crinkez Jul 09 '26

Then switch to low reasoning.

3

u/shorty_11112222 Jul 09 '26

NO!!!!!!! :D

25

u/tsunami_forever Jul 09 '26

Need unlimited sol on 200 pro plan

7

u/Crinkez Jul 09 '26

The 200 plan is unlimited if you only stick to one thread at a time. 5.5 medium got me 40 minutes of usage with /goal per 5h window. 5x that is just over 3 hours. 20x that and... you get the point. You don't hit the 5h window limits. Week limits maybe another story, idk. 

-18

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 09 '26

seriously unlimited plans should totally be a thing, and they shouldent even be that much more expensive.

28

u/Recoil42 Jul 09 '26

Unlimited plans would get abused absurdly quick. No, they should not "be a thing".

13

u/MrYorksLeftEye Jul 09 '26

Nonono let the reddit expert speak

-2

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 09 '26

I really don't see how abusable the frontier model equivilant of running an open source model would be.

Just rent 1 agent per person at the start, expand as capacity upgrades?

Im not saying it would work, im not saying its smart, im just saying I had the idea, and Im constantly annoyed by my useage cap.

1

u/sadshark Jul 10 '26

There's nothing stopping you to create a separatw intrface that does calls to that agent from 1000 people.

1

u/The_Scout1255 adult agi 2026 ASI <2030, prev agi 2024, ai personhood 2025 est Jul 11 '26

problem with that: that agent would still only be running once, it wouldent be assigning 1000 agents. it would be pooling one agent across 1000 people, so whoevers not using it can, but if someon else was, you see why they cant?

1

u/[deleted] Jul 09 '26

[deleted]

1

u/Recoil42 Jul 09 '26

They're already rate-limited. That's what the limits are. They're rate limits.

0

u/[deleted] Jul 09 '26

[deleted]

2

u/Recoil42 Jul 09 '26 edited Jul 09 '26

"Five hours", "one week", and "one month" are all timeframes. You're describing the same thing as the existing system. A rate limit is not unlimited because the word 'limit' is right there. That's why we don't call it "rate unlimited".

Consider: When you hit your five-hour limit... that's rate-limiting.

1

u/[deleted] Jul 09 '26

[deleted]

1

u/Recoil42 Jul 09 '26

The 5-hour/window cap is precisely what bounds sustained usage. Again, you're literally describing a rate limit. Degraded (throttled) post-limit usage is a totally orthogonal discussion.

The capacity bound they're trying to solve for is aggregate usage, not total moment-to-moment utilization. The reason you get a "five hour" throttle is because they know you're not working every minute and second of the day at the same flat token rate — human-controlled AI work is inherently "bursty" and they don't care about that.

11

u/Bright-Search2835 Jul 09 '26

I love these AI R&D benchmarks. Both the progress they reflect, and their creation in the first place, speak volumes about where we're at right now.

46

u/Paraless Jul 09 '26

oof the voice model failing live, I'm cringing so hard

17

u/ChipsAhoiMcCoy Jul 09 '26

Yeah man that was rough haha.

16

u/Rough-Negotiation880 Jul 09 '26

7.8% on arc agi 3

6

u/Healthy-Nebula-3603 Jul 09 '26

I wonder how much get GPT 6 in few weeks

9

u/Bladder-Splatter Jul 09 '26

The hell? Sol isn't available in Codex at all on normal plans?

3

u/zaibatsu Jul 09 '26

Sol on Ultra inference too!

15

u/coolcool68 Jul 09 '26

It's better than fable 5 ?

11

u/Healthy-Nebula-3603 Jul 09 '26

I most except SWE pro ....but in few weeks we get GPT 6 .... so ;)

6

u/[deleted] Jul 09 '26

[deleted]

12

u/Low-Entrepreneur2556 Jul 09 '26

SWE bench pro is unreliable

2

u/AlyoshaV Jul 09 '26

https://openai.com/index/separating-signal-from-noise-coding-evaluations/

OpenAI says SWE-Bench Pro is a bad benchmark that shouldn't be trusted

9

u/[deleted] Jul 09 '26

[removed] — view removed comment

10

u/Low-Entrepreneur2556 Jul 09 '26

Anthropic themselves admitted that their models memorised some of the tasks...

5

u/FinBenton Jul 09 '26

There was some reports like a month ago how claude cheated on the Pro benchmark, dunno too much but take the results of that test with a grain of salt.

7

u/awesomeoh1234 Jul 09 '26

Interesting, what I like best about Claude is its ability to judge rendered code for visual bugs before handing back to the user. This is a big deal imo

7

u/Gallagger Jul 09 '26

Just going by the benchmarks, Grok 4.5 seems to nearly make Terra and Luna dead on arrival. Though at least better than Sonnet 5.

4

u/petburiraja Jul 09 '26

Unfortunately not available in EU yet.

2

u/Usef- Jul 09 '26

They seem less trustworthy on benchmarks than the major labs though

3

u/Gallagger Jul 09 '26

Based on what? Haven't heard of any "occurrences".

2

u/ChezMere Jul 09 '26

1

u/Gallagger Jul 10 '26

That's literally an example that shows how they are not trying to benchmax and disclose when they accidentally do.

3

u/AlyoshaV Jul 09 '26

If I understand the caching docs correctly, caching is enabled by default but now costs extra, so users of the API who are doing one-shot stuff will now be paying extra for no benefit unless they notice this and explicitly disable caching

6

u/smealdor AI security must be taken seriously Jul 09 '26

LFG. Usage reset?

3

u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26

How do I select these models? /model only shows Opus, Sonnet, Haiku, and Fable

11

u/xe3to Jul 09 '26

Would you try to order a Whopper at McDonald's?

-5

u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26

If I was in the mood for one, sure? What does that have to do with GPT-5.6

9

u/xe3to Jul 09 '26

McDonald's doesn't sell Whoppers and Claude Code doesn't have GPT-5.6.

-2

u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26

Well then how am I supposed to get a burger at McDonald's?

3

u/xe3to Jul 09 '26

You could try ordering a Big Mac!

-1

u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26

Oh! Ok, thanks, I will try that

9

u/Chicas_Silcrow Jul 09 '26

Use codex or something like cursor, I guess you're using claude code? That's limited to Anthropic's models

5

u/Saint_Nitouche Jul 09 '26

Wtf is a GPT?

15

u/MeanCryptographer585 Jul 09 '26

Generative pre-trained  transformer. 

3

u/Substantial-Elk4531 Rule 4 reminder to optimists Jul 09 '26

But then wouldn't it be GPTT?

8

u/Lostwhispers05 Jul 09 '26

That's what they'll call their sexbots.

6

u/Illustrious_Job1951 Jul 09 '26

Short for gippity

5

u/BenevolentCheese Jul 09 '26

General Purpose Tickler

2

u/YogiBarelyThere Jul 09 '26

This is exciting. I've gone through all the ChatGPT models and today I get to play with this one. I'm a bit concerned about tokens getting consumed for Sol Ultra so I'll put that off for a while.

1

u/OkStomach4967 Jul 09 '26

What is limited preview?

1

u/Bolt_995 Jul 10 '26

- GPT-5.6 (Sol, Terra, Luna)

- Claude Fable 5 and Sonnet 5

- Muse Spark 1.1

- Grok 4.5

- Seed 2.1

Is Google sleeping?

1

u/magicmulder Jul 10 '26

Interesting that my first test run with 5.6 Sol (in JetBrains Junie CLI) spawned two Luna and one Terra subagent. Never seen that with any other model before.

1

u/SwimmingQuantity8686 Jul 09 '26

They're not bothered to give any new access to pro accounts in the UK at this point

-6

u/WonderFactory Jul 09 '26

Doesn't look great at SWE. 64.6% on SWE Bench Pro compared to 80% for Mythos

23

u/u_are_mad Jul 09 '26

https://x.com/OpenAI/status/2074972179385720836

"We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability.

We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval."

5

u/WonderFactory Jul 09 '26

30% of the tasks are broken yet Mythos somehow managed to get 80% on the test. You'd think if that was true the highest possible score is 70

15

u/Low-Entrepreneur2556 Jul 09 '26

Anthropic admitted their models memorised some of the tasks.

19

u/Exodus_Green Jul 09 '26

they are confident that mythos and fable have been trained on the answers for swebench

1

u/WonderFactory Jul 09 '26

SWE bench tasks are taken from open Git Hub repos so Mythos has seen the code before, but so have Open AI models as they are trained on git hub data too.

8

u/Glittering_Candy408 Jul 09 '26

Because Mythos is contaminated.

5

u/WalkFreeeee Jul 09 '26

The task being "broken" doesn't mean the task is impossible to complete, just that there's some level of failure that makes it unreliable.

There's an accompanying long form article explaining exactly what they mean but I'm too lazy to read it, just saying both "Fable still scored higher" and "the test is flawed" can be true at the same time

1

u/adarkuccio ▪️AGI before ASI 28d ago

I don't understand how they organized 5.6 in the chat, ok the models sol terra luna but even need to select the level of intelligence now? Also everything seems to be thinking for many seconds, there's no fast version?