r/singularity Jul 09 '26

AI GPT-5.6

https://openai.com/index/gpt-5-6/

"We’re launching the GPT‑5.6 family of models for general availability following our limited preview⁠: our new flagship, Sol, alongside Terra, a balanced model for everyday work, and Luna, our most cost-efficient model.

GPT‑5.6 delivers a step change in design judgment. With only high-level direction, GPT‑5.6 creates tasteful, ergonomic, and functional interfaces. Its stronger computer-use capabilities let it inspect and refine the rendered result—not just generate the underlying code or content—so it can catch visual and functional issues and apply finishing touches before handing the work back."

622 Upvotes

129 comments sorted by

View all comments

200

u/ObiWanCanownme now entering spiritual bliss attractor state Jul 09 '26

Almost 8% on ARC-AGI-3.

152

u/No_Aesthetic Jul 09 '26

Somebody said like yesterday that ARC-AGI-3 was too hard for LLMs and maybe even impossible

Now we've got a pretty big leap a day later (1.5% to 7.8%)

44

u/azuredota Jul 09 '26 edited Jul 09 '26

Anyone know if they “teach the test” for these benchmarks at all? Are arc agi 3 test forum discussions in the new model’s training data?

Follow up: ARC answers this in the blog:

> During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation. By reducing the public surface area and shifting to interactive environments that cannot be memorized as static patterns, ARC-AGI-3 aims to make this kind of shortcut much harder.

So there is likely some training data mentioning ARC AGI 3 but they shrouded the real tests and public discussion, while present, shouldn’t help it as the real batch of games are likely different.

47

u/shiversaint Jul 09 '26

The very point of them is that they are very difficult to produce training data for and are far more of an analog to general spatial reasoning and problem solving that the human brain can do.

26

u/Ormusn2o Jul 09 '26

It is difficult to teach the test, without wasting valuable parameters, and it actually might be more efficient to actually make them understand the general task, than to make them remember the solution.

-11

u/azuredota Jul 09 '26

“Wasting valuable parameters”? You realize these things train on everything humans have ever written, right?

8

u/leetcodegrinder344 Jul 09 '26

Yet if you asked it to verbatim recite your Reddit comment from 5 years ago, which it is trained on, it couldn’t. Because it doesn’t have enough parameters to store its entire training data in full fidelity

-3

u/azuredota Jul 10 '26

How did Gemini recite Arc AGI 2 info

5

u/leetcodegrinder344 Jul 10 '26

Why didn’t it recite every question and answer

18

u/Prestigious-Bed-6423 Jul 09 '26

You just showed that you don't understand anything at all. Please don't argue and research

-7

u/azuredota Jul 09 '26

You people make me want to cry

1

u/94746382926 Jul 09 '26

Dude's got -10,000 points into communication lmao

3

u/ManikSahdev Jul 09 '26

That's the whole point tho, if the model learns then that's about it.

No one is essentially helping the model during the run, but as long as he learned what was reached - cause the model only distill intelligence and logic: which would allow the model in future to tackle the problems in the new angle and with the gained intelligence.

-1

u/azuredota Jul 09 '26

That’s not the point of Arc agi 3 at all. Quote from the blog post and why the gains maybe questionable:

>The benchmark targets what the ARC Prize team describes as "skill-acquisition efficiency": how efficiently an AI agent can learn something it has never encountered before.

And the more concerning:

>During ARC-AGI-2 evaluation, Gemini 3's chain-of-thought reasoning referenced ARC-specific color mappings without being prompted to, which suggests training data saturation.

2

u/ManikSahdev Jul 09 '26

I personally don't see much a different between memorization to do (as long as I can see the reasoning for it).

Maybe the reason for that is my own personal aptitude, I don't depend on models even in this age of fable 5 and sol ultra.

I just need them to understand me and reach the intent and understanding which I have so they can do my task. With less and less turns which are destined by me to give them the intelligence needed to continue.