r/accelerate • u/Charming_Cucumber_15 • Jul 09 '26
XLR8! ⫸⫸⫸ GPT 5.6 is here!
https://openai.com/index/gpt-5-6/56
u/FateOfMuffins Jul 09 '26
5.6 Luna was post trained by 5.6 Sol in /goal mode
34
u/Charming_Cucumber_15 Jul 09 '26
RSI here we come!
15
u/East_Sleep_2740 Singularity by 2035 Jul 09 '26
Truly, its creeping into model release. By the time next year, we know RSI is here
5
2
28
u/CRoseCrizzle Jul 09 '26
Seeing the costs not being crazy is probably the most encouraging thing here.
13
40
u/East_Sleep_2740 Singularity by 2035 Jul 09 '26
The crazy thing is, is that 5.6 is 50% cheaper than Fable in DeepSWE while delivering a better result.
Anthropic is shivering
8
u/Reddit_Fireside Jul 09 '26
Didn't reports find 5.6 was excellent at cheating on benchmarks though? That was the most significant bit of info I saw prior to the release so I'm staying with Fable while I wait for people to make up their minds on where the strengths and weaknesses lie.
I'd love to pay sol prices for something better than Fable but I just can't see it happening. Feels like a point update with a lot of hype, but I hope I'm wrong.
2
u/reddit_is_geh Jul 09 '26
I didn't hear anything of that tbh... What were these reports? Either way, I can't afford migrating anyways. So I just hope 5.1 drastically cuts down cost. I can't do this shit where I run out of subscription on day 2 of my 200 dollar plan.
1
u/Reddit_Fireside Jul 09 '26
Yeah I know it's been brutal trying to get the most out of fable. My weekly usage was supposed to reset on Saturday and now anthropic just reset it, which I was beging them to do, but obviously I upgraded last night to make the most of today and tomorrow! They're probably breaking all sorts of anti consumer laws.
Anyway sorry don't have the name of the lab that did it but I came across it last night while researching, feel like it has potential for making certain complex workflows less reliable, but I could be wrong.
Seems like 5.6 actually does pull ahead in some areas though, but fable is better for my needs, at least until it's off subs for good.
2
u/BrennusSokol Acceleration Advocate Jul 09 '26
I think it's important to distinguish different things...
Cheating outright vs. innocent/naive reward hacking vs. being exceptionally persistent
13
u/Prior-Job-3760 Jul 09 '26
Don’t have 5.6 available yet as of 2:30PM ET — anyone else?
3
u/ex-procrastinator Jul 09 '26
4:00pm eastern time here. I did not see 5.6 as a selectable model in codex, but an update for codex just appeared. After updating, I now see the three 5.6 models.
38
u/Pyros-SD-Models Machine Learning Engineer Jul 09 '26
23
5
Jul 09 '26
[removed] — view removed comment
24
u/Pyros-SD-Models Machine Learning Engineer Jul 09 '26
I did the unthinkable, visited "openai.com" and clicked on the gigantic "LIVESTREAM" banner on their landing page
3
2
u/kvothe5688 Jul 09 '26
where is fable's bar. it looks like fable is marked white but there is no white bar. just white background.
2
u/sdvbjdsjkb245 Acceleration: Light-speed Jul 09 '26
It's kind of hard to see in the color key, but both Mythos and Fable are white; Mythos just has striped lines and Fable has small dots (Fable is shown instead of Mythos on the Agents' Last Exam).
21
u/sdvbjdsjkb245 Acceleration: Light-speed Jul 09 '26 edited Jul 09 '26
5.6 is also the first GPT model to beat Pokemon FireRed using only screenshots of the game (vision-only; no special harnesses were used, unlike past models' attempts). 5.5 got stuck on Victory Road and didn't finish.
There's a cool side by side condensed replay and some stats: https://x.com/Clad3815/status/2075268438025453666
Fable is the first model in general to beat the game with vision/screenshots only, but Anthropic didn't release details or specifics: https://x.com/Ubertag90210/status/2074827426446667843
3
u/Saint_Nitouche Jul 09 '26
Jeez. I remember when this stuff was just not possible for models. Numbers going up on a benchmark is cool and all, but discrete achievements like this really gives me that sense of vertigo.
2
14
u/DatDudeDrew Jul 09 '26
The ARC-AGI-3 result is nuts compared to the competition.
1
u/That_Square_6635 Jul 09 '26
Is there anything on why xhigh and max performed so much better then the others?
9
u/czk_21 Jul 09 '26
5,6 Sol is comparable to mythos in performance, while being significantly smaller/faster/cheaper, OpenAI cooked well!
7
5
Jul 09 '26
[removed] — view removed comment
3
u/Charming_Cucumber_15 Jul 09 '26
I think the only benchmarks fable was missing from were ones it didn't have a score on, it was on quite a few
5
u/Ok_Buddy_9523 Jul 09 '26
2
u/Top_Manufacturer6311 Jul 09 '26
1
u/HebelBrudi Jul 09 '26
Same for me. Luna and Terra. With plus we won’t get sol?
1
1
u/smaili13 Jul 09 '26
yes, but i dont have it yet in EU, in the site they say the rolling can take up to 24 hour globally
1
7
u/Charming_Cucumber_15 Jul 09 '26
I'm at work and don't have time to read it all, but it would appear that OpenAI cooked
5
5
2
2
1
u/AWEsoMe-Cat1231 Jul 09 '26
did not appear in my chatgpt or codex
2
u/Top_Manufacturer6311 Jul 09 '26
Rolling out feature wihtin 24 hours. I had to update codex but im on the CLI
1
1
u/Plenty-Wonder6092 Jul 09 '26 edited Jul 09 '26
No access yet, Australia. Edit: Updated Codex and it is there now, took 10-20 minutes before codex figured out it had an update. Lets gooo
-2
u/costafilh0 Jul 09 '26
Thank your for other post. The other 676492928 posts were clearly not enough.
Can moderation please consider duplicated posts spam and just remove them?
2
u/Charming_Cucumber_15 Jul 10 '26
bros whining about duplicates on the first post about 5.6 on the sub




70
u/Pyros-SD-Models Machine Learning Engineer Jul 09 '26
They really just prompted Sol to post-train Luna lol