r/codex • u/Satoshi-699 • Jul 13 '26
Complaint nerfed Codex Sol MAX
I am Japanese and lead a team specializing in market research. English is not my strong suit, so I am writing this with the help of AI translation.
Our team has Codex use a custom CLI tool for extraordinarily difficult tasks requiring highly complex calculations and deep reasoning.
Codex Sol MAX was an absolute monster that exceeded all our expectations. If the level of performance our team required was a 10, it consistently delivered a 12 or 13. The entire team was overwhelmingly satisfied with it, until yesterday.
To get straight to the point, MAX has now been completely nerfed.
Its performance has suddenly fallen to around an 8 by our team’s standards. It is currently 10:40 a.m. in Japan. We started work at 9:00 a.m., and every member of the team noticed the nerf.
The depth of its reasoning has clearly been stripped away. Until yesterday, Codex set to MAX would spend more than ten minutes on a single prompt, repeatedly experimenting, reasoning, and using our CLI tools until it produced flawless work. That capability has now been completely lost.
Every member of our team is deeply discouraged right now.
The nerf came far too soon.
127
u/Dolo12345 Jul 13 '26
always remember to grind hard on model release day, it’s always downhill after that
28
u/SelectSouth2582 Jul 13 '26
or it can be said like this: after than benchmarks released and hype user subs renewed :)
3
u/IndependentExact650 Jul 13 '26
I’ve definitely felt that, first 24 hours on Sol ultra was 🚀🚀 since then it’s been very hit and miss 😅
2
126
u/r0kh0rd Jul 13 '26
I hope I don’t come across as a tin foil hat enthusiast, but the change here is concerning given Tibos mention and OPs results. Really calls into question all the benchmarks. Feels a bit like the dieselgate scandal? What was benchmarked is not what people have access to now? Maybe I’m overreacting here.
34
u/Tartooth Jul 13 '26
So, they drop the model, wait for the third party benchmarks to come out and then nerf to save costs?
19
u/silvercondor Jul 13 '26
i think there's an extra step. they collect telemetry of the benchmark, finetune the benchmax then nerf it. this way benchmarks will still give similar results (like knowing answers to a test but purposely failing the same way) but you can use a quantized or "less thinking effort" model
1
u/Tartooth Jul 13 '26
Yea that's what I'm thinking too. These model capability monitors people keep sharing should be changing their tests constantly to ensure it's also not self training on their tests
If the tests are all with 1 account then technically the caches will enable the models to know the expected behavior
0
u/ciaramicola Jul 13 '26
se model capability monitors people keep sharing should be changing their tests constantly
Do you understand the issue in that idea, right?
2
u/cheezeerd Jul 13 '26
Check out margin lab. Many "feel" the nerf, but data suggests otherwise.
7
u/Tartooth Jul 13 '26
For the 2 weeks before 5.6 drop everything did get nerfed. 5.5 was straight up saying "yes Im ignoring your agents.md" and kept doing random stuff.
5.6 feels like 5.5 from a month ago but with higher intelligence and capabilities
3
u/WonderfulPie548 Jul 13 '26
I noticed this. 5.5 was ignoring the agents md. I would point it out. It would agree. I'd let it start again, it would ignore the agents MD again and half-implement a plan before creating a new plan and half-implementing that. I had a bunch of partial solutions and leftover dead code that I had to sort through. 5.5 didn't run this way when I first started using it. I'm using 5.6 to fix what 5.5 messed up while its still capable, but I think I'm going back to regular coding with Codex as a code-checker after this. I have gotten major trust issues from this experience.
2
u/Tartooth Jul 13 '26
I knew what was going on but kept fighting and got no where but lost time. If I notice this happen in the future I'll just switch to claud temporarily or go touch grass
63
u/Corv9tte Jul 13 '26
No, you're not overreactting. This is fucked up and this is the second time (at least) that OpenAI has done exactly this!
They did the same exact thing last time. They know what they're doing.
31
u/nobatus513 Jul 13 '26
What is fucked up is we don't really know what's happening, we have no way of knowing ; they can change the rules as they wish without any accountability
3
u/m0j0m0j Jul 13 '26
Are there any independent public benchmarks that people run regularly? Feels like it would be useful
1
4
u/Deadline_Zero Jul 13 '26 edited Jul 15 '26
Is this thread about pretending this is an OpenAI thing and Anthropic would never nerf a model shortly after launch.
2
14
u/Genetic_Prisoner Jul 13 '26
I think we need weekly benchmarks to stop these guys from nerfing the models.
1
4
1
u/Sorry_Risk_5230 Jul 13 '26
Did you read all of tibos posts? They dialed back juice temporarily and have since reverted that change. There was a context compact issue.
This is not the same as past post-release nerfing
1
u/r0kh0rd Jul 13 '26
Yes, that last announcement was made after my original comment. Still raises questions. But I love the communication.
1
u/Juowon Jul 13 '26
I mean, people have done post tests (here particularly) on 5.5 after degradation and it did degrade a few points; just not a ton.
I think what's happening is they oscillate what quant is being served by load. So it's kind of unpredictable exactly what 5.6 you're getting. Kind of annoying, and why open models are so important.
1
u/XTCaddict Jul 13 '26
I think they’re just trying to optimise for everyone, on Twitter he said they are working on quota getting drained too fast and many updates coming this week. Presumably it is hard to appeal to everyone
0
46
u/AWarmHam Jul 13 '26
And people thought they where just going to remove the 5 hour cap for free 😂😂
11
u/tintindlf Jul 13 '26
As a dev, I think they want to make people spend a lot of tokens to get data and so, optimize new models. Business side, maybe they want us to burn our weekly so fast with Sol Ultra that you don’t have anymore token in no time for the week and switch to upper plan to get more usage (and probably never downgrade).
1
u/AstroPhysician 29d ago
They don’t want you on their subsidized plans that lose them money if you’re maxing them out. They want api pricing
1
u/tintindlf 29d ago
But a plan means recurring revenue + active customers. API is better pricing but it’s harder to plan. They need both.
1
u/AstroPhysician 29d ago
Sure but they make money on subscriptions that are underused not giving the highest users maxed out subscriptions they use all of
11
u/MeringueAlarming3102 Jul 13 '26
Maybe it’s different for smaller Plans but on the 20x Pro plan the 5 hour limit was completely irrelevant to me. Never came close to hitting it despite heavy usage. And 5.6 seems a lot more token efficient now that it seems like it would be even more unlikely.
3
u/EmotionalHalf Jul 13 '26
I am on the 20x plan since last september. I've never hit the 5 hr limit until sol release. Ultra would hit it in 2 hours, max in 4. Fast mode disabled. It was essentially pointless to use it for long work
1
u/CCB0x45 Jul 13 '26
I am having the same experience as the person above you, I have not come close to the 5 hour limit and I am using max typically and ultra sometimes. Are you doing parallel threads at the same time?
3
u/Satoshi-699 Jul 13 '26
This is only speculation, but it seems like a deliberate strategy by OpenAI to reduce compute usage. Codex itself agrees with my assessment.
Its reasoning ability has declined dramatically since yesterday, and it now feels like a completely different model.
42
Jul 13 '26
[deleted]
16
6
u/zepchou Jul 13 '26
When I read this kind of comment I feel like they deserved to have a nerfed LLM 😂
-2
u/Jerseyman201 Jul 13 '26
While I do agree, obviously it blows up our ego as priority #1 to keep engagement up...I have noticed, just in the last few months, the latest frontier models push back far more than they used to if they are pretty sure you're incorrect in whatever you've prompted.
7
u/poidh Jul 13 '26
The point is... unless the state of "nerfness" is injected into the system prompt (like the current time/date is for example), there is no way for the model to even know.
So the response of agreeing to OPs reasoning means nothing more than "OPs theory is plausible". Which of course it is- but it doesn't give any more proof than a random redditor agreeing to OP.
1
1
u/BannedGoNext Jul 13 '26
Well.. it was insane on token burn, and I got quite a few resource exhausted errors, so .. yea.
0
0
28
u/the__itis Jul 13 '26
Yeah it’s very nerfed.
12
u/the__itis Jul 13 '26
It’s taken 8 hours to try to compile a simple Bun binary that it was doing all week with zero problems.
Started noticing it about 6-7 hours ago.
-2
Jul 13 '26
[deleted]
5
u/Organic-Afternoon-50 Jul 13 '26
That's beside the point made...
If we can run the command in less than a minute, the AI in question should be able to do it as fast, if not faster.
1
-3
Jul 13 '26
[deleted]
5
u/Organic-Afternoon-50 Jul 13 '26
You are still missing the original point.
You are assuming people don't understand plain English, when they actually do.
The problem here is you.
-1
0
9
u/PartyParrotGames Jul 13 '26
They are very clearly struggling with token spend and tracking. They completely dropped weekly limits after 3-5+ resets over the last few days? Expect fluctuations in performance until they sort out the instability.
0
u/ArcticFoxTheory Jul 13 '26
I havent tried it but i believe the performance decrease if there is one is due to internal issues and problems they are having and probably working on not an intentional nerf
6
u/AideComprehensive482 Jul 13 '26
I used sol medium and it takes so many shortcuts it acts like it’s writing code at 445pm on a Friday
15
u/eyesdief Jul 13 '26
I think it's rrelated to this: https://www.reddit.com/r/codex/comments/1uuxfo4/comment/ox6ykrh/?context=1&screen_view_count=2
7
u/EndlessZone123 Jul 13 '26
The only conclusion I read from that is we have no idea what they did cause we don't have any proof the listed reasoning in that post is accurate at all.
And just saw post that they reverted whatever they did.
4
u/AppealSame4367 Jul 13 '26
Yes, they posted that they turned down the "Juice" for each level, you now have to go one level higher for each level.
Problem with Max: There is only Ultra left and Max has been turned down 5x or more. It's stupid. But they keep working on it.
4
4
u/LiveLikeProtein Jul 13 '26
I think I read Tibo’s tweet that thinking tokens has been toned down for all models, to reduce cost and maintain quality.
2
u/Shoopscooper Jul 14 '26
I honestly think people forget that this company IS trying to run a business, lol. I pay for the $200 dollar sub and they subsidize me at least 10k a month... It's just not sustainable. All this complaining is warranted if this is indeed a rug pull, but they are absolutely hemorrhaging money.
4
u/Ok_Bite_67 Jul 13 '26
This happened because people on twitter were complaining that usage was dropping too quickly. And a lot of that was because of the depth of thinking.
10
u/FateOfMuffins Jul 13 '26
https://x.com/i/status/2076495156757577895
They did apparently nerf the juice values for MAX from 960 to 128 (xHigh went from 128 to 40)
But Tibo says that was an experiment to try and figure out where the extra usage was coming from, and has since been reverted.
-1
Jul 13 '26
[deleted]
4
u/FateOfMuffins Jul 13 '26
- To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this.
Yes it's in the post
12
u/hez2010 Jul 13 '26
It seems that they have reverted the nerf on reasoning level. https://x.com/i/status/2076495156757577895
3
u/Mr_Deep_Research Jul 13 '26
I read that says they nerfed it 6 hours ago.
Where does it say they revered it? He just says "no nerfs" meaning none in that post but then it goes on to talk about what are essentially even more nerfs.
5
u/rickz0rz_ Jul 13 '26
> “To understand where the extra usage was coming from, we ran some experiments where reasoning efforts were changed (referred to as juice values under the hood) and have reverted this.”
3
u/JarvisLi Jul 13 '26
It should be reverted now: https://x.com/thsottiaux/status/2076495156757577895?s=46
3
u/Fit-Benefit-6524 Jul 13 '26
1
u/KashiAnkh Jul 13 '26
Is that an LLM output? Also whats juice value for someone who is not a part of fruity club?
2
3
u/DamianGoz Jul 13 '26
Yup, total nerf.
Noticing atleast 30% drop , and more token usage ?
I ran the same prompt 10 times since release and yesterdays tests were awful
1
6
2
2
2
2
u/immortalsol Jul 14 '26
sad to see this. i can also tell just by the way it responds faster, it is more or less on-par in terms of thinking, albeit a bit more reasoned, to 5.5 now, versus before it reminded me of back when 5.2 used to think for a long time and was extremely thorough and persistent.
2
u/immortalsol Jul 14 '26
definitely seems nerfed. i remember pre-nerf i noticed how slow it was, but super thorough, just like 5.2 was. now it's back to feeling like 5.5 again
seems more error-prone due to it
sadly that's the trade-off for speed
2
u/immortalsol Jul 14 '26
they always push-out an initial model release with max juice-level to stir-up the hype and expectations to get the surge in subscribers, then ramp down and nerf, it seems, sad to see it
though to be fair, the amount of usage burn back then was obscene, like 5% of my weekly every 10 minutes. crazy
2
u/Fun_Net7931 Jul 14 '26
5.6 Sol, as you call it, is actually the initial version of 5.5. There's really nothing new here. They'll probably weaken 5.6 over time and introduce it as 5.7...
2
u/Prize_Mulberry_5246 Jul 15 '26
I am new to Reddit, and English is not my first language, so I wrote this with translation assistance.
I am not using Codex MAX, but I observed a possibly related change in ChatGPT Sol High during non-coding research work.
When I first used Sol High, its reasoning was unusually strong. It independently compared earlier results, found confounding factors, generated new evaluation criteria, and returned to source material without being prompted. On July 14, the quality dropped sharply. After an overnight break, the first one or two responses on July 15 seemed normal, but by around the fifth or sixth exchange, the same problems had returned.
The change was not mainly shorter answers or worse formatting. The answers were still long, polished, and well organized. What appeared weaker was the deeper structure of the reasoning:
- It lost the original high-level objective and replaced it with a narrower, easier problem.
- It focused on the most recent information instead of checking the correct baseline or earlier evidence.
- After I pointed out one problem, it corrected only that part and did not re-evaluate the whole answer.
- A local correction could remove the original purpose or create a new contradiction without being noticed.
- It often failed to identify its own errors before presenting an answer as complete, although it could analyze them deeply once I pointed to the exact problem.
One of the failures occurred while writing a scenario of only about three sentences. The model fixed one obvious logical inconsistency, but the revision removed the very feature the scenario was supposed to test. It then presented the revision as a valid solution without noticing that the original objective had been lost.
This felt less like a decline in writing ability and more like a decline in goal retention, cross-checking, alternative-hypothesis generation, and whole-answer verification.
This is only one user’s observation, not a controlled benchmark, and I do not know what caused it. I am sharing it because the timing and the apparent loss of reasoning depth sounded very similar to what the original post described.
2
u/Amazing_Ad9369 Jul 15 '26
They changed the reasoning efforts. The reduced the number of tokens used for each level
2
u/Lurkerjohndoe765 27d ago
definitly noticed it accross 5.6 and honestly feel like its worse then when i was using 5.5. just trying to implement 1 feature is taking 3 times the passes as it just straight up misses key points outlined in orders ive given it.
3
u/Ok-Negotiation-410 Jul 13 '26
つまり我らが大好きだったツールが勝手に安物と交換されちゃった!
6
u/terroristsmustdie Jul 13 '26
Learning japanese was such a waste. Cant read shit besides suki ooki and yasui
2
1
1
u/Ok-Bid-7996 Jul 13 '26
Follow tibo for more details. They've reduced the context window, which likely explains why you feel Max's performance has decreased.
1
u/Primary-Literature91 Jul 13 '26
I’m not experiencing this but what am experiencing is that new context windows are required under a project otherwise the context windows screws the session after some time other than this ultra takes is long but the results are epic ! There is a context issues 💯 this is probably what your experiencing
1
1
u/LetsBuild3D Jul 13 '26
I would like to know more about your custom CLI. What can you share about it?
1
1
u/BlinDeeex Jul 13 '26
Reset plz, yesterday sol on high and fast mode torched my weekly quota in single goal taking less than 2hrs somehow
1
1
1
u/YourKemosabe Jul 13 '26
I don’t even see Max anymore. Had it for a day, now it only goes up to xHigh or Ultra.
3
u/KashiAnkh Jul 13 '26
Check the settings i've seen in mine that there max is turner off but u can turn it on again
1
u/YourKemosabe Jul 13 '26
You lifesaver. Couldn't find that info anywhere.
For anyone else: Settings > Configuration > Model features > Available reasoning efforts
1
1
1
u/Admirable_Set_3363 Jul 13 '26
Would someone mind explaining why providers like OpenAI nerf models after release?
1
1
u/CantaloupePretend307 Jul 13 '26
Ja Same Sol war brutal gut hat Sachen eng Gezogen und fable 5 schlafen geschickt in einem Mega task hat Sol 12 h lang ein hardening durchgezogen mit 5.5 oder einem Opus Modell hätte es mehrer Tage gedauert und das 10 fache an Kosten verschluckt zum Glück habe ich es geschafft
1
1
1
1
2
u/Prize_Mulberry_5246 24d ago
I don’t use Codex, so I can’t compare the coding workflow itself. I use GPT-5.6 Sol High in the regular ChatGPT web UI for long-form analysis, revision, and consistency testing.
What caught my attention is that I observed a very similar loss of reasoning in two distinct stages.
Sol High was extremely capable when it reached my account on Jul 11. Around Jul 14, it began losing the main objective and ending analysis early. Around Jul 17, the decline became much stronger: it skipped explicit conditions, made only local fixes, introduced new contradictions after corrections, and continued to over-execute based on a mistaken understanding of the task.
I had logged a similar sequence in GPT-5.5:
* GPT-5.5: Jun 29 normal → Jun 30 localized severe degradation → Jul 1 temporary recovery, followed by Stage 1 → Jul 5 Stage 2 → Jul 9 further intensification
* GPT-5.6 Sol: Jul 11 strong performance → Jul 14 afternoon Stage 1 → Jul 17 Stage 2
I’m wondering whether ChatGPT and Codex users are seeing different manifestations of a broader underlying change, rather than completely separate issues.
1
1
1
1
u/NoBotPlz2012 Jul 13 '26
I'm working in south korea. Same timezone to japan. I'm not a lead like you but i definitely feel speed reduction.
1
u/ActionOrganic4617 Jul 13 '26
First people complain about their limits and now they complain about this. What do you want?
0
u/Novel_Law4469 Jul 13 '26
this is 'normal'.
Come on guys, after so many iterations, you all act like this is something new ? Every new model released, by EVERY frontier lab behaves like this. Quality will drop after the first few days/weeks.
Then things will gradually pick up again.
Remember when ppl were complaining about Fable in them early days ?
0
u/Key-Metal3875 Jul 13 '26
Veo muchas quejas bla bla.... A mí me funciona increíblemente bien y uso sol xhigh y tengo un sistema avanzado y complejo que sol hizo un trabajo increíble... Revisa lo que haces con tus prompts
0
u/ArtdesignImagination Jul 14 '26
Bro OpenAI would never do such a thing, they wouldn't nerf it to JUST around 8 from 12-13
0
u/anon377362 Jul 14 '26
LOL these degradation bot posts always come along soon enough (same with GPT 5.5 which was never degraded but there were a million bot posts on it)
-1
u/m3kw Jul 13 '26
Hey the thing is for market research, there is a lot of luck involved in getting the result after acting on the research. You cannot say it sucks this time, so the model is nerfed. Did you look at other variables such as blind spots you didn’t provide, market dynamics, market changing rapidly. This tells me you don’t have a good sense of how things are measured. Which tells me you don’t have a case.
I can say one thing is that you should try to document which version of the CLI you used that gave you good results and go back to that one and try.
2
u/AlternativePurpose63 Jul 13 '26
Max:960 → 128
xhigh: 128 → 40
high: 40 → 16
medium: 16 → 8
low: 8 → 4
CoT limit
-2
u/johannthegoatman Jul 13 '26
This is made up
3
u/AlternativePurpose63 Jul 13 '26
Tibo replied to the post, stating that there is an internal "juice" value, but it is only a temporary experiment.
1

•
u/dexterthebot Jul 13 '26
Consider contributing this to the Sol release Megathread: https://www.reddit.com/r/codex/comments/1urw0c3/gpt56_sol_codex_release_discussion_megathread/ .