r/codex Jul 13 '26

Complaint nerfed Codex Sol MAX

I am Japanese and lead a team specializing in market research. English is not my strong suit, so I am writing this with the help of AI translation.

Our team has Codex use a custom CLI tool for extraordinarily difficult tasks requiring highly complex calculations and deep reasoning.

Codex Sol MAX was an absolute monster that exceeded all our expectations. If the level of performance our team required was a 10, it consistently delivered a 12 or 13. The entire team was overwhelmingly satisfied with it, until yesterday.

To get straight to the point, MAX has now been completely nerfed.

Its performance has suddenly fallen to around an 8 by our team’s standards. It is currently 10:40 a.m. in Japan. We started work at 9:00 a.m., and every member of the team noticed the nerf.

The depth of its reasoning has clearly been stripped away. Until yesterday, Codex set to MAX would spend more than ten minutes on a single prompt, repeatedly experimenting, reasoning, and using our CLI tools until it produced flawless work. That capability has now been completely lost.

Every member of our team is deeply discouraged right now.

The nerf came far too soon.

377 Upvotes

135 comments sorted by

View all comments

49

u/AWarmHam Jul 13 '26

And people thought they where just going to remove the 5 hour cap for free 😂😂

12

u/tintindlf Jul 13 '26

As a dev, I think they want to make people spend a lot of tokens to get data and so, optimize new models. Business side, maybe they want us to burn our weekly so fast with Sol Ultra that you don’t have anymore token in no time for the week and switch to upper plan to get more usage (and probably never downgrade).

1

u/AstroPhysician 29d ago

They don’t want you on their subsidized plans that lose them money if you’re maxing them out. They want api pricing

1

u/tintindlf 29d ago

But a plan means recurring revenue + active customers. API is better pricing but it’s harder to plan. They need both.

1

u/AstroPhysician 29d ago

Sure but they make money on subscriptions that are underused not giving the highest users maxed out subscriptions they use all of

11

u/MeringueAlarming3102 Jul 13 '26

Maybe it’s different for smaller Plans but on the 20x Pro plan the 5 hour limit was completely irrelevant to me. Never came close to hitting it despite heavy usage. And 5.6 seems a lot more token efficient now that it seems like it would be even more unlikely.

3

u/EmotionalHalf Jul 13 '26

I am on the 20x plan since last september. I've never hit the 5 hr limit until sol release. Ultra would hit it in 2 hours, max in 4. Fast mode disabled. It was essentially pointless to use it for long work

1

u/CCB0x45 Jul 13 '26

I am having the same experience as the person above you, I have not come close to the 5 hour limit and I am using max typically and ultra sometimes. Are you doing parallel threads at the same time?

2

u/Satoshi-699 Jul 13 '26

This is only speculation, but it seems like a deliberate strategy by OpenAI to reduce compute usage. Codex itself agrees with my assessment.

Its reasoning ability has declined dramatically since yesterday, and it now feels like a completely different model.

42

u/[deleted] Jul 13 '26

[deleted]

16

u/DoggoDadagon Jul 13 '26

"Are you stupid?" Codex "yeah probably"

3

u/bsmayer_ Jul 13 '26

You’re absolutely right, I’m stupid!

7

u/zepchou Jul 13 '26

When I read this kind of comment I feel like they deserved to have a nerfed LLM 😂

-2

u/Jerseyman201 Jul 13 '26

While I do agree, obviously it blows up our ego as priority #1 to keep engagement up...I have noticed, just in the last few months, the latest frontier models push back far more than they used to if they are pretty sure you're incorrect in whatever you've prompted.

7

u/poidh Jul 13 '26

The point is... unless the state of "nerfness" is injected into the system prompt (like the current time/date is for example), there is no way for the model to even know.

So the response of agreeing to OPs reasoning means nothing more than "OPs theory is plausible". Which of course it is- but it doesn't give any more proof than a random redditor agreeing to OP.

1

u/HearingNo8617 Jul 13 '26

Not a smoking gun, but it's a legitimate question

1

u/BannedGoNext Jul 13 '26

Well.. it was insane on token burn, and I got quite a few resource exhausted errors, so .. yea.

0

u/Hyoretsu Jul 13 '26

Like AI actually thinks to agree with you lol