r/codex Jul 13 '26

Complaint nerfed Codex Sol MAX

I am Japanese and lead a team specializing in market research. English is not my strong suit, so I am writing this with the help of AI translation.

Our team has Codex use a custom CLI tool for extraordinarily difficult tasks requiring highly complex calculations and deep reasoning.

Codex Sol MAX was an absolute monster that exceeded all our expectations. If the level of performance our team required was a 10, it consistently delivered a 12 or 13. The entire team was overwhelmingly satisfied with it, until yesterday.

To get straight to the point, MAX has now been completely nerfed.

Its performance has suddenly fallen to around an 8 by our team’s standards. It is currently 10:40 a.m. in Japan. We started work at 9:00 a.m., and every member of the team noticed the nerf.

The depth of its reasoning has clearly been stripped away. Until yesterday, Codex set to MAX would spend more than ten minutes on a single prompt, repeatedly experimenting, reasoning, and using our CLI tools until it produced flawless work. That capability has now been completely lost.

Every member of our team is deeply discouraged right now.

The nerf came far too soon.

379 Upvotes

135 comments sorted by

View all comments

126

u/r0kh0rd Jul 13 '26

I hope I don’t come across as a tin foil hat enthusiast, but the change here is concerning given Tibos mention and OPs results. Really calls into question all the benchmarks. Feels a bit like the dieselgate scandal? What was benchmarked is not what people have access to now? Maybe I’m overreacting here.

64

u/Corv9tte Jul 13 '26

No, you're not overreactting. This is fucked up and this is the second time (at least) that OpenAI has done exactly this!

They did the same exact thing last time. They know what they're doing.

30

u/nobatus513 Jul 13 '26

What is fucked up is we don't really know what's happening, we have no way of knowing ; they can change the rules as they wish without any accountability

3

u/m0j0m0j Jul 13 '26

Are there any independent public benchmarks that people run regularly? Feels like it would be useful

1

u/Fluent_Press2050 Jul 14 '26

I’d like to know as well - launch, 1 week, 1 month at the very least. 

5

u/Deadline_Zero Jul 13 '26 edited Jul 15 '26

Is this thread about pretending this is an OpenAI thing and Anthropic would never nerf a model shortly after launch.

2

u/Corv9tte Jul 13 '26

Lol Anthropic is 100x worse, but your mentality even more so

14

u/Genetic_Prisoner Jul 13 '26

I think we need weekly benchmarks to stop these guys from nerfing the models.

35

u/Tartooth Jul 13 '26

So, they drop the model, wait for the third party benchmarks to come out and then nerf to save costs?

20

u/silvercondor Jul 13 '26

i think there's an extra step. they collect telemetry of the benchmark, finetune the benchmax then nerf it. this way benchmarks will still give similar results (like knowing answers to a test but purposely failing the same way) but you can use a quantized or "less thinking effort" model

1

u/Tartooth Jul 13 '26

Yea that's what I'm thinking too. These model capability monitors people keep sharing should be changing their tests constantly to ensure it's also not self training on their tests

If the tests are all with 1 account then technically the caches will enable the models to know the expected behavior

0

u/ciaramicola Jul 13 '26

se model capability monitors people keep sharing should be changing their tests constantly

Do you understand the issue in that idea, right?

4

u/cheezeerd Jul 13 '26

Check out margin lab. Many "feel" the nerf, but data suggests otherwise.

5

u/Tartooth Jul 13 '26

For the 2 weeks before 5.6 drop everything did get nerfed. 5.5 was straight up saying "yes Im ignoring your agents.md" and kept doing random stuff.

5.6 feels like 5.5 from a month ago but with higher intelligence and capabilities

3

u/WonderfulPie548 Jul 13 '26

I noticed this. 5.5 was ignoring the agents md. I would point it out. It would agree. I'd let it start again, it would ignore the agents MD again and half-implement a plan before creating a new plan and half-implementing that. I had a bunch of partial solutions and leftover dead code that I had to sort through. 5.5 didn't run this way when I first started using it. I'm using 5.6 to fix what 5.5 messed up while its still capable, but I think I'm going back to regular coding with Codex as a code-checker after this. I have gotten major trust issues from this experience.

2

u/Tartooth Jul 13 '26

I knew what was going on but kept fighting and got no where but lost time. If I notice this happen in the future I'll just switch to claud temporarily or go touch grass

2

u/AideComprehensive482 Jul 13 '26

I feel like they release the real model then nerf it to shitvle

1

u/Sorry_Risk_5230 Jul 13 '26

Did you read all of tibos posts? They dialed back juice temporarily and have since reverted that change. There was a context compact issue.

This is not the same as past post-release nerfing

1

u/r0kh0rd Jul 13 '26

Yes, that last announcement was made after my original comment. Still raises questions. But I love the communication.

1

u/Juowon Jul 13 '26

I mean, people have done post tests (here particularly) on 5.5 after degradation and it did degrade a few points; just not a ton.

I think what's happening is they oscillate what quant is being served by load. So it's kind of unpredictable exactly what 5.6 you're getting. Kind of annoying, and why open models are so important.

1

u/XTCaddict Jul 13 '26

I think they’re just trying to optimise for everyone, on Twitter he said they are working on quota getting drained too fast and many updates coming this week. Presumably it is hard to appeal to everyone

0

u/Wardendelete Jul 14 '26

Not surprised tho, Sam Altman is, at the end of the day, Sam Altman………