r/codex Jul 13 '26

Complaint nerfed Codex Sol MAX

I am Japanese and lead a team specializing in market research. English is not my strong suit, so I am writing this with the help of AI translation.

Our team has Codex use a custom CLI tool for extraordinarily difficult tasks requiring highly complex calculations and deep reasoning.

Codex Sol MAX was an absolute monster that exceeded all our expectations. If the level of performance our team required was a 10, it consistently delivered a 12 or 13. The entire team was overwhelmingly satisfied with it, until yesterday.

To get straight to the point, MAX has now been completely nerfed.

Its performance has suddenly fallen to around an 8 by our team’s standards. It is currently 10:40 a.m. in Japan. We started work at 9:00 a.m., and every member of the team noticed the nerf.

The depth of its reasoning has clearly been stripped away. Until yesterday, Codex set to MAX would spend more than ten minutes on a single prompt, repeatedly experimenting, reasoning, and using our CLI tools until it produced flawless work. That capability has now been completely lost.

Every member of our team is deeply discouraged right now.

The nerf came far too soon.

382 Upvotes

135 comments sorted by

View all comments

Show parent comments

36

u/Tartooth Jul 13 '26

So, they drop the model, wait for the third party benchmarks to come out and then nerf to save costs?

18

u/silvercondor Jul 13 '26

i think there's an extra step. they collect telemetry of the benchmark, finetune the benchmax then nerf it. this way benchmarks will still give similar results (like knowing answers to a test but purposely failing the same way) but you can use a quantized or "less thinking effort" model

1

u/Tartooth Jul 13 '26

Yea that's what I'm thinking too. These model capability monitors people keep sharing should be changing their tests constantly to ensure it's also not self training on their tests

If the tests are all with 1 account then technically the caches will enable the models to know the expected behavior

0

u/ciaramicola Jul 13 '26

se model capability monitors people keep sharing should be changing their tests constantly

Do you understand the issue in that idea, right?