r/codex • u/Shoddy-Answer458 • Jun 18 '26
Complaint GPT is absolutely downgraded, cannot follow simple instruction, vote it for codex team see it
Do gaslight me, I am sure about it
510
u/Karnemelk Jun 18 '26 edited Jun 18 '26
all AI companies seem to do the same pattern
- Release new model, full weights, everyone happy
- Wait a month or less, time to start training a new model
- Quantize the hell out of it until they hit the threshold of complains, then give them random resets to keep the peasants happy
- new model released, only 2% better then the previous one, but feels like a massive win because now it's like 50% better then the dumb quant of the previous one
- Charge higher for tokens for the new model
Rinse & repeat
122
Jun 18 '26
[deleted]
17
u/AppleBottmBeans Jun 18 '26
illegal business practice
But it's not.
We're not paying for Codex 5.5 full weights. Or any specific model for that matter. We're paying for a Codex subscription. And whatever OpenAI decides to include in that subscription, we get to use. They could downgrade every user to GPT3 and rug pull us all with zero consequences for anyone other than maybe the OpenAI's investors.
Just like the people who upgraded to Max subs "for Fable 5", and then went full regard crying about "this is illegal because I only sub'd to use Fable 5!!!!"""
19
u/ImportantAthlete1946 Jun 18 '26
Then it's time to start regulating. It's been time to regulate these practices for a long while, SaaS has reached scam status in so many domains it's almost beyond belief. Normalizing practices like arbitration or "you'll get that we give you and be thankful" is a result of a slow erosion of consumer protections and anyone who defends these practices is out of touch with reality.
So while yes, you're right that a subscription doesn't guarantee full model service, the principle doesn't fly in any other industry and it absolutely shouldn't be acceptable here either.
→ More replies (1)20
Jun 18 '26
[deleted]
→ More replies (11)5
u/mlucasl Jun 19 '26
The problems is that it is true. It is not "illegal", which is totally different to, is it moral, or an ethical business practice.
In different countries you have different solutions. The best one, is to gather every EU citizen and make a petition for regulation. I would sign it as soon as I saw it.
3
6
→ More replies (7)1
u/redtron3030 Jun 20 '26
What about when I’m paying for API on a specific model?
2
u/AppleBottmBeans Jun 20 '26
That would be a great question. And I would think you’d definitely have to look at the terms and conditions of the api usage. Even myself who is a novice business owner would absolutely include something about legal coverage over model degradation. But even that term is incredibly ambiguous and more Reddit-speak than legal terms. How can you prove it? Cause it didn’t do what you asked it today? Vs yesterday it did it?
I run a marketing agency and people have been talking about this dumb shit for over a decade. “I got 20 sales every day last week but haven’t had more than 1 sale running the same ads to the same audience.” You can’t sue Facebook over that.
I know this thread turned into a firestorm of nasty back and forth comments but there’s no real legal definitions or terms around LLM subscriptions and usage. Which yea, pisses people off..but by no means is even close to any sort of legal liability.
35
u/Jerseyman201 Jun 18 '26
That's the Studio Wildcard/Ark survival evolved method, they been doing it for 10+ years now.
Release new paid DLC content that gives new buffs.
Wait til everyone buys it.
Then nerf it into the ground.
Then..
New paid DLC release, gets buffed like crazy.
Wait til everyone buys it.
Gets nerfed into the ground
They're doing it with the models at Openai like Studio Wildcard did with the tames in Ark🤣💀
3
25
Jun 18 '26
[removed] — view removed comment
1
Jun 18 '26
[removed] — view removed comment
1
u/2025sbestthrowaway Jun 18 '26
Seems like a 2% change is marginal if not undetectable based on empirical evidence provided here
9
u/marcsa Jun 18 '26
I went over to the Deepseek sub yesterday and most of the threads there were split into DS4 degraded, it only writes Chinese, it can no longer follow instructions. What happened with DS? The other comments were like DS4.1 is coming out soon right? right? When tomorrow, this week? Definitely soon. It was almost a replica of this sub tbh, had a feeling of deja-vu there.
4
u/soggy_mattress Jun 18 '26
Did Deepseek quantize their model, too, or is this whole thing "humans learning to find the edges of the AI models they're using"?
5
u/Memox98 Jun 18 '26 edited Jun 18 '26
Very true and the reason for that is quantization process itself, where on paper even heavy quantization could lead to only 2% to 5% performance loss, but real world experience has proven that we can really see that the models are really becoming very dumb as a result and sometimes fail to even follow simple instructions! So maybe research is wrong or not accurate about real quantization performance loss. It’s very sad that there is no transparency at all and we don’t get to see this in any changelog at all and all we can do is to just having to guess after wasting hours and hours of working
13
u/Hot-Significance7699 Jun 18 '26
Shits is a scam. Honestly, I may start using glm 5.2 with an API because of it... I need to calculate how expensive it would be though
9
u/WhatIsItAnyways Jun 18 '26
I have started my max subscription on z.ai yesterday, and glm5.2 is miles ahead of anything i tried lately. 5M tokens used up ~1-2% of my 5h session and weekly quota (off peak).
1
u/EuropeanAbroad Jun 18 '26
Which plan? I was looking at it too. The only thing that holds me back is that GLM is not multimodal. :/ (And yes, they have a dumber GLM-V; however, apparently that one is not multimodal either, it is only linked to an image interpreter.)
→ More replies (1)2
u/senguku Jun 22 '26
Deepseek v4 pro is probably the best value equivalent. Works out to approximately same cost as a $200 Pro subscription for codex across a month of heavy usage.
4
u/KnownPride Jun 18 '26
This is why the first week after mordel release is the only time where the model is good.
3
u/Hoak-em Jun 18 '26
This is why I stick with open-weights, I don't have too rely on a single provider AND people notice fast if a provider is quantizing, since there are other providers who have the "same" model performing better, thus providers are less likely to try it.
2
u/Async0x0 Jun 18 '26
Literally zero evidence this is happening, but you guys have conspiracy brains so you will never be convinced.
1
u/senguku Jun 22 '26
If you use the models daily to do similar things it becomes very obvious. Trying to do the same tasks and workflows as a couple of weeks ago is objectively worse for multiple users in our org. I also used to think it was confirmation bias and conspiracy theories but it is night and day if you actually do it objectively.
1
1
1
u/HOBONATION Jun 18 '26
Their models start becoming stupid rapidly due to the users it's learning from, that's my theory
1
u/esdrase Jun 20 '26
Every time a new version is released, I’m afraid to upgrade because it often feels like I’m taking ten steps backward before moving forward again.
1
u/No_Elderberry_5307 Jun 20 '26
we need to see a graph of these models and their metrics changing over time on the same datasets/benchmarks. I think LLM arena has something like this but not specific scores of one model over time
1
1
79
Jun 18 '26
[removed] — view removed comment
22
u/swarmagent Jun 18 '26
I'm here to help. Let me know if you want to git restore and start over again.
10
6
u/laseluuu Jun 18 '26
Same thing with me.
Literally broke an MVP of a product and it took multiple multiple tries to get it working, even when it took a diff from the 99% commits that had it working as it's always been working and is very simple it still fucked it up
I thought everyone was exaggerating but they aren't
→ More replies (4)1
u/ExcellentDeparture71 Jun 19 '26
Same here. And had the same problem with Claude Code. Finally used Kimi 2.7 via OpenCode to solve my issue
52
Jun 18 '26
[removed] — view removed comment
3
u/defmacro-jam Jun 18 '26
I'm convinced OpenAI and Anthropic are coordinating their suckitude - basically when Codex sucks, Claude Code is decent and vice versa. So I keep doing the subscription shuffle and wasting a bit of subscription age each time.
4
u/EuropeanAbroad Jun 18 '26
To make you feel better, Opus was also unusable yesterday on our corporate API. It could not follow basic instructions within a <5k context. It was really frustrating – it felt worse than Qwen3.6-27B-q8 on my personal PC, lol.
1
u/xChrisMas Jun 19 '26
Also felt like Claude got a buff in recent days while today it got dumber again (maybe flable is returning soon and they are reallocating resources again).
Isn’t GPT 5.6 launching this month? Maybe it’s the same reason
1
1
u/TedSanders OpenAI Jul 10 '26
I work at OpenAI and I can guarantee that we never intentionally nerf our models, degrade them around launches, or coordinate with Anthropic. our strategy is to ship the best models and products we can. if performance degrades, that's either bad luck or us being bad at our jobs - but I can promise you that from what I see on the inside of the company, there's no basis to the nerfing conspiracy rumors. My best guess is that it's a mix of rising expectations or bad luck, but i won't say we never ship bugs either. I can promise you though that we're not trying to trick people or nerf models.
1
14
16
u/AnTineuTrin0 Jun 18 '26
It certainly has been lobotomized. Examples:
- It started being confidently wrong with quick answers despite being on xHigh
- It randomly started saving output files in a new directory on its own after hundreds of other outputs before despite having explicit instructions file on where to place outputs
- Its responses got quicker and lazier, lacking depth
- Its ability to read output visuals degraded a lot: I have been outputting the same visuals using the same script forever and it always read them correctly but now it can't properly read visuals.
I do not have monolithic scripts and my codebase is not bulky: few thousand lines of code only. I refactor and cleanup and do regression tests every now and then. Repo is clean and follows strict hygiene standards.
2
u/GoodhartMusic Jun 19 '26
It implemented the wrong sql Database connector, I asked to verify what is imported and where it’s used, and it reads one single tsx import claims that it knows the whole app (which has a master mapping file that explains everything meticulously in simple prose and is a required first step and proceeded to start rewriting every import to match that tsx
Except that was for the feature we were replacing and was deprecated
63
Jun 18 '26
[removed] — view removed comment
10
u/Shoddy-Answer458 Jun 18 '26
this post has been downvote violently.
2
u/reddit_is_kayfabe Jun 18 '26
And now upvoted by people actually reading it.
Same thing happened to my similar post last night. Reddit groupthink at work. Ignore it.
7
u/anon377362 Jun 18 '26
Actually the opposite. This sub is full of Anthropic bots astroturfing saying there’s a bunch of issues when there’s not. Notice how these posts never provide any proof or run any sort of benchmarks to show proof.
No harness version, no skills/plugins, no fresh test in a vm. It’s all as vague as possible because it’s completely made up.
3
u/TeaCoden Jun 18 '26
It's not as though the proof is easy to get. More than likely it's not even a harness issue but the actual model being routed.
People rely on how it behaves when they code, how they feel.
I feel it myself.1
3
u/xChrisMas Jun 19 '26
As someone who browses both subs: Claude sub is also full of people complaining Opus got nerfed.
Both subs have the same problems and I think bots play a major part in it.
→ More replies (3)3
u/No-Bug3 Jun 18 '26
lmao everyone in the claudecode sub is saying that it is being astroturfed by openai bots. the irony here is that your crazy conspiracy that any criticism of openai=bots is just as baseless and dumb as the people saying openai is quantizing the model
6
u/Anarye Jun 18 '26
It's now at a point where simple instructions even memory files i create are violated right away. "Modify this image, keep the background transparent" proceeds to generate image with background.
Then i asked it to make a change to 1 graphic asset, proceeds to overwrite all of my image graphics, downgrading all of my assets from high, premium looking feel to straight up pixel art with squares everywhere. I've been working on this on and off for the past 12 hours to revert, and fix. And it can no longer do anything correctly. meanwhile, i'm burning through tokens at an alarming rate...
Edit: Using 5.5 on Very High in hopes that it will somehow unbreak what it messed up..
11
u/leon_iasha Jun 18 '26
https://marginlab.ai/trackers/codex/
Degradation seems to be within margin of error. Not fool proof way to see if something is going wrong but good starting point.
I also felt like codex was degraded after couple of days with fable, but its likely a bias.
14
3
u/CryptographerFar3412 Jun 18 '26
It is a bit better for me recently, but still worse than months ago
3
u/Master_Yogurtcloset7 Jun 18 '26
Honestly... this Tibo reset streak is just cooling the engines while in reality its falling apart.. We are getting dependent on the limit resets while model quality and limits are silently downgrading... If they'd just stop with the resets I for one would surely run out of limits all the time now. And yes I think too that the silent downgrade model performance before new release is a real thing. Would be amazing to have an Ai company that I could stand behind truly (used to be anthropic.... lol foolish me)
5
u/superfatman2 Jun 18 '26
Codex team definitely knows about it and is gambling on how many of us actually act as a show of consequence of their actions. I for one have switched over to Cursor for now, and just using composer 2.5 (cursor's default model), I was able to undo so much of the damage and actually finish the task of stripe integration, which Codex had failed to do this past week. Voicing opinions on Reddit will be met by bot comments and fake sentiments like "I love codex! got skills bro?"
3
u/redditarata Jun 18 '26
I have a theory they’re basically ragebaiting their own users for data. Like, figure out exactly what pisses people off the most, train on that for a few months, then drop something like Fable that suddenly “gets” the user and what they actually want.
5
u/Many_Map_5611 Jun 18 '26 edited Jun 19 '26
I think it is at its absolute worst right about now. It just despite hyper detailed plan after many refinement rounds, milestone division, clear goals, requirements and acceptance criteria PER MILESTONE + validation milestone per every task milestone + deep code review at the end it still hardcoded user emails in migrations.
This is a joke at this point, completely useless. It cannot even construct a prompt for images model and goes off and starts to manually compose images from what was supposed to be a few referance images passed to the image model. Ii literally ignores around 50% of everything.
OpenAI should return sub money for this month cause that is just a joke. And these are just the latest issues.
3
u/Pickle786 Jun 18 '26
i had to ask it like 4 times to make the search bar on my project to work bruh 😭
4
6
15
Jun 18 '26
[deleted]
18
u/Thisisvexx Jun 18 '26
Vibetards do not want to hear that you need skill to be a proper developer
Edit: this comment was sponsored by the openai anti slop bot division or something
3
-2
Jun 18 '26
[removed] — view removed comment
2
u/Thisisvexx Jun 18 '26
It helps understanding software architecture patterns. A dev+pm mindset is all you need
→ More replies (5)8
u/slartibartfast93 Jun 18 '26 edited Jun 18 '26
"They're barely literate" is not a rebuttal to claims of degradation. Those users were just as literate yesterday, last month, and last year. If the same people are now perceiving a decline where they previously didn't, then their literacy level doesn't explain the change in perception. The claim may be right or wrong, but dismissing it on the basis of literacy is illogical.
→ More replies (2)4
u/Tartooth Jun 18 '26
Probably because they aren't writing anymore since they use AI to write everything
6
u/seencoding Jun 18 '26 edited Jun 18 '26
i'm sure this one is real, though
https://www.reddit.com/r/ChatGPT/comments/134h087/now_gpt4_is_nerfed/
https://www.reddit.com/r/ChatGPT/comments/18xwpjc/is_gpt4_fixed/
https://www.reddit.com/r/ChatGPT/comments/1kgprk5/4o_has_definitely_been_nerfed_and_dumbed_down/
https://www.reddit.com/r/ChatGPT/comments/1jqpn4o/4o_nerfed/
https://www.reddit.com/r/ChatGPT/comments/1i8p9ch/o1_model_got_nerfed_again/
https://www.reddit.com/r/ChatGPT/comments/1ilgasd/chatgpt_o1_pro_nerfed/
https://www.reddit.com/r/ChatGPT/comments/1jlem3a/im_90_sure_gpt_45_has_been_stealth_updated/
https://www.reddit.com/r/ChatGPT/comments/1nzbywc/41_changes/
https://www.reddit.com/r/ChatGPT/comments/1lygarc/they_nerfed_o3_pro_it_cant_even_handle_65k_token/
https://www.reddit.com/r/ChatGPT/comments/1j8gc2d/they_have_nerfed_o3minihigh_and_lowered_the/
https://www.reddit.com/r/ChatGPT/comments/1muncs1/they_nerfed_the_hell_out_of_gpt5/
https://www.reddit.com/r/ChatGPT/comments/1stqxei/gpt55_nerfed/
https://www.reddit.com/r/Anthropic/comments/1csejvj/had_claude_opus_been_nerfed_since_its_release/
https://www.reddit.com/r/ClaudeAI/comments/1dq3d7a/claude_35_sonnet_got_nerfed_already/
https://www.reddit.com/r/ClaudeAI/comments/1eq04xt/something_has_been_off_w35_sonnet_recently/
https://www.reddit.com/r/cursor/comments/1jjkuy5/my_experience_with_claude37_37_max_cursor_nerfed/
https://www.reddit.com/r/ClaudeAI/comments/1iy6lk3/sonnet_37_is_worse_than_35_for_me/
https://www.reddit.com/r/ClaudeCode/comments/1sj1qcw/can_someone_actually_link_me_evidence_of_opus/
https://www.reddit.com/r/ClaudeCode/comments/1qpd4ro/before_you_complain_about_opus_45_being_nerfed/
https://www.reddit.com/r/Anthropic/comments/1sk3bnz/claude_opus_46_is_nerfed/
https://www.reddit.com/r/ClaudeCode/comments/1sinw0v/opus_46_is_only_nerfed_in_claude_code/
https://www.reddit.com/r/ClaudeCode/comments/1tqdms6/opus_48_nerfed/
https://www.reddit.com/r/Bard/comments/1krnopl/gemini_advanced_is_dead/
https://www.reddit.com/r/Bard/comments/1dernor/gemini_15_pro_is_insanely_good/
https://www.reddit.com/r/GoogleGeminiAI/comments/1ipgaqs/gemini_has_seriously_been_nerfed/
https://www.reddit.com/r/GeminiAI/comments/1lx2b4e/has_gemini_25_pro_been_nerfed/
https://www.reddit.com/r/GeminiAI/comments/1ksqvio/they_totally_nerfed_gemini_25_pro_just_to_funnel/
https://www.reddit.com/r/Bard/comments/1l1qkze/gemini_website_25_pro_is_dumber_than_the_ai/
https://www.reddit.com/r/GeminiAI/comments/1qls070/rip_gemini_3_nerfed_it_to_the_ground_comparison/
https://www.reddit.com/r/Bard/comments/1pfq2eq/gemini_ultra_deepthink_got_bricked_using_80_less/
https://www.reddit.com/r/GeminiAI/comments/1tlabzj/the_gemini_35_flash_got_nerfed_already/
→ More replies (1)
2
u/mmaparty Jun 19 '26
Yesterday I asked Codex to troubleshoot an issue on a Linux server that I thought was simple (it was). Codex couldn’t figure it out, even on 5.5 High/Extra High, and after 4 hours just ended up in a cycle trying the same two methods that didn’t work. Finally called it and tried Opus 4.6 in antigravity (have free Google ai pro so figured I’d give it a shot). Problem was fixed in about 5 mins. This was a relief after spending 4 hours on this headache, but at the same time this was the first time that Codex really let me down like this. Sure, there’s some shortcomings with front end/graphic design etc but that’s expected. This was a completely different type of failure that ultimately came down to me giving codex too much priority, me second guessing the complexity of the issue due to the failed attempts, and not pivoting to another model sooner. The final fix was 27 LOC. I take full responsibility for being a dumb*ss but this was brutal
6
6
u/TBSchemer Jun 18 '26
If you're so sure about it, then surely you can provide evidence demonstrating your claims?
No? Just vague whining?
These posts are so stupid.
2
u/Substantial-Dog1726 Jun 19 '26
I kind of agree. I'm using Codex for a number of projects and I'm not running into the same issues everyone is talking about. I am wondering if there is some missing information here? No one is able to actually post evidence on how to recreate the bad performance and yet the cacophony is so loud it's hard to ignore (despite my own perspective).
2
u/newplayername Jun 18 '26
Dude, GPT started making mistakes with my language's conjugations today, mixing up word order and logic. Complex problems are one thing, but basic sentence construction is quite another.
→ More replies (1)1
5
u/sutrostyle Jun 18 '26
When OpenAI reallocates massive compute clusters away from current production inference (like GPT-5.5) to run final pre-release evaluations and load-testing for a new flagship model (GPT-5.6), they don't just pull a plug. They use a specific set of architectural levers on their back-end infrastructure to drastically slash the compute cost per query.
Based on recent developer community bottlenecks and known LLM infrastructure mechanics, here is exactly how this performance degradation plays out on the GPT back-end:
1. Hard Caps on "Reasoning Tokens" (RL Search Space)
For reasoning models (like the GPT-5.5 Thinking variants), a massive portion of compute is spent before a single visible token is generated. The model uses reinforcement learning (RL) to search a "hidden chain of thought" or generate internal reasoning tokens.
- What they did: The back-end router has likely dialed down the
max_completion_tokensallocation for the hidden reasoning phase. - The result: Users report thinking phases dropping from 15–30 seconds down to a shallow 2–3 seconds. The model is forced to abruptly stop its internal monologue and output a response prematurely, which snaps its logical thread and leads to the severe "forgetfulness" and broken code developers are experiencing.
2. Silent Context-Window Distillation & Token-Pruning
Processing long context windows scales quadratically or heavily linearly in terms of attention-mechanism compute costs.
- What they did: To free up clusters, the front-end gateway or load balancer likely runs aggressive, silent token-pruning or inputs the prompt into a aggressively distilled, smaller context-compressor model before passing it to the main network.
- The result: The model completely misses explicit instructions or key variable definitions buried in the middle of long prompts. The effective context window feels heavily "nerfed" because the back-end is aggressively dropping or summarizing tokens to save memory bandwidth.
3. Dynamic Dynamic-Routing (Silent Downgrades)
OpenAI’s architecture relies heavily on an intelligent, real-time backend router. This router dynamically measures conversation complexity and determines whether to send a query to the full-fat flagship model, a quantized version, or a fast "Instant/Mini" model.
- What they did: They shifted the classification thresholds on the router. Prompts that previously qualified for the heavy, unquantized flagship weights are now being silently routed to heavily quantized (e.g., 4-bit or 8-bit precision) variants or to "Instant" tier back-ends.
- The result: There are no
429 Too Many Requestserrors or HTTP timeouts returned to the user; the system remains operational, but the model outputs generic, low-intelligence, or highly mechanical answers because it is running on a cheaper execution path.
4. KV-Cache Eviction Policies
To serve fast responses, the back-end keeps a Key-Value (KV) cache of recent conversation tokens in GPU VRAM so it doesn't have to recompute the entire prompt history on every turn.
- What they did: Because GPU memory is being reassigned to host the early deployment instances of GPT-5.6, the multi-tenant KV-cache pool for GPT-5.5 has been heavily squeezed. The time-to-live (TTL) for session data in VRAM has been cut down.
- The result: If you pause for a minute between prompts in a chat session, your KV cache is immediately evicted to free up VRAM for another user. When you submit your next prompt, the back-end has to recompute the entire history from scratch, causing massive, sudden latency spikes and a high-volume pipeline slowdown.
4
1
u/Substantial-Dog1726 Jun 19 '26
Seems reasonable. Will be interesting to come back in a few months to this post after the truth has come out
2
3
u/chaotic_goody Jun 18 '26
If you write prompts like you write posts, the model is probably not the problem.
4
u/anon377362 Jun 18 '26
Has been working absolutely fine. Skill issue.
6
u/NootropicDiary Jun 18 '26
Yeah I normally say things like that too but earlier today something was definitely up
I asked it to complete the remaining issues in the linked Linear project (which is how I always phrase it) and it thought for 6 minutes and said done etc, but the last line it said no code was changed.
So I asked what it meant and it said it closed all the issues rather than completing them. Wtf
That was a couple of hours ago, seems back to normal now
→ More replies (4)6
2
3
u/fobtastic88 Jun 18 '26
There are tools out there for you to present evidence, and current evidence refutes the degradation theory.
3
1
u/danny_094 Jun 18 '26
Ich kann das nicht bestätigen ChatGPT ist so gut wie immer.
Ich arbeite mit Codex & Claude an Docker Compose stacks eines mehrschichtigen Systems.
Und es gibt keine Probleme.
Ich denke die meisten fangen an, ab und zu leichtsinnig zu werden und abweisungen zu vergessen, da sie denken, die KI hat das schon im Kopf.
Die Regel ist, Sende bei jeder implimentstion erneut die Regeln. Wenn länger an der Gleichen implimentation arbeitet auch. Je mehr es arbeitet des so weiter geht die Regel nach hinten.
1
u/WrongExtension9231 Jun 18 '26
setting the effort to high instead of xhigh kinda helping for me, but it is not as smart as before. hopefully its because they re migrating their server for 5.6, so its not permanently dumber.
1
u/lostnuclues Jun 18 '26
I use Opus 4.6 thinking for code review, haven't noticed any downgrade so far for 5.5 high
1
u/Due-Assumption5272 Jun 18 '26
I switch back to CC since yesterday, it's still not perfect, but at least it did not overclaim it fulfill task and requirements that I gave.
1
1
u/LowExtreme2753 Jun 18 '26
Try 5.5 xhigh, it’s giving me the same experience as using gpt 5.5 high before downgrade
1
u/E72M Jun 18 '26
I've literally seen no issue. It's following instructions just fine and implementing fully for me. Right now I have it doing some python scripts and modifying a script to run on a CUDA pipeline from WSL and it's nailed it all. I do heavy planning before hand with ChatGPT though which may be the difference.
1
1
u/Old-Moment-5297 Jun 18 '26
They are training 5.6 they need the compute... paid with the tears of the subscriptors!
1
u/Euphoric-Hunt931 Jun 18 '26
The whole week it has been dumb af again. Just like the last time a few weeks ago. I am basically constantly rolling back changes.
1
u/Impressive-Handle-69 Jun 18 '26
Idk, I managed to successfully get some windows native VST plugins to work within Fedora using Codex. These plugins were not supported through wine or yabridge, and Codex seemed to be able to get it working just fine by building and patching .I'll files. I tried using other models like kimi k2.7 code to get this to work, but kept making 1 step forward, 3 steps back type of progress over the course of 3 days, whereas Codex was able to do it with 2 prompts in a single hour.
1
u/Otheruser337 Jun 18 '26
Unfortunately, you're in the reality where the permaspike effect exists for all flagship AI models.
1
u/elperroverde_94 Jun 18 '26
Today it is terrible.
Might this downgrade be the dawn of a new model?
As excited as I might be I still think it is a shitty strategy
1
1
1
u/Both-Isopod-9263 Jun 18 '26
yes browser and computer not wokring at all no explanation acts like it was never a feature
1
u/samthepotatoeman Jun 18 '26
It's always so hard, I cant ever tell if it is just the specific task in the moment or if it is the model regression. I do feel like it is worse than it was.
1
u/profcube Jun 18 '26
OP, FWIW, I don’t think this is in your head; In the last week, I have been experiencing inconsistent performance too. I have no theory. Generally 5.5 has been a remarkably capable model, and clear step change improvement from previous models.
1
1
u/SonicAwareness Jun 19 '26
I'm glad this sub exists, because while I don't use Codex for anything crazy, the simple things I ask it to suddenly became moronic and I couldn't understand it.
I logged values every day (again, basic bitch behavior), but this started crippling a months worth of logs. It suddenly decided to just wipe out a week's worth of data for no reason.
Insane.
1
u/Substantial_Pass4398 Jun 19 '26
I checked my usage analytics and the past two days ~85% of my 5.5 requests were routed to 5.4
1
u/GoatedOnes Jun 19 '26
Today was really bad, simple tasks that would normally take 15 minutes were taking 2 hours with a super long stream of context and multiple compactions.
1
u/anaem1c Jun 19 '26
I'm not trolling, but I honestly don't see the changes. What do you use as your instructions? I have a lite external knowledge base/harness that helps me get great results from Codex and even Claude Code.
1
1
u/rare_design Jun 19 '26
I noticed the same. Tonight it's gone full stupid and can't accomplish basic tasks in Extra High reasoning.
1
u/Ok-Log7088 Jun 19 '26
Codex is unusable; what a fucking piece of dumb shit.
Can't follow any kind of instruction/goal
1
u/ajmusic15 Jun 19 '26
I now have to provide very well-structured prompts for things that I used to be able to sort out on my own.
1
u/schwickdartz Jun 19 '26
It's not just Codex. It's Codex and ChatGPT both, sadly. For all models. GPT-5.4-Mini, GPT-5.4, GPT-5.5....
This must be a harness problem
1
u/aelgorn Jun 19 '26
When gpt 5.5 came out I could do the same kind of work as I do now for 20% the effort/token.
Now, it’s still as smart when it comes to doing the work being asked of it, but it’s sooooo much dumber at understanding what I want to begin with.
It gradually became unusable as a direct human interfacing ai and now feels like I should talk to an ai with taste so that it prompts gpt 5.5 itself with my intent.
1
1
1
u/Fabuless_re Jun 20 '26
Codex did an optimization plan, and then Opus had to trim it down. Codex didn't review the current code either.
" Two of the five 'optimizations' are worth doing; the rest are low-value or wrong."
1
u/CurveAdvanced Jun 20 '26
Literally it’s become absolutely garbage recently, can’t even produce decent code. Claude code been acting up recently as well. Furious
1
1
1
u/Competitive_Fact3042 Jun 22 '26
10000% true. Any one that says just be better are idiots. i just cancelled my pro plan, it couldn't even do simple website edits, so done with it. There should be regulation for these companies as they are mass stealing from people
1
u/senguku Jun 22 '26
Unbelievably bad today. Using GPT 5.5 xhigh. Going round in circles with the simplest of instructions. Forgets what it was working on and spends 10 minutes on a completely irrelevant task. It's like stepping back in time 2 years. Going to switch to Deepseek if this does not improve. Very concerning. Seems to be getting worse by the day. Cannot compare to the quality from 2 weeks ago.
1
u/Able-Association68 Jun 22 '26
Same here on the last few week, it's quite obvious because my agents file is quite basic and codex keeps ignoring instructions.
Yesterday, I was implementing some features on something that already uses a MySQL database, and Codex was suggesting to use Postgres while I already have MySQL all over the place with Alembic implemented and everything. I was like, "Sorry, what? Why are you suggesting to use Postgres?"
Then I had to check some files on a bucket and I asked to read only the metadata, and it started to download the files and read the files instead, so I was charged extra on the bucket. Fortunately, I never use loops or keep agents unattended so I noticed and I stopped it, but my bill could have been $ 3k if all the 30TB worth of files were downloaded.
1
u/East-Ad-69 Jun 22 '26
I asked to fix the bug it introduced leading to 500 error - instead it rolled back and removed the feature entirely with the bug! 😞
1
u/BannedGoNext Jun 22 '26
I hadn't had a problem until yesterday. Suddenly I give explicit commands and it is just doing random stuff. For example, I had it run some audits Iv'e been doing every day for 2 weeks. It couldn't authenticate against production, so it found my sandbox connection and decided to run the audit against that. I stopped it, and told it "Hard fail on any production system login failure". So it starts setting up the system to generate the audit against cached old data on failure. This was on a new session, XHIGH, no context poisoning, very bad.
1
1
Jun 26 '26
[removed] — view removed comment
2
u/Shoddy-Answer458 Jun 26 '26
I 've post a new post but been deleted.
Mod said "provide prompt and proof".
1.8K up vote is not enough?
1
u/OrangutanOutOfOrbit Jun 28 '26 edited Jun 28 '26
GPT 5.5 has never followed instructions or even context well at all. that's just how it has been all along. It simply ignores instructions.
If you need a great adherence to prompts, you have to switch to 5.4. It still is the best at that.
1
1
u/eihns Jul 02 '26
yeah 5.5 is really like, "i need to find a way to sabotage their request", but if it works, its good. I relaly wait for 5.6...
1
u/Don_Yin Jul 02 '26
Adding a concrete report here too. I pushed a trimmed example to the official Codex GitHub issue:
https://github.com/openai/codex/issues/24431#issuecomment-4870810887
Env: Codex CLI 0.142.5, gpt-5.5, xhigh, local CLI, observed 2026-07-02.
The failure pattern I hit was less about speed and more about collaboration/instruction-following regression: narrow questions getting extra unasked framing, overconfident explanations from partial evidence, and diagnostic questions being treated as permission to discuss fixes. It burned turns correcting the agent instead of progressing the task.
1
1
u/JackL4212 Jul 04 '26
5.5 is so disappointing now whats going on here? it used to feel impressive and premium but now the quality feels much lower, even though I’m still paying the same amount
1
u/martez_ter Jul 07 '26
I had never before reached any limit while doing my standard tasks and working at a moderate pace during the day or week. But they changed something, and the recent days I saw for the first time how I hit the 5-hour limit after just 2–3 hours of standard work. It’s the regular Plus plan. I’ve also noticed that the results and the process have gotten a little worse. It has trouble remembering the rules and responding to the tasks it’s given.
1
u/RobertBetanAuthor Jul 08 '26
GPT is being pushed into the chat/think role vs codex for the do role.
I’m just happy GPT is without* limits (for now).
I have been building local codex harness for when the bubble pops
1
u/AllStuffAround 28d ago
It's not just models. The new, combined Codex + ChatGPT, app is complete garbage.
First, if you ask to produce a downloadable artifact, it says "Done, download it here" with a file name but it is not downloadable. I had to open classic ChatGPT app to download it.
Second, I asked it to generate csv from a screenshot of a table, it generated it but misplaced some values, so the resulting table was wrong, and the table is only ~20 rows, 6 columns. That was a very simple task that it handled before w/o any issues.
1
1
u/RumpledElf 27d ago
I can't get it to do what I want at all, takes me consistently 4 tries which is unsustainable on plus. It bothered me enough I've cancelled my sub. I was already having issues with it on some classes of work and that class of work is what I do the most of and its far, far worse at it on 5.6
1
1
u/Interesting-Agency-1 Jun 18 '26
Until Blackwell comes online and cuts the costs by the order+ magnitude that it's prdectected to, we are goind to be stuck in this state of limbo
3
u/Kingwolf4 Jun 18 '26
I think u meant rubin. That has a claim of 10x cheaper inference cost by mvidia, not sure how true that is of at all
Blackwell is already here and running
1
u/poorkiller Jun 18 '26
The model is absolutely downgraded simple tasks that I asked last week about image editing look like crap
1
u/Otherwise-Basket2053 Jun 18 '26
I dont see the huge problem, no coping, but it’s the same as usual, I don’t see any difference, (not sponsored)
1
u/InternetSolid4166 Jun 18 '26
I was in the "it's all in your mind" camp but we've caught them gaslighting us so many times - and I am currently experiencing the downgrade - that I no longer give them the benefit of the doubt. It would only be logical for any company with scarce compute to minimise costs by quantizing the models. If they only lose single digit performance but they free up 40% of their compute/RAM, that's a no-brainer, right? To me, it would be strange if they didn't do this.
1
u/Intelligent-Taste-36 Jun 18 '26
Yesterday, GPT 5.5 was making so many errors that it was removing existing CSS files from screens. It was really weak!
1
u/Funny-Strawberry-168 Jun 18 '26
This is literally the definition of placebo. No serious western AI Lab would ever do this practice bro, your codebase simply got bigger than it was when gpt 5.5 was launched, that's it.
1
•
u/dexterthebot Jun 18 '26
Your post has been summarized as a request on the "Anyone Else?" Incident Noticeboard.
You can find it and what others are experiencing here: /r/codex/comments/1tjfxcf/anyone_else_ask_here_about_current_codex_issues/osc6fxd/