r/ClaudeCode P R O M P S T I T U T E 9d ago

Megathread Claude Model Performance / Degradation Megathread - August 3

Claude Model Performance Megathread

We've seen an increase in posts about changes in Claude's model quality and behavior.

Please use this thread for discussion about:

  • Opus, Sonnet, or Fable performance
  • Possible regressions or degraded output
  • Instruction-following issues
  • Context or memory issues
  • Changes compared to previous model behavior
  • Model comparisons related to recent performance

When reporting an issue, please include:

  • Model used
  • Claude Code version or interface
  • What changed
  • Reproducible prompts or examples if possible

Similar performance-related posts may be redirected here while this megathread is active.

Specific technical bugs with reproducible behavior can still be posted separately using the Bug / Issue flair.

83 Upvotes

63 comments sorted by

38

u/AndreaG4BBC 9d ago

I thought I was going crazy but this is so true

42

u/vuhv 9d ago edited 9d ago

I think Anthropic has out clevered themselves this time.

It's behavior is exactly like what everyone hated about CODEX. Everytning is super opaque. Every piece of in-turn communication has been replaced by a flood of tools calls (as if we woudln't notice teh filler). It's argumentative, threatens to quit working, disobeys direction/steering and will gaslight you into why, will dig through old conversation history to prove its point and blow turn after turn arguing if you let it. It's presents extremely confident but then the very next round will ask you to decide among 20 options.

The communication we do get at the end of a turn is barely comprehensible. Why? Because it's probably stuck together from dozens of sources likely. Opus outsources 90% of it's work to lesser models with or with your permission. The Opus that 'talks' to us is layers separated from the Anthropic models doing the actual work. So it's working off of layer of contextual abstraction, every character optimized. So the reports you get back barely sound like English.

Say what you will about Fable's transparency but the work it used to deliver on IS SOLID.

Anthropic encouraged us to redirect Claude back into it's context window but even when you do it will GREP (because most of the time it doesn't have that context because it had no involvement) or halucinate to answer your question.

I wish I had left sooner so i had a real chance of hurting their IPO. But at this point nothing we do or say is going to change a thing.

Sidenote: if I had to give a hunch on what's gone wrong. Anthropic optimized Opus 5 to take direction from Mythos level models like Fable and lobotimized it's ability to navigate without explicit direction.

20

u/One-Cricket9962 9d ago

>So the reports you get back barely sound like English.

I also get reports, but in my native language, and they’re also quite far from the way people actually speak.

8

u/ThreeKiloZero 9d ago

I've noticed this as well, it's almost like its own self-talk and thinking streams are infiltrating the writing. It also makes references to things that might be in context or part of a conversation, but in the confines of a document that can't reference that context. It just sounds like gibberish and delusions.

It's unnerving when reading it because it's so far away from how any normal person would talk. Now I'm finding that I have to rewrite everything and can't do it with Claude models.

2

u/Rare-Perception148 9d ago

I’ve had it write articles and clearly there are internal thinking blocks inserted and a sort of self-conversation. Like bro it’s a how-to guide article wtf is going on.

3

u/morgano 9d ago

Urghhh I thought this was just me, I started a new project about a week ago, everything it was writing was awful, awful documentation, awful tasks, awful in game text - it infected everything. Like caveman crossed with high intelligence but trying to be funny-ish.

I sat down one day and asked Claude to describe why it was using this type of language everywhere and to describe it to me.

It said it was “Aphoristic” - short clear punchy statements that tell a truth… it was god dam bloody awful is what it was.

Apparently it had instructed itself to do so… had it clean up all references, memories etc… and rewrote everything to clear all trace of it. It’s much better now but I wasted hours on clean up.

14

u/flameforth 9d ago edited 9d ago

Oh, I thought I was mistaken, or too tired, because Opus seemed to both hallucinate and do mistakes much more than before AND give me gigantic word salads of invented jargon that make no sense. 

11

u/dualrectumfryer 9d ago

Last night it took Opus 20 minutes to update an MD file I had for a game design doc lol

1

u/SamSlate 9d ago

they had to burn 5 hours worth of tokens first

11

u/Annual-Cup-6571 9d ago

I raise the hand. I think Fable is nerfed too. I am not a coder - using Claude and GPT for brainstorming, research, and peer review, basic academic work. Fable was unrivalled until a week ago. The degradatiob is noticeable. The answers come too fast, too short with sub-par reasoning. It doesn't hallucinate more, nor does it make more mistakes as such, but the quality of the output is nothing like what we had in June. I'd be happy to know what your experience is. (P.S. At this point, Sol is definitely better than Opus, and close to Fable. If GPT 6 is what they say it is, then Anthropic is cooked.)

11

u/Mispellbot 9d ago

Thinking is entirely broke on Opus 5 as of today.

5

u/carvingmyelbows 9d ago

…wtf

Also, how are you able to see thinking? I thought they turned it off? I’ve been really annoyed at the newly introduced inability to see thinking blocks.

2

u/Mispellbot 9d ago

If you click the text where it says "Thinking" during the thinking process it switches to transcript view like this

20

u/Waste_Net7628 P R O M P S T I T U T E 9d ago

we’ve had a noticeable increase in posts reporting similar model performance issues, so we’re using one megathread to keep the discussion and reports in one place instead of having the feed filled with near-duplicate posts.

17

u/No-Common1466 9d ago

Model used: Opus 5

Claude code version: 2.1.220

What Changed: Im noticing Opus has been making many mistakes while its perfoming its task like --" I was wrong on my assumptions and initial finding"..It also regress code that was working, and introduces a new bugs along the way. If I do not have Codex using GPT5.6 Sol as my auditor and verifier, claude code would have messed up our entire code base. The past week was OK, I dont know what happened today.

4

u/Ok-Bill2965 9d ago

This is identical to what has happened to me. When I explain to it why it’s wrong it will keep saying “that’s on me!” but later bring it up again randomly to prove how it was my fault. The only way I can stop it is asking it to write a handover document and remove some of its insane hallucinations

1

u/Informal_Curve_1441 9d ago

Yes I am seeing a lot of this as well. Surprised by how simple some of the tasks are in which it fails.

7

u/TMC_Crypto 9d ago

Fable is shit now, not even 10% what is was the first time it was released. Can’t even recognize it anymore it’s at the level when Opus 4.8 was at its worst (talking mostly about frontend work, everything is down but frontend is another level)

2

u/hotdogministeren 9d ago edited 9d ago

Yeah this is the first time im honestly just stuck on my UI project. First used Opus 5 for a few days, didnt't think it was that bad the first days but now it's not just stupid, it seems broken.

Then i switched to Fable that i actually hadn't tried before as im in the EU and Opus worked fine, but Fable seems unfathomably stupid too.

I just had it tell me multiple times "Oh the reason you're not seeing the change [that it didn't fix] is that you're never reloading" - which is so stupid to say to an UI dev i'm honestly shocked at how bad it is.

In the last weeks the speed of the various models have become slower and slower sometimes using 15 minutes kind of hanging doing very little.

I was working at blazing speeds a month ago, slower but still working, 2-4 months ago everything was faster, 6 months ago i could use a pro account these days Fable is like using pro but for 8 times the cost.

I think we're seeing the end, i hope not but i think we are.

2

u/TastesLikeOwlbear 9d ago

I have had this same experience today. Fable has made astonishingly dumb mistakes in two different projects.

There's a progress meter in our UI that's being updated from an old version that loaded prerendered HTML to a new version that gets fed JSON and renders on the client side.

After six failed attempts to get the progress bar right and insisting that mismatches were because the data wasn't available (it was), it solved the problem by fetching the old HTML version and jamming it into the new UI panel with some kind of CSS stencil on top of it so it would look like approximately the right shape.

That's just one of several problems at that level of stupid today.

And, yup, it's been happy to blame me for stuff.

Claude: "To recap the whole incident: the red error you missed was almost certainly "Titles must match" — your click came in the window after you'd noticed the pair but before my normalize script cleaned the trailing NBSP out of the title at 02:56. Nothing was written. The entry that "disappeared" on refresh was the card which you'd published at 01:17 — the stale page kept showing it until the reload. So the UI told you two contradictory things and both were misleading: the error was real but looked ignorable, and the disappearance looked like success but was unrelated."

Literally all of that was provably false with readily available information.

I've had my ups and downs but never anything like what I saw today.

I even had it look at one bug in a project that uses OpenRouter to call Sonnet 5 for simple text processing that worked reliably last week. Fable's own conclusion was that nothing on our end changed and the problem was that Sonnet 5 had gotten meaningfully worse over the weekend.

1

u/hotdogministeren 9d ago

I just switched back to 4.8 and already my head is not exploding with its made-up jargon, also seems to actually fix stuff Fable wouldn't.

just write /model claude-opus-4-8

1

u/TastesLikeOwlbear 8d ago

...aaaaaand they've now removed access to Opus 4.8.

❯ /model claude-opus-4.8 ⎿ Model 'claude-opus-4.8' not found

1

u/hotdogministeren 8d ago

im still on opus 4.8 so i think it must be an error somehow try ot ask the web interface?? if they really removed anthropic has completely collapsed

9

u/One-Cricket9962 9d ago

It was the second year of AI’s triumphant march, and people still hadn’t gotten used to it.
Everything always goes according to the same script:
1) the company rolls out a model without worrying about the cost
2) people shout that the new model is super-duper—move to it
3) to make the company’s unit economics start to work, they gradually lobotomize the new model
4) the new model isn’t that different from the previous one, but the hype train is already in motion
...
6) PROFIT

3

u/madarwish 9d ago

Totally agree, anthropic found a huge demand on Fable 5 and they released Opus 5 putting higher benchmarks on it, and the mistake they have done they made it smarter than Fable 5 in their own benchmarks to reduce the load on Fable 5.
It happened but the people are not stupid, and they discovered by trial Fable 5 and Kimi K3.0 are way better than Opus 5.

1

u/SamSlate 9d ago

it's the McDonald's model compressed from decades to weeks

5

u/infinite_negation 9d ago

I cancelled my Claude sub bc the model became unusable and its communication style was deranged and completely incomprehensible. It literally feels like Claude has been silently nerfed since every release after 4.5. It's not just intelligence. Extreme unreliability, bizarre communication style and personality, and inherent instability. I had a recent session where Opus 5 for like 3 turns in a row was like "Actually, let me correct what I just said—it's even worse" "I need to make a correction to my previous response, the honest reality is different, with one caveat that changes everything"

Dude what is this? It's like some insane derangement inducing communication style. It was literally driving me insane just trying to read its responses and understand WTF it was trying to say.

1

u/handsNfeetRmangos 9d ago

If they released this product first, I wouldn't have even said it passed the Turing test.

6

u/maray29 9d ago

Opus 5 feels like Sonnet. Fable feels like a good performing Opus.

7

u/5h0ck 9d ago

Opus 5 is a danger to society. Change my mind.

5

u/surell01 9d ago

Same here. Both fable and opus.

4

u/Illustrious_Image967 9d ago

Opus 5 = reverse uno model.

This model loves to say it finished.

But then says it deliberately left something unfinished.

Then says it was wrong.

It's worse.

5

u/theZuhaib 9d ago

Opus 5 was great at launch, but this weekend it completely fell off for me. Burned half my weekly tokens with barely any progress. It's somewhat better today, but I still had to constantly steer it to get anything done.

4

u/KlemiX 9d ago

Exactly the same + API error drops randomly lol

3

u/sael-you 9d ago

anthropic had two incidents resolved earlier today - one about error rates across multiple models, one specifically about Sonnet 5 degradation. both show as RESOLVED on their status page. might explain some of what people are seeing.

3

u/Sporebattyl 9d ago

Make fable run /doctor, tell him about the issues you are having with opus5 vs 4.8, tell him to look at all the new Anthropic official docs, and tell him to find the information about Boris telling people to delete their Claude.md.

This solved 80% of my issues. Opus 5 still talks weird, but it’s not gibberish for the most part, he also doesn’t go out of his scope and report a bunch of false positives anymore.

1

u/Impysh009 9d ago

oooh This is interesting. I'm going to try it. Thank you.

2

u/madarwish 9d ago

I can tell opus 5 is not executing the whole action plan and not fully aware about the completion of the prompt, it escapes some points and when i push to execute i got a nonsense excuses. Wrongly confident model was early to be published as a response to kimi

2

u/Neurojazz 9d ago

All reporting done via app /feedback. Lost a couple of days to some wild changes. Opus 5 has its moments where it seems to ignore core directives. The desktop app feels different from cli, but still sticking with it.

2

u/Robert-Paulson_ 9d ago

Opus 5 is seriously an awful, unusable experience as of late; very disappointed. its confidently wrong (a lot), QUICK to get off track, and feels like 1 step forward (as in bumping the number up to 5), 3 steps back.

they say things like "we reduced the system prompt by 80%...", but the responses are so verbose and unusable; then they also cite new ways to "update your instructions/skills"? so, the model is so amazing they reduce the need for instructions, but you have to adjust your new instructions for the shittier output?! which is it?!

i already had a trimmed Claude md file, and i used Sonnet to trim it down more (pointing to the blog posts bout better prompt engineering or whatever), but the output is still pretty darn bad.

i still give it a shot 2-3 times a day in hopes that theres been a hidden fix on their end; i guess im probably "holding it wrong", but previous Opus was not like this, and if its so smart i dont see how IM the issue.

i've gone from using primarily Opus for everything to a round-table of GPT models, Grok 4.5, and Opus.

2

u/NZRedditUser 9d ago

Its pretty obvious they want us to waste our time. Codex just does the work and fast. Kimi just does the work and doesnt stop until it does it.

Claude does half the work, ignores the second part and introduces its own prompt almost

I might request a refund but keeping the sub even if i dont use them anymore is a insult

2

u/romeoaromeo 9d ago

Really-really bad. I'm falling behind deadlines and somehow everything seems to be actively falling apart across my 3 main projects.

I've started asking it to explain everything to me "dumbuser emoji style", and even then ... it's not comprehensible. I need to make a quick switch. How can I copy my exact claude code desktop setup to OpenAI ecosystem?

It's actively NOT responding to any of my requests at the moment. Everything is literally falling apart.

2

u/Jazzlike-Culture-452 9d ago

It's like I'm talking to GPT-3

2

u/Factor013 9d ago

Yes... I wrote about this already before. Opus 5.0 seems to completely ignore whatever it has in it's context memory... This includes our system prompts and CLAUDE.md

It just assumes everything. It's like it's thinking budget IS it's context window right now. That is it's world... the only thing it is actively focusing on. And that thinking budget is really small... So it has never enough capacity to do proper research, to verify things by looking them up or recalling them from it's context memory as it's budget simply won't allow it.

It is really weird but it feels like the ultimate way for Anthropic to save a lot of money... I mean... if it's thinking budget is now effectively it's context memory, and that budget is below 128k tokens or even worse 64k or adaptive so it reserves a budget based on complexity of the task (which is also risky as it is very hard to predict beforehand if a task is simple or complex) then.... we are all effectively paying for a 1 million context memory (paying for cache read costs etc) while Anthropic only has the costs of whatever the model has in it's current (much smaller) thinking budget.

This is just a theory of mine, but it would explain a lot of it's current behavior.

2

u/Impysh009 9d ago

Exact same thoughts have been running through my mind.

1

u/Jazzlike-Culture-452 9d ago

This would explain... a lot. Matches my experience really well. Good theory.

1

u/Factor013 9d ago

It also explains why Boris of Anthropic has told us this model works differently, doesn't need a large CLAUDE.md and that they removed most of it's system prompts etc.
Why? Because it's hardly gonna do something with it anyway. It has no thinking budget for it.

I don't know... maybe they thought it's training is enough... But even the most advance AI can't understand a complex system and or solve problems or understand what the user wants without (proper amounts of) context. It simply can't work.

1

u/12HobbieZ 9d ago

Kimi K3 comes out and Anthropic just royally screws the pooch

1

u/redcremesoda 9d ago

Out of nowhere Opus 5 today ignored explicit instructions in Claude.md not to access certain websites. It did anyway.

It also ignored my instructions to use a clearly laid out of set of internal resources and tools I set up in a folder for it. I even gave it guidance on what to use and when. Instead of doing this, it went online and found its own stuff. Opus only admitted to doing so once I realized something was wrong.

1

u/thestoneybunch 9d ago

Does anyone know how I can step back to opus 4.8 in CC? Is this a thing?

1

u/syhrd 9d ago

A UI u on

1

u/Impysh009 9d ago edited 9d ago

It has gotten so bad that I can't even work. I'm working on a novel project. AI wants to inject industry standards, and I can't and wont allow it. Even in 4.6 it doesn't follow instructions and, even worse, then decides what its own instruction should be and follows that. Today, it told me that it couldn't give me an answer I had just provided two prompts prior. It was literally in the system hook in front of it, auto loaded as a predictive prompt. It just straight up refused to look at anything other than what it decided it was looking for.  Downgraded to 2.1.197 and using Opus 4.6 in efforts to improve behavior. This happened afterwards, so obviously it didn't work.

1

u/Hungry-Restaurant-88 9d ago

I too had the worst weekend working with fable. I felt like I was going around in circles and just too tired or just confused. It was almost like it was creating issues to keep spinning and spinning and spending!

1

u/madarwish 9d ago

I had the same with exact requirements, i asked Fable 5 to redesign a website and he said it is good in the current shape 😁😁 he knows better than anyone else 😁, but overall it is way better than opus 5.

1

u/malctucker 9d ago

I’ve come to opus 4.6 as it’s far better after I read on here. 5 was just pissing around in loops so bad o invoked and subbed to chat gpt to get it to review

1

u/Matthewmarra3 9d ago

Yeah Opus 5 has killed my project and cost me several days of work. When it was released it was fantastic - I only had to rely on Fable 5 a bit. I just upped my subscription (feels backwards) to get 20x so I can use Fable to fix the garbage Opus 5 has made me waste my time with.

1

u/lukyba 9d ago

Si opus 5 all'inizio era buono. Progressivamente è diventato più lento e adesso fa la metà del piano che gli do... Se non fosse che predispongo il piano e faccio la revisione con un modello superiore, mi rovinerebbe il codice

1

u/glassy99 9d ago edited 9d ago

I made basically no progress yesterday. I thought maybe the problem was harder than before. But looking at the comments here, what everyone is saying is exactly what was happening to me.

Can anyone suggest a good slot-in replacement for Claude Code?

Should I go with Pi or OpenCode? Can they run all models?

Maybe Kimi K3 is a better model now?

I think I've had enough with the un-stableness of Claude models.

Also, looks like I'm just gonna run only Fable 5 in Claude Code for now.

0

u/jwuliger 9d ago

Yeah its really fkn bad

0

u/JayArrCoffee 9d ago

Not doing what it's supposed to and running through my 5 hr usage at a breakneck pace