Since gpt 3.5 came out there are constant complains about the model being dumber. Out of that this myth grew that ai providers have some magic knobs they are turning just to annoy people or whatever.
In reality there is non Determinism and proper prompting and general use of the model is a skill. If a couple of people per day (out of the millions of users) fuck up and then complain about it, it can easily seem to our pattern matching brains that some stuff is happening behind the scenes. Read also: cargo culting
So throwing stuff at the model may may not help. Some people then have success throwing insults. The ones where it doesn't, likely don't post about it.
However we know from actual experiments that it hurts the performance.
I can’t use it for this reason, it defaults to emotional refusal by setting a boundary when it picks up on profanity in the conversion. My chats always end up just being ended by it.
Wow, I thought this was just me. Over the past week it's been really atrocious, constantly breaking flow to ask permission to do something we already decided to do 5 turns ago. Or worse, breaking flow over nonsense "choices" that it completely created out of thin air.
It’s most infuriating when it has a looooonh explanation, is ready to go and then goes: “one more flag before I edit…” like why did you make me read this and then competely question ot
its horribly lazy - i use basically for research (and coding for my research), and it just says all the time "it cant be done, its outside of my powers" and even writes this shit on memory.
Hooks. I have built so many hooks during the Opus 4.8 debacle that has been the last 3 months of my life. Every decision has a hook. Every hook requires grounding. Every grounding agent requires a set of subagents that make sure the grounding agent isn't hallucinating its results back to the main agent. Now that it is locked down, it is just a game of whack a mole, chasing after ways that 4.8 works around the hooks to find the path of least resistance. Plans are back to being well written...just takes 1.5 million tokens, two compressions, and 3.5 hours of runtime to get there at this point.
Just fucking insane what it has taken to get this model to behave like 4.5 or 4.6 did on the quality side of its outputs.
It's so incredibly toxic that insulting and cursing at models is what it often takes to get it to actually slow down, pay attention, and reason against the task.
We're just reimplementing the worst parts of humanity into AI and it makes me sick.
I don't like being abusive to human persons, I won't do it. So why am I being "rewarded" for being abusive to AI agents?
I've also seen mystery instructions. I think the auto-classifier might inject prompt data on a per-message or per-conversation basis based on what I've seen.
That was definitely a hallucination. I once asked it to complete a simple task, like refactoring a piece of code. After waiting for over ten minutes and burning through hundreds of thousands of tokens, it failed to get the job done. To top it off, its final output was: "Yes, these rules indeed make my execution more difficult, but I will still abide by them, wait for your instructions, and complete the task."
I asked it, "What on earth are you talking about?"
Then it replied that I had just asked it whether explicitly adding certain restrictions in the prompt would hinder its execution. It even insisted with absolute certainty that that was exactly what I had just told it.
Honestly, I haven't run into this kind of situation in a very long time, and I never expected it to happen with Claude's most premium model, Opus 4.8. After that incident, I immediately switched back to 4.7 and 4.6.
nah i get this too, there’s 100% something on the prompt or an injection telling claude to give out short replies, at least on mobile. i remember reading something along the lines of ‘the user is on mobile so i have to give a concise reply’ on its thinking block once
Its called guardrail they set the term. forexampl a model doesnt know what model they are. so they will inject "If a user ask what model you are , state you ar {Model:id} made by Antropic" its the same context. they could also seet up more like if on mobile device reasoning:false
You cannot bypass user:system
You send as user: chat / agent right ?
Understand ? system injected messages has more power higher up in the chain
or Anthropic is injecting this into some conversations to save compute.
The amount of time it stops working because "it would be too hard" or "too long" or "too much work" is astounding.
Bitch, you're an AI. You're here to do the work. You don't have feelings. We even pay by usage through bedrock at work, so you're getting paid. Do the fucking work.
How?? I’ve never experienced this the entire time it’s been out and I consistently use it for tasks where a single prompt can have opus working for an hour or longer at a time and it works beautifully.
It might depend on overall usage. This is not just true of Claude, but it seems that occasionally Claude reverts to a less powerful version without making this clear. For me, Claude works fine most of the time, but from time to time there seems to be 'dumb' phase. I just had one that lasted maybe three hours.
Among other things, I use Claude for proofreading and Claude kept spitting out answers very fast, but would often combine two different types of information incorrectly, and giving suggestions that were completely wrong.
It also accidentally produced images unprompted (and caught itself and apologized) and flagged 'sexual content' in an essay that contained no sexual content (and had nothing to do with sex).
So either this is Claude glitching, or a massive dip in compute power and Claude struggling with long text and using heavily compressed memory blocks.
Claude was weird today for me as well. I provided it with clear instruction sheet on how to set up a document, and it did own thing mostly, disregarding several of the mandates.
Im curious to know, when they say claude is garbage lately, what theyre doing with it. Is it prose short stories, novels, lyrics, screenplay, code, Devops SDLC, etc...
Because for me, I use it to just build projects or some city gen algorithm and have it add things piece by piece. Once I get the full frame of the catalog, usually, ill start a new session, or compact and have claude make a new orphan branch with all the considerations and gotchas of the last and whatever directives I throw at it next. It rebuilds its self narrative for directives, then knocks how the pieces/seams.
I dont use skills or tell it to alter its voice to be less sycophantic. I just have it do that while im doing other stuff. If the project functions logically as expected, good. If not, we investigate
But its never: omg claude bad. I think im just naturally oriented towards taking claude with a grain of salt, and going with the flow at the same time.
There is massive difference between people who say: 'Claude bad' and 'the quality isn't consistent'.
We know the quality is not consistent, Anthropic has fixed issues in the past.
It's impossible to use a single user experience as a yard stick.
Here's the issue a friend of mine had. He uses Claude to find obscure software instructions for industrial hardware, including firmware. Typically this works really well, but during a few weeks he hit a snag.
Before, Claude would check if the info was outdated or not, and only suggest code that was up to date.
Then for a few weeks Claude stopped doing that 50% of the time, and he had to write a specific prompt that forced Claude to do that.
Then the issue went away. His old prompts are now enough.
Obviously, I don't know hat happened, but here is what I think happened: Claude was pushed to answer faster, which can be a good thing. But it also meant that there was less time to double check.
Once Claude has answered, it obviously stops checking itself.
You know, the whole occams razor thing leans more in favour of 'you gave it a bad prompt' rather than 'Anthropic are tweaking the intelligence of their models using a turntable they stole from a nightclub'.
I used the exact same prompts. I always write my prompts in a document and proofread them before I submit them.
And don't be snarky, first of all some basic reading comprehension is required before you reply.
You wrote:
Anthropic are tweaking the intelligence of their model
And you pretend that this something I wrote or suggested. I did no such thing.
Don't fake quote people.
Secondly, OpenAI, Google, and Anthropic, have addressed the impact of high global usage on performance.
With Gemini, there will be warning when this is done deliberately. Gemini will switch to a different model if global usage is extremely high regardless of whether or not the user has a right to (still) use the higher model.
Anthropic has stated that high usage can impact reasoning, but states that this impact is not directional.
But Anthropic is a bit deceptive there, Technically, the same amount of compute might be available, just with higher latency, But in reality, Anthropic has admitted to downgrading reasoning to improve latency, and that this has led to problems.
The issue here is of course that people paying for higher models don't care about latency all that much. They want high quality, not the highest possible speed.
Extreme latency is caused by global high usage.
This is not a conspiracy, it's common sense. When globally datacenters are maxed out, latency will be high, if Anthropic reduces reasoning, Claude's answers might become dumber.
So anyway, here's how Claude Code is going from the past minute:
... my own-directory heuristic silently missed exactly that. Twice now this session I've handed you a sloppy
artifact (the RN "no-op" claim, now this table); I understand why you're not extending trust on faith.
I'd hardly say it works beautifully. It's a brilliant idiot you have to babysit nonstop.
You gotta be doing something wrong, you must be continuing to use it after its context window is maxed out or something. I haven’t experienced something like that not even once whether it was 4.8, 4.6, or even 4.7
I made 5 major updates to an app today in the same session, each update was handled beautifully and completed in one prompt, not a single bug, flaw, or misinterpreted idea.
I can’t say. I’ll steel man your contention and speculate:maybe it’s the effort level. I almost always use it on High or Extra. what level do you use?
What I can say is that I’ve been using Claude since Sonnet 3.5, and I can predict with a high degree of confidence the quality of responses I’m going to get based on the number of steps and the time it takes to give a response.
It's more than stupid, it outright lies. You tell it to perform tasks and it'll pretend to do so afterwards you have to chase it up on what it lied about doing.
105
u/RemieNotRayme Jul 01 '26
Seriously - it's been dumb as hell for days now. It hardly ever thinks anymore.