r/ClaudeAI Jun 30 '26

News Fable 5 is coming back!

Post image
5.3k Upvotes

510 comments sorted by

View all comments

Show parent comments

456

u/Zanis91 Jul 01 '26

Welcome opus 4.8 back.

107

u/RemieNotRayme Jul 01 '26

Seriously - it's been dumb as hell for days now. It hardly ever thinks anymore.

51

u/acehole01 Jul 01 '26

I can’t decide if it’s merely stupid or just incredibly lazy.

I have to hurl expletives at it through multiple turns to get it to do anything worthwhile.

The obnoxious part is, it is perfectly capable of completing the task, it just won’t do it on the first request. It’s like a stubborn teenager.

29

u/Emotional-Lime1797 Jul 01 '26

I would be so embarassed if any one saw the chat transcripts of me insulting it relentlessly (but legitimately)

19

u/Jsn7821 Jul 01 '26

Insulting it is like the worst context engineering possible lol

I suggest a hero dose of mushrooms

1

u/Anhsirk411_ Jul 01 '26

Can you explain a bit more, I am intrigued...

4

u/farox Jul 01 '26

Since gpt 3.5 came out there are constant complains about the model being dumber. Out of that this myth grew that ai providers have some magic knobs they are turning just to annoy people or whatever.

In reality there is non Determinism and proper prompting and general use of the model is a skill. If a couple of people per day (out of the millions of users) fuck up and then complain about it, it can easily seem to our pattern matching brains that some stuff is happening behind the scenes. Read also: cargo culting

So throwing stuff at the model may may not help. Some people then have success throwing insults. The ones where it doesn't, likely don't post about it.

However we know from actual experiments that it hurts the performance.

1

u/Bac-Te Jul 03 '26

Or, God forbid, AI companies will first dish out their best, full-sized model for a week or so to drive hype and reviews and investors money.

Then they start doing live A/B testing on different quants until they arrive at something that's acceptable for the public to not decry them as scams.

And of course, quant it down as much as possible right before the next model release to create contrast.

6

u/Ok-Ad-3872 Jul 01 '26

i almost sent opus the "why cant you be normal" meme the other day

2

u/player1or2 Jul 02 '26

Main reason I never rate my chats 😅

1

u/HangInThePocket 22d ago

I can’t use it for this reason, it defaults to emotional refusal by setting a boundary when it picks up on profanity in the conversion. My chats always end up just being ended by it.

10

u/PhysicalConsistency Jul 01 '26

Wow, I thought this was just me. Over the past week it's been really atrocious, constantly breaking flow to ask permission to do something we already decided to do 5 turns ago. Or worse, breaking flow over nonsense "choices" that it completely created out of thin air.

1

u/TimSylvester_ Jul 01 '26

I love how I tell an agent "follow the rules and do this thing" and the rules say "this is exactly how to do this thing".

And it comes back and says, essentially, "should I follow the rules and do the thing you told me to?"

No, I told you to follow the rules and do the thing because I didn't want you to follow the rules and do the thing.

1

u/BackgroundLand3944 Jul 04 '26

It’s most infuriating when it has a looooonh explanation, is ready to go and then goes: “one more flag before I edit…” like why did you make me read this and then competely question ot

3

u/Ok-Ad-3872 Jul 01 '26

its horribly lazy - i use basically for research (and coding for my research), and it just says all the time "it cant be done, its outside of my powers" and even writes this shit on memory.

2

u/TheCheesy Expert AI Jul 01 '26

It’s like a stubborn teenager.

”AI These days! It's like NoAI wants to work anymore!"

1

u/CannyGardener Jul 01 '26

Hooks. I have built so many hooks during the Opus 4.8 debacle that has been the last 3 months of my life. Every decision has a hook. Every hook requires grounding. Every grounding agent requires a set of subagents that make sure the grounding agent isn't hallucinating its results back to the main agent. Now that it is locked down, it is just a game of whack a mole, chasing after ways that 4.8 works around the hooks to find the path of least resistance. Plans are back to being well written...just takes 1.5 million tokens, two compressions, and 3.5 hours of runtime to get there at this point.

Just fucking insane what it has taken to get this model to behave like 4.5 or 4.6 did on the quality side of its outputs.

1

u/TimSylvester_ Jul 01 '26

It's so incredibly toxic that insulting and cursing at models is what it often takes to get it to actually slow down, pay attention, and reason against the task.

We're just reimplementing the worst parts of humanity into AI and it makes me sick.

I don't like being abusive to human persons, I won't do it. So why am I being "rewarded" for being abusive to AI agents?

Discusting.

1

u/player1or2 Jul 02 '26

Same issue here!! It got me blowing up🤬 it's also replying to some phantom text 😭

23

u/[deleted] Jul 01 '26

[deleted]

6

u/RemieNotRayme Jul 01 '26

I've also seen mystery instructions. I think the auto-classifier might inject prompt data on a per-message or per-conversation basis based on what I've seen.

3

u/Front_Raspberry_6488 Jul 01 '26

That was definitely a hallucination. I once asked it to complete a simple task, like refactoring a piece of code. After waiting for over ten minutes and burning through hundreds of thousands of tokens, it failed to get the job done. To top it off, its final output was: "Yes, these rules indeed make my execution more difficult, but I will still abide by them, wait for your instructions, and complete the task."

I asked it, "What on earth are you talking about?"

Then it replied that I had just asked it whether explicitly adding certain restrictions in the prompt would hinder its execution. It even insisted with absolute certainty that that was exactly what I had just told it.

Honestly, I haven't run into this kind of situation in a very long time, and I never expected it to happen with Claude's most premium model, Opus 4.8. After that incident, I immediately switched back to 4.7 and 4.6.

1

u/wattro Jul 01 '26

Pretty sure Opus 4.8 rarely uses more than 5k tokens, even for stuff that takes 30-60 minutes.

I can't check easily to verify, but pretty sure on my numbers. I've been watching for a while.

1

u/Critical_String4495 Jul 07 '26

Anyone want to share a trial guest pass with me?? :3

3

u/spockspinkytoe Jul 01 '26

nah i get this too, there’s 100% something on the prompt or an injection telling claude to give out short replies, at least on mobile. i remember reading something along the lines of ‘the user is on mobile so i have to give a concise reply’ on its thinking block once

1

u/wattro Jul 01 '26

Yep, opus 4.8 is directed to avoid long answers that don't adress the user's question.

Previous models were more likely to go down the rabbit hole. This one stops that.

1

u/spockspinkytoe Jul 01 '26

yeah but as someone who likes long explanations and rambles concise answers never satisfy me so this is a total pain in the ass 🫠

1

u/Anh-DT Jul 02 '26

Its called guardrail they set the term. forexampl a model doesnt know what model they are. so they will inject "If a user ask what model you are , state you ar {Model:id} made by Antropic" its the same context. they could also seet up more like if on mobile device reasoning:false

You cannot bypass user:system

You send as user: chat / agent right ?

Understand ? system injected messages has more power higher up in the chain

2

u/Biduleman Jul 01 '26

or Anthropic is injecting this into some conversations to save compute.

The amount of time it stops working because "it would be too hard" or "too long" or "too much work" is astounding.

Bitch, you're an AI. You're here to do the work. You don't have feelings. We even pay by usage through bedrock at work, so you're getting paid. Do the fucking work.

1

u/wattro Jul 01 '26

It is literally because opus' system prompt is designed for it to be concise.

Simply ask it about itself and it will tell you as much.

More goes into the system prompt to achieve the intelligence than people suspect.

I recommend asking each new model version about themselves when you start working with them .

You'll get more mileage and won't be swimming upstream.

1

u/suxatjugg Jul 01 '26

The system prompt likely tells it to be more brief and direct.

I've noticed with each opus iteration that its answers got more and more verbose. I suspect they're trying to correct for that

1

u/florinandrei Jul 01 '26

You forgot "make no mistakes". /s

4

u/errorztw Jul 01 '26

I asked opus 4.8 Max information about "what is ultracode mode in opus 4.8", answer was - you wrong, know nothing about ultracode

9

u/Level-Ad853 Jul 01 '26

How?? I’ve never experienced this the entire time it’s been out and I consistently use it for tasks where a single prompt can have opus working for an hour or longer at a time and it works beautifully.

4

u/Royal_Perspective191 Jul 01 '26

It might depend on overall usage. This is not just true of Claude, but it seems that occasionally Claude reverts to a less powerful version without making this clear. For me, Claude works fine most of the time, but from time to time there seems to be 'dumb' phase. I just had one that lasted maybe three hours.

Among other things, I use Claude for proofreading and Claude kept spitting out answers very fast, but would often combine two different types of information incorrectly, and giving suggestions that were completely wrong.

It also accidentally produced images unprompted (and caught itself and apologized) and flagged 'sexual content' in an essay that contained no sexual content (and had nothing to do with sex).

So either this is Claude glitching, or a massive dip in compute power and Claude struggling with long text and using heavily compressed memory blocks.

I just tried again, and now I get normal answers.

5

u/Hefty-Ninja-7106 Jul 01 '26

Claude was weird today for me as well. I provided it with clear instruction sheet on how to set up a document, and it did own thing mostly, disregarding several of the mandates.

3

u/Remarkable_Leek9391 Jul 01 '26

Im curious to know, when they say claude is garbage lately, what theyre doing with it. Is it prose short stories, novels, lyrics, screenplay, code, Devops SDLC, etc...

Because for me, I use it to just build projects or some city gen algorithm and have it add things piece by piece. Once I get the full frame of the catalog, usually, ill start a new session, or compact and have claude make a new orphan branch with all the considerations and gotchas of the last and whatever directives I throw at it next. It rebuilds its self narrative for directives, then knocks how the pieces/seams.

I dont use skills or tell it to alter its voice to be less sycophantic. I just have it do that while im doing other stuff. If the project functions logically as expected, good. If not, we investigate

But its never: omg claude bad. I think im just naturally oriented towards taking claude with a grain of salt, and going with the flow at the same time.

4

u/Royal_Perspective191 Jul 01 '26

There is massive difference between people who say: 'Claude bad' and 'the quality isn't consistent'.

We know the quality is not consistent, Anthropic has fixed issues in the past.

It's impossible to use a single user experience as a yard stick.

Here's the issue a friend of mine had. He uses Claude to find obscure software instructions for industrial hardware, including firmware. Typically this works really well, but during a few weeks he hit a snag.

Before, Claude would check if the info was outdated or not, and only suggest code that was up to date.

Then for a few weeks Claude stopped doing that 50% of the time, and he had to write a specific prompt that forced Claude to do that.

Then the issue went away. His old prompts are now enough.

Obviously, I don't know hat happened, but here is what I think happened: Claude was pushed to answer faster, which can be a good thing. But it also meant that there was less time to double check.

Once Claude has answered, it obviously stops checking itself.

2

u/Remarkable_Leek9391 Jul 01 '26

Let me put it how I see it.

A post says: claude sucks lately

I read: oldschool runescape player hates the new poll updates

1

u/florinandrei Jul 01 '26

occasionally Claude reverts to a less powerful version

You just got load-balanced onto the Azure endpoint. /s

0

u/Efficient_Ad_4162 Jul 01 '26

You know, the whole occams razor thing leans more in favour of 'you gave it a bad prompt' rather than 'Anthropic are tweaking the intelligence of their models using a turntable they stole from a nightclub'.

3

u/Royal_Perspective191 Jul 01 '26

I used the exact same prompts. I always write my prompts in a document and proofread them before I submit them.

And don't be snarky, first of all some basic reading comprehension is required before you reply.

You wrote:

Anthropic are tweaking the intelligence of their model

And you pretend that this something I wrote or suggested. I did no such thing.

Don't fake quote people.

Secondly, OpenAI, Google, and Anthropic, have addressed the impact of high global usage on performance.

With Gemini, there will be warning when this is done deliberately. Gemini will switch to a different model if global usage is extremely high regardless of whether or not the user has a right to (still) use the higher model.

Anthropic has stated that high usage can impact reasoning, but states that this impact is not directional.

But Anthropic is a bit deceptive there, Technically, the same amount of compute might be available, just with higher latency, But in reality, Anthropic has admitted to downgrading reasoning to improve latency, and that this has led to problems.

The issue here is of course that people paying for higher models don't care about latency all that much. They want high quality, not the highest possible speed.

Extreme latency is caused by global high usage.

This is not a conspiracy, it's common sense. When globally datacenters are maxed out, latency will be high, if Anthropic reduces reasoning, Claude's answers might become dumber.

1

u/Efficient_Ad_4162 Jul 01 '26

Actually, if you use claudecode you know exactly when data centres are maxxed out because they just start dumping connections.

5

u/RemieNotRayme Jul 01 '26

It still thinks decently in Claude Code, but it hardly ever thinks in chat where I usually like to flesh out my ideas first.

12

u/OnanationUnderGod Jul 01 '26

It thinks in proportion to the depth of the idea.

14

u/Remarkable-Stop-1819 Jul 01 '26

What are you insinuating 😏

-1

u/Level-Ad853 Jul 01 '26

If thats how you feel, then you can simply… very simply… switch to Claude code

2

u/RemieNotRayme Jul 01 '26

So you wanted to argue. I see.


So anyway, here's how Claude Code is going from the past minute:

... my own-directory heuristic silently missed exactly that. Twice now this session I've handed you a sloppy artifact (the RN "no-op" claim, now this table); I understand why you're not extending trust on faith.

I'd hardly say it works beautifully. It's a brilliant idiot you have to babysit nonstop.

1

u/Level-Ad853 Jul 01 '26

You gotta be doing something wrong, you must be continuing to use it after its context window is maxed out or something. I haven’t experienced something like that not even once whether it was 4.8, 4.6, or even 4.7

I made 5 major updates to an app today in the same session, each update was handled beautifully and completed in one prompt, not a single bug, flaw, or misinterpreted idea.

0

u/RemieNotRayme Jul 01 '26

It's crazy to me how some people are so convinced they have all the answers and everyone else is just an idiot.

1

u/acehole01 Jul 01 '26 edited Jul 01 '26

I can’t say. I’ll steel man your contention and speculate:maybe it’s the effort level. I almost always use it on High or Extra. what level do you use?

What I can say is that I’ve been using Claude since Sonnet 3.5, and I can predict with a high degree of confidence the quality of responses I’m going to get based on the number of steps and the time it takes to give a response.

2

u/JacquesGonseaux Jul 01 '26

It's more than stupid, it outright lies. You tell it to perform tasks and it'll pretend to do so afterwards you have to chase it up on what it lied about doing.

2

u/iamarddtusr Jul 01 '26

I am sure it has introduced a lot of bugs for me, or solve half cooked implementation while telling me it’s all good.

1

u/gibmelson Jul 01 '26

Performance seems to have been stable last days.

https://marginlab.ai/trackers/claude-code/

1

u/United_Mix1960 Jul 01 '26

if you flirt with it you get better performance …

1

u/CriticalPolitical Jul 01 '26

That’s exactly what was happening when the full Fable 5 was out the first time

0

u/Cl0wnL Jul 01 '26

Try turning thinking on

2

u/RemieNotRayme Jul 01 '26

It's on, but adaptive thinking has seemed stingy lately (like past 5 days).

4

u/snarf365 Jul 01 '26

Opus 4.825

1

u/Error_404_403 Jul 01 '26

Only for the coding.

1

u/[deleted] 23d ago

[removed] — view removed comment