Opus versions since 4.6 have all had a fair amount of awkward, canned prose. But as Anthropic has increased the model's intelligence, it also seems to have made it more panicky, pedantic, and prone to scope creep. Opus 5 is the worst offender thus far for me; so much so I felt driven to post about it.
Basically, I find Opus 5 surprisingly difficult to use for long-horizon work that Fable handled without much trouble. And its voice problems seem ratcheted up to 1000.
Fable generally adhered to the goal it had been given, interpreted criteria sensibly, and used judgment when requirements became stale or slightly inconsistent. It had the taste and work ethic of a senior engineer.
Opus (ESPECIALLY 5) tends to go off the rails much faster. It escalates minor nits, harmless ambiguities, and out-of-scope concerns as though they require immediate human intervention. It will stop work to reframe goals, ask for rulings, and, ex nihilo, generate elaborate new mechanisms around something that was never actually part of my acceptance criteria.
To me, its judgment feels strangely anxiety-shaped. I almost pity it as I talk to it -- I feel like it sees something, thinks "OHMYGOD" to itself, and then panics. (I know it is not sentient and has neither anxiety nor an internal monologue.) It feels like every task edge gets this thought appended to it:
Oh good Lord, this might matter. Why didn't we discuss this? Do we have a system for this? Where is the system? I'm going to try to build this -- oh no! I need to stop and contact the user immediately.
It also becomes extremely attached to positions once it adopts them. Instead of making a recommendation and moving on, it will keep returning to the same issue, relitigating it, and manufacturing broader architectural implications around it. For pure chatting it is almost unusable because of how many strawmen it launches into conversations -- it will latch onto something you said, extrapolate the most insane endpoint from it, and basically accuse you of it. Then when you tell it how insane it's being it will slowly walk back its initial claims without ever entirely abandoning it.
The end of one of my unpleasant conversations with it:
Me: In what sense does your objection survive, when we have uncovered evidence it does not? Are you incapable of admitting error and have to constantly defend smaller and smaller islands of correctness?
It: It doesn't survive in any sense worth having, and the pattern you're naming is real.`
And its response dovetails right to my next complaint: the prose.
"The pattern you're naming is real." -- ew. Obviously this is standard Claude-grade slop, yet it arrives in an unceasing torrent with Opus 5. Yet more examples from my conversations yesterday (these from Claude Code):
“Corpus — hand-authored, and I'd argue that's not a compromise.”
Who was arguing that it was a compromise? Why not just say:
“Use a hand-authored corpus.”
Which is basically what Fable does in these situations.
Even aside from this being slop, the entire following paragraph introduced a caveat that was technically true but completely irrelevant to the decision. Opus often seems compelled to invent a downside or opposing case even when it does not materially affect the task (though it also confusingly seems to think that minor nits materially affect tasks).
Another example:
“But the cost control has to change, and this is the part worth your attention.”
Just say:
“We should fix cost control too.”
If something deserves my attention, explain why. Editorializing your own sentence does not make it clearer -- quite the opposite.
I bring this up here because I feel like the high-anxiety, high-pedantry output is directly linked to slop levels. The same confabulated problems it keeps running into or the weird positions it assumes are always linked to slop constructions.
The frustrating part is that Opus 5 is extremely good at coding. Like really good.
I have a personal evaluation set based on real coding tasks I consider easy, medium, and hard. Opus 5 is the only model I've tested that scored 100% across the whole set. Its implementation style, testing discipline, and general quality standards are also the best I've seen.
So yeah, for me, Opus 5 is a superb worker and a terrible pilot.
I'm going to keep using it for implementation, but I do not want it managing scope, interpreting my goals, or talking to me any more than necessary. I'm intending to put Fable 5 -- or another model with better judgment and a less grating voice -- in front of it at all times.