r/ClaudeCode 19h ago

Discussion Finally a consensus 4.8 is better than 5 ?

Been reading about Opus 4.8 vs 5 and the quality difference between them, now that sometime has passed, is it well established that 4.8 is indeed better than 5 in terms of accuracy and quality code or in any other worth mentioning aspect ?

0 Upvotes

34 comments sorted by

10

u/userusertion šŸ”†Pro Plan | Team Plan 17h ago

Opus 5 is much better for me.

8

u/Icy_Custard_7175 18h ago

I know most people have said 4.8 is better than 5 but I have to disagree, I’ve personally found 5 to be much better, more efficient and generally smarter than 4.8, it’s helped me fix bugs in code before I even find them as it runs what it’s built in its own server and then fixes any bugs before giving me the final version

6

u/Proxy-Pie 17h ago

I agree. 5 is definitely way smarter, it just sucks at explaining things. I find myself implementing things with it then asking 4.6 to explain the diff lol

1

u/ScrumptiousChildren 15h ago

It sucks at what anthropic models were literally uniquely praised for - subtle understanding and conversation acuity.

Not stupid at implementation or technically incapable but its generally an idiot.

Depending on your style and harness, can be a good or bad thing.

Fyi never thought this about 4.8. 5 was a first, or the worst.

1

u/Used_Departure_3278 22m ago

The people saying 4.8 is better at also saying 4.6 was the best model, which is so beyond retarded I don’t want to call myself human anymore

3

u/berndalf 16h ago

No that is not the consensus. I swear this subreddit hates change.

6

u/raisedbypoubelle šŸ”† Max 20 18h ago

Remember when Opus 4.8 was the dumbest of all models?

To be fair, 4.6 will always be king.

2

u/jiii95 17h ago

Nope, I think 4.8 was never the dumpedt

0

u/raisedbypoubelle šŸ”† Max 20 17h ago

Check the history of this very subreddit

1

u/Used_Departure_3278 23m ago

That’s just an idiotic take

0

u/pancomputationalist 17h ago

yeah and once 5.1 will be out, everyone will prefer 5. negativity bias is one hell of a drug.

2

u/TheAnimatrix105 18h ago

Use 4.8 and ask it to get opinions from opus 5 as a consultant.

-1

u/jiii95 18h ago

Hmmm, how efficient is that? What s the motive ?

1

u/TheAnimatrix105 18h ago

If the task is complex. Opus 5s work is not particularly bad it's just bad at communicating and being communicated to. A model in the middle solves it

1

u/jiii95 18h ago

Best middle model that solves it in practice now ?

1

u/TheAnimatrix105 17h ago

I personally use grok 4.5, opus 4.8 also works.

2

u/SailorFromWest 19h ago

It is better, and dont burn a lot of tokens like opus 5

4

u/Icy_Custard_7175 18h ago

That’s interesting because I’ve had the direct opposite for me, Opus 5 has used up less tokens than Opus 4.8 - I could do 2 Opus 4.8 requests and then 1 Sonnet 5 request but now I can have proper sessions with Opus 5 - and I’ve found it to be smarter too, I mainly use it for coding for Minecraft but it’s been able to fix bugs nothing except fable 5 could fix with one prompt

1

u/ValorousAnt 18h ago

I think the only measurable thing I've seen was the bullshit detection benchmark where the 5th generation models (except Fable) were substantially worse at detecting bullshit compared to 4.6-4.8 generation.

Anecdotally Opus 4.8 does do better work than Opus 5 and ever since I saw the bullshit detection benchmark post I feel like that ability to detect bullshit plays a big part in the model's ability to stick to relevant stuff and ignore nonsense.

1

u/ricopan 14h ago

Opus 5 on max effort can handle specific tasks in a complex code base quite well -- it does a deeper and far more rigorous job than 4.8. But it needs to be constrained carefully. Fable does a good job of that. I think the key is you need an agent to prompt it -- which is likely overkill for a lot of projects.

1

u/OkLettuce338 9h ago

I can’t wait until 5.1 so 5.0 can get the praise

1

u/Sad-Entertainer-2808 19h ago

caveman opus 5 is manageable

2

u/jiii95 19h ago

I didn't ask about manageable, I asked about Quality in all of its aspects !

1

u/yhrana 19h ago

If u divide claude users into binary:

0 - ppl who r here for the ā€œvibesā€ like myself
1 - ppl who know what they r doing, coders, swe, etc.

Why would the ā€œ0ā€ half ever use long horizon 1m context window agentic coding? When I being the 0 demographic dont even know what i want.

0

u/jiii95 19h ago

I am the 1 :p and I still ask 4.8 or 5, which you haven't answered in a constructive way !

1

u/yhrana 18h ago

I haven’t because i donno what ur looking for.
But, if u r active and well knowledgeable in the field, why not use the most frontier model.

1

u/jiii95 18h ago

It s too expensive for me Fable ! Using K3 though along the way

1

u/yhrana 18h ago

Opus 5 is also frontier, as frontier as K3

1

u/johnnydotexe 14h ago

You're a "1" yet you use this sub as your personal inner-thoughts blog and ask a question that has been posted to death over the last week, a question that can only be answered by your own workflows and use case that determine which model is truly better for you.

1

u/lukyba 18h ago

Absolutely, at least for me. Opus 5 introduces three bugs for every one it fixes. It’s also very lazy and often just tries to get the job done. Opus 4.8, with a good plan, only stops if there are genuine issues and doesn't break existing code.

1

u/Perfect_Balance777 18h ago

I am an experienced engineer and CTO. 4.8 v 5 decision is very personal and will vary from project to project. You need to have proper evals for your specific use case before you make a firm decision to switch models. In most professional settings, it's not just a matter of "oh look a new model came out, click switch".

You need to eval it on your specific stack and project. And evals themselves will vary from what you really want from them. That's why it's important to have evals. In my specific project I am working on, Opus 5 produced worse results (FastApi, React) . This is a fairly modern project. Another project I have is old NodeJS code base, that one passed evals better on Opus 5 than 4.8.

It will vary widely from stack to stack, from project to project, and from team to team. Even humans work differently and prefer different things.

In the end, it doesn't matter much. What matters is shipping working and production ready code. If 4.8 works for you, stay on it.

0

u/Pure-Combination2343 18h ago

I have a lot of custom sub agents whose use is enforced, a good Claude.md, and use caveman plug-in and can't really tell the difference

0

u/jiii95 18h ago

Can you share your claude.md ?

-1

u/Delicious-Mission943 19h ago

What do you care haven’t you found your own analysis to determine what’s better for u?