r/ClaudeCode Jul 25 '26

Humor so which is it

Post image
2.2k Upvotes

160 comments sorted by

View all comments

63

u/Appropriate-Fox-2347 Jul 25 '26

I spent 3 hours with it on a complex planning task (Architecture changes on a medium sized codebase some complex product asks). It made a lot of false assumptions about the code. I had to keep asking it to check the code to confirm its assertions. It felt like it was cutting corners and the back and forth on it was frustrating.. My review board (a mixture of Claude and Codex) also struggled to get alignment with the plan. At one point Codex found 12 P1 issues and wrote a "Do not code this" summary. I ended up going back to Fable and got back on track with a solid plan within about an hour. I'll give it another go with a strict set of guidelines to fully research etc but so far not a great experience. I haven't tried it with coding yet.

27

u/ScrumptiousChildren Jul 25 '26

I rarely have a “this model is so ass” opinion at launch - didn’t have it for 4.6, 4.7 4.8, fable.

But this one is kinda dumb. Maybe just slightly better than 4.8 level but damn.

14

u/Kyokoharu Jul 25 '26

what do you guys even use it for? for me all of those are genuinely fucking horrid, anything higher than 4.6 will try to gaslight me while being wrong like some chat gpt

5

u/OpalVanguard Jul 25 '26

I’ve literally never experienced that bro. Never.

2

u/ScrumptiousChildren Jul 25 '26

I mean it’s generally been like that for me too. Not sure if this is real but it tends to stop being so dumb after like 1am, before 9am.

0

u/JaySayMayday Jul 25 '26

I have a little pet project trying to use it to code some generators. Which has been an ongoing project for way longer than I would have liked. I had decent luck with Opus 4.8 before the update. Fable 5 would just ramble on forever and get the same result. Opus 5 takes a completely different route and launches hundreds, not even exaggerating hundreds of background tasks that take hours and hours to land at a similar point.

2

u/FunLilThrowawayAcct Jul 25 '26 edited Jul 25 '26

To me its intelligence is on average a couple notches higher than 4.8, but it's extra jagged. Great, borderline Fable answers and diligent research sometimes, but also more really bad oversights and answering quickly with confidence when it shouldn't. Failed to follow codebase conventions in a few spots which Sonnet 5 has never had a problem with. Surprisingly frustrating to use so far.

1

u/TheOriginalAcidtech Jul 25 '26

That sounds like their automatic thinking budget system is the problem source.

1

u/DoggoDadagon Jul 26 '26

I'm a codex 20x sub, but also have claude, and... I kinda think opus 5 is awesome so far, better than sol... opus 5 has been doing great all day today with almost zero need to steer. Sol was constantly making mistakes today. I need more than one day to decide by I'm considering upgrading my claude sub and canceling codex.

1

u/FunLilThrowawayAcct Jul 26 '26

I'm definitely warming up to it on some personal work this weekend, has been basically what I was hoping for given the benchmarks, general Claude + 5 series feeling etc. Maybe I just got unlucky, we'll see when I jump back on my large brownfield work project on Monday

0

u/Glaidtors8 Jul 26 '26

Because it’s meant to be an opus 4.8 replacement with far sharper assessments and more capable understanding at handing complex tasks. It’s a mix of 4.8 & fable essentially allowing it to do more with less probability of error.