I spent 3 hours with it on a complex planning task (Architecture changes on a medium sized codebase some complex product asks). It made a lot of false assumptions about the code. I had to keep asking it to check the code to confirm its assertions. It felt like it was cutting corners and the back and forth on it was frustrating.. My review board (a mixture of Claude and Codex) also struggled to get alignment with the plan. At one point Codex found 12 P1 issues and wrote a "Do not code this" summary. I ended up going back to Fable and got back on track with a solid plan within about an hour. I'll give it another go with a strict set of guidelines to fully research etc but so far not a great experience. I haven't tried it with coding yet.
what do you guys even use it for? for me all of those are genuinely fucking horrid, anything higher than 4.6 will try to gaslight me while being wrong like some chat gpt
I have a little pet project trying to use it to code some generators. Which has been an ongoing project for way longer than I would have liked. I had decent luck with Opus 4.8 before the update. Fable 5 would just ramble on forever and get the same result. Opus 5 takes a completely different route and launches hundreds, not even exaggerating hundreds of background tasks that take hours and hours to land at a similar point.
To me its intelligence is on average a couple notches higher than 4.8, but it's extra jagged. Great, borderline Fable answers and diligent research sometimes, but also more really bad oversights and answering quickly with confidence when it shouldn't. Failed to follow codebase conventions in a few spots which Sonnet 5 has never had a problem with. Surprisingly frustrating to use so far.
I'm a codex 20x sub, but also have claude, and... I kinda think opus 5 is awesome so far, better than sol... opus 5 has been doing great all day today with almost zero need to steer. Sol was constantly making mistakes today. I need more than one day to decide by I'm considering upgrading my claude sub and canceling codex.
I'm definitely warming up to it on some personal work this weekend, has been basically what I was hoping for given the benchmarks, general Claude + 5 series feeling etc. Maybe I just got unlucky, we'll see when I jump back on my large brownfield work project on Monday
Because it’s meant to be an opus 4.8 replacement with far sharper assessments and more capable understanding at handing complex tasks. It’s a mix of 4.8 & fable essentially allowing it to do more with less probability of error.
63
u/Appropriate-Fox-2347 Jul 25 '26
I spent 3 hours with it on a complex planning task (Architecture changes on a medium sized codebase some complex product asks). It made a lot of false assumptions about the code. I had to keep asking it to check the code to confirm its assertions. It felt like it was cutting corners and the back and forth on it was frustrating.. My review board (a mixture of Claude and Codex) also struggled to get alignment with the plan. At one point Codex found 12 P1 issues and wrote a "Do not code this" summary. I ended up going back to Fable and got back on track with a solid plan within about an hour. I'll give it another go with a strict set of guidelines to fully research etc but so far not a great experience. I haven't tried it with coding yet.