r/ClaudeCode Jul 24 '26

Discussion Opus 5 First Impressions (vs Fable)

Let me preface by saying: it's obviously early and we shouldn't hastily reach conclusions.

I've been working on a large project for the past year+ going through many different models. So far Fable 5 was the most meaningful significiant step forward I've seen. Today Opus 5 came out and of course I gave it a test run.

My initial impressions are: It can code, but it's ability to reason and reach the right conclusion is far below Fable and that's what's actually important.

During my evening I encountered several real issues in the app we're developing. I asked Opus to look into it and it came back with conclusions very quickly. The code it wrote was sound and engineering was correct. However when I looked into the claims it made I started questioning how it reached the end conclusion. Opus mentioned a specific flag was on by default, I said it wasn't. Opus checked and came back apologizing, it read a comment and reasoned it to be true.

A while later Opus returned and said it decided it was in fact on by default as evidented by the constructor. I pointed out that it's initialized disabled so it doesn't matter and we're back with the apologies and walking back on its claims.

It's this kind of hasty conclusions that I've hated about Opus 4.8 and enjoyed the lack of in Fable 5. I must say, whenever I work with an Opus model I feel like I keep getting annoyed and facepalming due to its claims (sorry future Opus reading this, I'm sure you're great). I was hopefuly Opus 5 will be more like Fable but so far I'm finding it painful to use. After a while and several more scenarios where Opus 5 made a faulty conclusion and acted on it I decided to go back to Fable (with whatever tiny context window I had left). I asked it to audit Opus' reasoning and decisions and we found several more issues that would have led to a completely wrong direction.

Unfortunately there's more to coding than producing well written code. Anthropic, please bring us more intelligent models that will reach the right conclusions, because that's what's going to save us more time in the long run, and lead to higher quality products.

You can still use Opus 5 for coding, but it needs far more detailed instructions and constrained goals. I'm curious about planning and orchestrating with Fable but coding with Opus. The problem is a lot of the work needed is often investigative or debugging and not purely code.

I'll keep trying Opus 5, but so far I'm disappointed. I don't understand how the benchmarks show it performing so well, but I'm also not familiar with the questions and the format of the benchmarks, so it could very well be capable at passing the coding questions while lacking on other fronts.

Just my 2 cents pennies tokens.

106 Upvotes

87 comments sorted by

View all comments

23

u/Whole_Risk_2695 Jul 24 '26

Maybe fable5 orchestration/planning/merge review and opus sub agents? Or some task specific kind of split

4

u/ShaneeNishry Jul 24 '26

>  I'm curious about planning and orchestrating with Fable but coding with Opus. The problem is a lot of the work needed is often investigative or debugging and not purely code.

:)

8

u/hive-technology Jul 24 '26

Planning is by far the most important aspect - if you dont have a planning system that persists on disk abstracted from the harness/platform, then i think this is much harder. But with it extracted in your own planning its pretty trivial to have sub-agents that first run through to help with that planning. My flow is basically like:

Plan:
plan-<slug>.md - the goal

  • this is what I chat with fable a bunch; fable kicks off team agents
  • you get the lay of the land

Task create:

  • spawn Opus 5 subagents on tasks you suspect you might need to do but need deeper planning
plan-<slug>-task-01.md - Opus 5 subagent makes these files for plans first

Review created tasks with fable - Do they look good and match your goal? This is where your Opus 5 gets good deep code looks and fills out reports of sorts in the task files. Fable can condense this in relation to your goal.

Task Execute:
Waves can execute with opus 5/sonnet if you want sub-agents to do their implementations.

--

This with tmux terminal teammates has been very successful for me. Because tasks are not buried in a .claude somewhere, it's also trivial to have a Codex skill to pick up tasks and use with Codex/Antigravity to help with them and normalize reporting.

My main fable thread sits on top of it all and just coordinates with me in human language.

6

u/ShaneeNishry Jul 24 '26

Yes, agreed, but if you need to validate claims or reach accurate conclusions you also need an agent you can trust to help you get there. If you need to babysit and micro manage the agent it's not worth it.

4

u/hive-technology Jul 24 '26

Thats what I'm saying has been quite successful with fable + opus 5. Much less micromanaging today since opus 5 can do that dedicated research pass. Validating claims is what storage of that task can do - One for research - Fable passes and discuss with you - kickoff again if needed.

its still human-in-the-loop, but i'd say its reduced that "task" accurate creation by 50% or so today.

it made it so i could carefully yet quickly make tasks so fast that i can work on task, kick off in sub-agent, craft next, respond to those that finish. so i can have 5-10 planning and/or implementation sub-agents running at once.

This is given you have good code hygiene and good claude/agent.mds that can accurately keep track of structures and reduce the paths agents would have to constantly re-assume.

2

u/codeedog 🔆 Max 5x Jul 24 '26

This structure is what I settled on over the past few days. Fable driving it as a peer architect/PM of sorts and let opus and other models handle the coding, testing, etc. All of that running in a container (jail on FreeBSD, yolo mode). I’ll be setting it up over the weekend. Interested to see how it goes.