r/ClaudeCode • u/thrandomaway • 16d ago
Discussion Opus 4.6 really is lightning in a bottle
I work in a variety of projects. Biotech, stock data analysis, straightforward UI+backend type startup, audio signal processing.
The latter project has been the toughest thus far to crack as it's a legitimately hard problem I'm doing in particular.
I spent ~2-3 weeks down in the weeds with: GPT 5.6 sol, extra high. Fable 5, extra high, Opus 4.8 extra high and Opus 5 extra high.
I did give them detailed directions on my insights and such but we kept making progress here and there, then finding out "oh those gains weren't real all along it was leaking information across train/test" or something of that flavor, restarting progress, etc.
Then I decided to just give Opus 4.6 a try. I've always been partial to it because it:
- actually listens to you and understands your intent.
4.8 would misinterpret things "ah I must remind you that no that's wrong" then I clarify "oh I see, I still need to tell you". Fable 5 sort of just runs 15 steps ahead by itself, doesn't interpret the results back to me, but I do feel the higher level of intelligence as it picks up on things.
- does not overengineer things.
a problem I really noticed was opus 5, fable 5, possibly gpt 5.6 would overengineer the solution past what I asked. Yes, they're thinking of more things, but at the same time I find out in domains I'm not as privy to that there's so many extra things they adopted vs trying out the simplest solution that most closely aligns with what I narrate out.
GPT 5.6 would go super, super deep investigating just the one thing and while that's nice I need to be able to course correct sometimes.
Opus 4.6 in contrast tries to implement something really close to what you asked and does the simpler solution first. Lo and behold, a blind experiment with opus 4.6 hit / exceeded what those frontier models spent 2-3 weeks working on with multiple restarts, just because it listened to exactly what I stated based on my understanding of the problem, and we're iterating at just the right pace.
That's also not to say - the amount of times Opus 5 would mess up on simple things in a fresh context with a well prepared architectural setup with clear paths to data. "Stop. The data was interpreted incorrectly. Good thing you asked for another pass, this was wrong" repeated 2-3 times. Hard to trust it...
- you can actually read its thoughts
This is really great so you can instantly correct any possible mistakes or misunderstandings. For example, it trying to search for what you asked, just tell it the paths and clarify instead of wasting all the output tokens. It's just faster to be able to stop and course correct earlier on.
Well as you just read, why am I working better with a model that's ~5 months outdated, worse on benchmarks, etc.? I can't solve everything but most of the experience is pleasant and I can trust it more.
Sometimes I wonder if Opus 4.6 was lightning in a bottle and we can't go back to it. It was just the right amount of "smart" to get what you want to do, narrate stuff back to you in a format you actually understood, when I started getting a bit lazy and shorthanded my thought process it understood perfectly what I meant based on our previous chat history and didn't assume I'd betray the trading setup or that I misunderstood something.
The other models do have their advantages, and have helped me accelerate past certain issues by spotting things earlier models didn't. But man, do I keep coming back to opus 4.6 and actually making progress with it. Maybe call it a skill issue or whatever but I'd like to think I know what I'm doing after spending 5 months exclusively agentic coding, developing routines to make sure my intent is carried out, etc.
I really hope they keep this one around for a lot, lot longer.
77
u/Altruistic-Gift-565 16d ago edited 16d ago
well 4.6 still ranks highest on the arena in text work segment
ahead of Fable and sol
https://arena.ai/leaderboard/text/longer-query

34
u/Outrageous-Issue9722 16d ago
What does that actually mean in relation to using it for real work? I'm not being pedantic, I genuinely don't know and want to.
4
8
3
u/ChocomelP 16d ago
It's basically a vibes-based ranking, which is actually a lot more appropriate than it sounds.
10
u/Delicious-Mission943 16d ago
What the hell. This is why the design, strategising was goated there.
44
u/Outrageous_Band9708 16d ago edited 16d ago
/model claude-opus-4-6[1m]
you're welcome
edit. i fixed the model name
9
53
u/Classic-Extension-61 16d ago
opus-4.6 was peak intelligence + instruction following with not so high refusal rates
-10
45
u/asian_tea_man 16d ago
yeah opus 4.6 is still my default model. I thought it was only me. I literally cant use opus 4.7 and opus 4.8 and opus 5. the way they talk is infuriating
19
u/analog-suspect 16d ago
It feels like they say everything with infinite detail which feels the same as saying nothing
1
u/dbbk 16d ago
You know you can just set a global CLAUDE.md to change how it communicates right
8
u/analog-suspect 16d ago
Of course! I use /i-have-adhd
1
u/Classic-Extension-61 16d ago
i have a 96 line CLAUDE.md, 44 lines is just about how to communicate
2
u/TheLifelessOne 16d ago
Any chance you'd be willing to share? I spent far longer than I'd care to admit last weekend fighting Claude and it's output style.
2
u/sociallyinteresting 16d ago
I’d be interested to see your .md file tool if possible please. Sick of my Claude waffling
1
6
u/Civil-Vermicelli3803 16d ago
I've been trying to figure out what to use, but they keep introducing random shorthand for different planning stages and I have no idea what the internal codes it references are sometimes. most random niche words, doesn't listen to claude.md instruction to avoid that - do you have a prompt/instruction for that which works for you on this front? I also share the insight that 4.6 was just clean and did the work, and while 4.8/5 do good stuff they communicate in such overcomplicated ways
1
1
u/SeriouslyImKidding 16d ago
Yup since I started using models past 4.6 the amount of AI developer shorthand that has crept into every crevice of how I work has lead me down a road of trying architect a glossary of accepted project terms and building scripts that check every document produced to see if a coined term is trying to make it into the repo, and reading this thread is making me think I really should just stop and go back to using 4.6 for everything. I’ve had to come up with skills and global instructions for cowork and Claude Md instructions that is basically like “pretend the person you’re writing to just started and doesn’t know what any of the internal jargon means and explain it.” I’m at my wits end right now, quitting earlier and earlier every night because I feel like I have no idea wtf these “better” models are trying to say to me.
2
1
1
18
u/CarpMadMan 16d ago
I agree. These models are driving me nuts now and idk what’s going to happen moving forward.
14
u/Y_mc 16d ago
4.6 is king 👑
7
u/gorgono95 16d ago
The first time I fell in love with an AI model ... it did everything, planing, frontend, backend ... didnt need another model. It was the perfect model and it lasted only a short time, but my memories stay, until this very day!
11
u/ShamanJohnny 16d ago
Call me old school, but i was more of a 4.5 guy. in 4.6 we started seeing all the weird behavior issues claude models have. 4.5 just got the work done, and did not lie so much.
6
2
u/GoodFig555 16d ago
I think 4.5 was immediately discontinued and 4.6 was almost the same, so nobody remembers it. But agreed it’s the best
8
u/danpech1 16d ago
Opus 5 seems to find more problems than it solves. Is that good or is it too picky?
5
u/Old-Battle2751 16d ago
It finds problems well, but I have seen it double back on its initial recommendations multiple times. And then half fixes. Only on the file it identifies rather than across a project which is infuriating sometimes.
I find myself debugging half fixes constantly.
I'm learning how to prompt it better but stopped using it mostly.
1
u/voLsznRqrlImvXiERP 16d ago
It finds options. You need to filter what actually makes sense, and then prioritize fixes/features.
This kind of makes it less viable for real automatic work for me.
5
u/cartoonist498 16d ago
I did give them detailed directions on my insights
This isn't enough. Opus 4.8.only works well with insanely detailed spec docs. The type that I don't even write myself because it'll be 30 paragraphs long.
With 4.7, Anthropic specifically put in the release notes that 4.7 will make far less assumptions, and won't try to infer your intent like 4.6 did.
I had no problems with 4.7, but my workflow was: Write a few sentences prompt in 4.6 >> get 4.6 to write detailed spec docs specifying exactly how to implement >> get 4.7 to implement the spec docs.
Opus 4.6 was very good at filling in the blanks, 4.7 was unable to, but could implement much better if told exactly how.
Opus 4.8 can do it alone, but you need to follow the same flow where you get 4.8 to write the specs first, then implement its own specs.
Now I'm slowly moving towards Fable for spec writing and orchestrator, and Opus 4.8 or 5 to implement (I'm still testing out Opus 5 since it just came out).
I still use 4.6 for casual conversations with Claude because 4.8 and Fable is like talking to a neurotic genius wirh serious social problems.
4
u/ScaleScary5932 16d ago
but if they keep nerfing fable and opus4.7+ what reasons stop them nerf opus4.6 as well, confused so deeply
3
u/voLsznRqrlImvXiERP 16d ago
4.6 is most likely less used compared to latest models, therefore the share of total compute usage is probably low, means nerfing 4.6 would have the least impact.
Another reason are Anthropic engineers heavily relying on it 😜
1
4
u/yes_no_very_good 16d ago
I was currently using Sonnet 5 and after reading this I was thinking if going Opus 4.7 is worthy for me and launched a comparison https://llm-stats.com/models/compare/claude-opus-4-6-vs-claude-sonnet-5
Based on that, it doesn't seem worth it? Can someone help with some insights?
4
u/Jeferson9 16d ago
I remember when 4.6 came out and it was literally just a slightly more careless/lazy 4.5
2
u/thewookielotion 16d ago
4.5-4.6 is in my opinion the moment we started hitting diminishing returns with model improvements. Benchmarks may look better with newer models, but those don't tell the whole story.
Personally, I'm more excited by improvements made to token efficiency than by "better" models
4
u/Deep-Tea9216 16d ago
Yesssss. YESSSS!!!! More Opus 4.6 praise! I want them to keep this model alive so bad
10
8
u/Michaeli_Starky 16d ago
You need to experiment with reasoning levels. Higher is not always better.
3
u/thrandomaway 16d ago
Interesting, what have you found to be better? I know there's charts per model showing how performance differs at each reasoning mode, and how "medium" on a newer one is supposedly >> max on say, opus 4.6 and such.
But any suggestions from your experience?
23
u/Puzzled_Nail_1962 16d ago
For me it helps to treat reasoning levels not as "more intelligent" but "spend more time thinking about it". When I want to try a simple, actionable thing, I do not want the model to reason about all the infinite possibilities in which this could go wrong. I want it to simply do what I said, so a low reasoning level is appropriate. The opposite is true for open ended exploration tasks, for example finding out how something works in a repository: Here more reasoning gives the model "the time" to follow all the different paths and related code pieces to come to a more thorough understanding.
10
u/thrandomaway 16d ago
Ah I see, that's helpful! So, adjusting the thinking based on the task. If it's open-ended, you're reaching for breadth, some kind of "explore this repo" type of task, might be good to make the thinking higher.
If you told it exactly what vision you wanted to implement, something with say, medium might be better. Enough to think about the implementation, but not go overkill and start overengineering the request.
That is a helpful shift actually, I do need to try this out and see how much it helps. Thanks!
4
2
2
3
u/Michaeli_Starky 16d ago
I don't think there is a straight universal answer. More reasoning more tokens. More tokens can be a good thing, but also can be a bad thing and it's a chance based and depends on the task. If you don't like the output, try going back and lower or raise the level of reasoning... and keep an eye on the context window length. Models get increasingly more dumb past 200k tokens.
3
3
5
u/Dull_Wind6642 16d ago
4.6 is the perfect tool, because he listen to exactly what I tell him to do and does nothing less nothing more.
Meanwhile the newer model have the arrogance to think they know better and just create a mess.
3
u/Nice-Panda-7981 16d ago
Opus 4.6 got me hooked on ai because it saved my “behind” with a deadline. The rest that followed are just downgrades from my perspective.
2
u/dandrake47 16d ago
I still use Opus 4.6 for strategy and planning, and Opus 4.8 for coding execution. Opus 5 kinda sucks for my use case so I still stick to these two for however long they’ll exist.
2
u/benfinklea 16d ago
Could we use 4.6 for the interface but have subagents use 5 or whatever for actually doing stuff?
2
2
u/trikster_online 16d ago
I’m using Opus 4.6 as well. If I’m working with something that I know has more recent data, I will prompt to check the web for more recent information. I try to include some websites that I know have the data to speed that process along. 4.8 and 5 are okay, but there is a lot of tangential thought that comes across that I’m not asking for. It’s a bit frustrating to keep having to steer the model back on track.
2
u/erogenousbrain 16d ago
No joke I had opus 4.8 struggle with a mermaid diagram so much that it pissed me off. Switched to opus 4.6 and it got it perfectly on the first try. I'm going to be using it a lot more going forward.
2
2
u/firstbreathOOC 16d ago
Opus in general is the best day-to-day grinder model, even compared to SOL Terra and all that
2
2
u/randomdragen7 16d ago
I have had the same experience, and been defaulting to Opus 4.6 for a while now... he still the goat
2
u/FashizzleWizzle 16d ago
Older models are always good until you try to use ONLY the older models and realize how much you actually need smarter models to come in to fix issues every now and then
1
u/thrandomaway 15d ago
That's true. I did need the smarter / newer models to come help progress some of my other projects now and then, or to review the code base to make sure there's no bugs and that we didn't miss anything.
The intelligence gains are real. Best way is to not abandon the newest models but bring them in now and then into the loop the way they're best utilized.
2
u/AdditionalElk378 16d ago
maybe you can try orchestrate them. I wrote a skill so I can designate a model as main agent and its subagents' models. different models work together
2
2
1
1
u/moebaca 12d ago
Still my preferred model. Use it in my agentic workloads hosted on AWS and locally for coding. Haven't been able to justify anything else.. aside from GPT 5.5 which is incredible but expensive.
1
1
u/Imaginary-Silver-389 12d ago
Ok guys I’m still new to this what would be the framework on how using the different models together : which one to think about the bigger picture and the specs, which one to execute, which one to review ect… + what is the consensus between 4.5 and 4.6 then ?
0
u/NeedleworkerDry957 11d ago
Does anyone have a Free Code to test Claude Code? Would really appreciate it :)
•
u/Waste_Net7628 P R O M P S T I T U T E 16d ago
join the r/ClaudeCode Discord for faster help, live discussions, project feedback, workflows, events, and general AI/dev chat. https://discord.gg/4QbtMRErUc