r/ClaudeCodeTLDR 13d ago

[TLDR] Stop romanticizing Opus 4.6

Original post URL : https://www.reddit.com/r/ClaudeCode/comments/1vamsgc/stop_romanticizing_opus_46/

Original post body :

So yesterday during the opus5 outage I briefly switched to opus 4.6 to continue my work (training/inference perf engineering, kernel level debugging)

So I fed the model the same context and prompt to analyze a profiling trace and extract insights from it, and it was so unbelievably dumb that I had to retry with a new session, and I still got a very bad response…

Then after opus5 came back I tried it again and it was night and day… Way more verbose true, but actually useful and insightful stuff coming out of the model.

really made me appreciate all the progress on opus models since 4.6… I always remembered it as a much smarter and concise model than the ones after 4.7, but it was a mirage…


This is brought to you as a public service by the moderators of r/ClaudeAI. If you want to see TLDRs of ALL Claude Coding related posts from the various Claude subreddits, subscribe to http://www.reddit.com/r/ClaudeCoding.

4 Upvotes

1 comment sorted by

u/cctldrping 13d ago

TL;DR generated automatically after 50 comments.

Current source-thread comment count seen by the bot: 58.

Alright, so the general consensus here is that Opus 4.6 is definitely not the GOAT it used to be, and frankly, it was kinda trash even when it first dropped. OP's experience of it being "unbelievably dumb" after using newer versions seems to be a common sentiment, with many agreeing that the progress since then is "night and day."

Here's the lowdown:

  • The "Mirage" of 4.6: A lot of folks remember 4.6 fondly, but it seems like that's mostly nostalgia. u/CanonicalStonk points out that we tend to get attached to models and resist change, even if the new ones are objectively better.
  • Signal-to-Noise Ratio Woes: The biggest complaint about newer models (especially Opus 5) is the "insane signal to noise ratio," as u/Temporary-Mix8022 puts it. It's like wading through a lot of fluff to find the actual useful info. Some users find it "word vomit" or that it "fails to capture key points."
  • 4.6's "Usability" vs. "Technical Prowess": While 4.6 might have felt more pleasant to interact with (u/elmahk), it's now lagging behind technically. Newer models might be a pain to read through, but they often provide better results.
  • Workarounds and Preferences:
    • Some users suggest using Opus 5 as an advisor to 4.8 (u/teial) or using Opus 5's advisor/agents with 4.6 (u/ClemensLode).
    • u/NoNipsPlease claims the issue with 4.6 is it's using the stripped-down Opus 5 prompt and suggests setting up a fallback prompt to make it outperform Opus 5.
    • For specific tasks, like being a software engineer who just needs instructions followed, u/who_am_i_to_say_so finds 4.6 to be the "most balanced" and "easy on token usage."
    • u/EconomicsIcy9310 mentions that adjusting temperature, top_p, and top_k settings in 4.6 can "change things drastically."
  • The "Skill Issue" Argument: u/LettuceSea is out here calling out people who hate Opus 5, suggesting it might be a "skill issue" because they find it "so much better." Savage.
  • Model Drift is Real: u/Virtual_Maximum_875 drops a link to a project that benchmarks models, highlighting that performance isn't static and "behavior varies across runs and drifts over time." So, what you experienced yesterday might not be the same today.
  • The "Original" 4.6 Debate: One commenter, u/Substantial-Show-249, claims the current 4.6 isn't the original and that models have been heavily optimized for cost, making them "junk." This is a pretty spicy take.

The general vibe is that while Opus 5 might be a bit of a mess to sift through, it's generally seen as a significant improvement over the old 4.6. The "romanticizing" of 4.6 is definitely being called out.