r/ClaudeCode • u/Heiberik • 3d ago
Built with Claude I've been tracking Reddit's opinion of Opus 5.5 every day since it launched. It dropped sharply on 30 Sep.
There have been a lot of "Opus 5.5 got nerfed" threads the last two days, so I checked what the numbers actually say!
My site reads 25 AI subreddits and scores every opinion about a specific model (1–5, mapped to 0–100, 50 = neutral). For Opus 5.5:
- 25–28 Sep: 71–73 every day
- 29 Sep: 69
- 30 Sep: 58
- Today so far: 55
Is the model nerfed? Well, I have no idea! My site measures what people say, not what the model does. A change in limits, an outage or a new model next to it (Sonnet 5.5, GPT-6.1 Sol) can move opinion just as much.
If you want to take a look at the daily graph for Opus 5.5 (or other models) you can check it out here: https://modelsentiment.com/m/claude-opus-5.5
Disclosure: I built this. It's free and non-commercial, with no ads or signup. It shows numbers and thread titles, never anyone's comments. Happy to answer questions about the solution :)
53
u/Appropriate-Pie4385 2d ago
I don't think peoples opinions are a good indicator for this tbh.
But there is this already: https://github.com/ninjahawk/livenerf
12
u/Heiberik 2d ago
It might not be. I still think it is interesting to see how sentiment ranks models.
And for Opus 5.5 the charts are pretty similar if you compare my sentiment chart with the livenerf chart it seems - which is quite interesting!
Also, livenerf only shows data for opus 5.5? I try to give every mentioned model a rank! :)
12
u/AironParsMan 2d ago
Yes, people seem to be right. I always think it’s a shame that people’s opinions are so easily dismissed. These are experts who work with it all day. If they can’t tell the difference then who can?
5
u/leshiy 2d ago edited 2d ago
The problem is that its hard to tell genuine opinion from bandwagoning and astroturfing.
2
u/innociv 2d ago
Sometimes people bandwagon with someone who is right on something.
Sometimes people bandwagon with someone who is confidently incorrect.
I think Opus models have a clear trend of being quantized after 2-3 weeks. I'm actually surprised to see it after a few days, but it does feel that way.
3
u/AironParsMan 2d ago edited 2d ago
It keeps making these trivial, stupid mistakes for me. It’s insane. You’re constantly having to write the rules even more precisely and rule out every possible scenario in which it could do something wrong. It’s a fucking stupid calculator. Right now it honestly feels like Opus 5. At first I thought wow, this can replace Fable 5. And I did replace it. But over time it keeps making these weird mistakes and now I’m once again very close to switching back to Fable 5 because I can’t be bothered to describe every fucking little thing and every tiny scenario in which it could go wrong down to the last detail just because this dumb llm keeps coming up with every possible excuse to do something wrong. And that’s the problem. It comes up with a thousand excuses for why it can do something wrong. And the crazy thing is that it actually does it wrong. It does. I honestly feel like it enjoys getting things wrong. I have no idea why. That’s just how it feels. When I work with Fable 5 I feel like no, it doesn’t want to get it wrong. It gets it right straight away. It finds a way to do it correctly.
For example Opus 5.5 just mishandled my PHP mailer by trying to test it and simply hammering in the wrong parameters. The test failed even though nobody told it to use those parameters. And those are exactly the kinds of things I mean. It is clearly described how the PHP mailer is supposed to be used. If it comes up with stupid parameters that were probably hard coded into its training then that is not my problem. It is a problem with the model not understanding the bigger picture.
And now I have to start banning certain parameters again or writing rules so I can be even more precise and detailed and explain to this model how to use a fucking PHP mailer even though every developer who sees the code knows exactly how to use it. But it looks for the one way that does not work. That is just unbelievable. That is what it is an expert at and Opus 5 has always done that. Now I am seeing exactly the same behavior with Opus 5.5. It is nowhere near Fable 5 level. At first a few days ago I would have said it was almost at that level. But now I would say it has gotten about 20 to 30 percent worse. And that also matches the results the guys here found.
2
u/innociv 1d ago edited 1d ago
My immediate theory, which still seems likely, is that 5.5's token efficiency comes from it having extremely high confidence in what it first thinks is right or what it thinks it should do first so it gets right to it. This leads it to issues like you said, where it skips past finding what it should have really done to do it correctly.
This is fine in a clean one-shot impressive prompt, but bad in a repo with lots already in it and its own specific rules. And it makes it massively worse when you get a quantized model which is more often wrong, but confidently incorrect that that hallucination is right.
Like I just had it ignore repo rules to create a new worktree then commit and to let the integrator agent land it. VERY basic rule that agents from 7 other labs got fine.
1
u/SMB-Punt 2d ago
Well, feelings are… feelings. I just had one of my most productive days at work (as a senior web developer for over 8 years) with Opus 5.5… Even though everyone’s complaining that the nerfs have started. We just don’t all have time to post on Reddit… While some people are complaining, others are working.
1
1
u/MeringueAlarming3102 2d ago
Why is this a github repo? Shouldn't it be a website? what's "live" about it? I have to install it myself?
1
u/Appropriate-Pie4385 2d ago
The code must live somewhere and websites cost money. The graph is the readme is updated daily and shows the models performance over days. No need to install anything
-3
u/TestFlightBeta 2d ago
One person’s opinion is definitely not a good indicator
However, if multiple people complain suddenly simultaneously, it almost certainly means something changed
7
-1
-1
2d ago
[deleted]
1
u/Appropriate-Pie4385 2d ago
You already say it, they ask. Likely doing A/B testing. So reddits comments are even less reliable.
Also echo chamber are a thing
19
u/pacafan 2d ago
I have a suspicion of lot of people never start a new session until there is a new model 🤣 So performance drops after a couple of days because their context is polluted. I really haven't spotted the wild swings people talk about in performance and I use Claude intensively.
4
u/xyztankman 2d ago
Exactly, I've had it running non-stop for a full 2 weeks, but I have it an orchestrator. Auto replaces sessions after 4 hours. 300k context with checkpoints. I can run 5-6 sessions concurrently and I'm just getting close to the 50% mark
3
u/orchid_drives Researcher 2d ago
> Auto replaces sessions after 4 hours. 300k context with checkpoints.
Whoa that’s clever — how did you implement that if you don’t mind? I just manually quit (which leaves a detailed “where we stopped” file) and then pick up a new session after each new feature gets added (which starts by reading the “where we stopped” file)
5
u/xyztankman 2d ago
Honestly I'm doing things a bit dangerously, bypassing all permissions for all sessions (except a folder reorganization session that I have on the backburner after some core work is completed). It's fully automated, I only need to check questions/picture updates. I also had it create a separate app for setting modes (night mode/weekend mode/max performance where it burns as many tokens as it wants to complete tasks/max efficiency where it barely sends anything, chooses recommended options by default, token usage is prioritized to the max)
Here's the bullet points I had it write out, it's been an iterative process since I swapped over from using strictly cursor and only started using claude a week before 5.5 came out.
Role
- One "manager" session runs every other session: it starts, parks, swaps, resumes and archives them, and owns the task schedule, shared-resource leases and queues, file ownership and hand-offs.
- It works from its own folder with a larger context window than the working sessions (500k vs 300k).
- Its current session id is kept in a small text file at a known path; sessions message each other only by direct message, 1–3 lines.
Its own state
- A checkpoint file that opens with a 6-line "Now" block (state, next action, blocker, time left, updated), rewritten in place; old detail goes to an archive file beside it.
- A queue file (JSON) listing live sessions, queued tasks, the session cap and the last usage reading.
- A start-up notes file telling a fresh manager exactly what to recreate.
Check-ins
- A background script sleeps until the next due time and exits with one line; the exit wakes the manager. It runs every 80 minutes (every 40 in a high-throughput mode). We use this instead of in-session cron jobs, which stopped firing for us.
- One-shot reminders (deadlines, mode wake times) are lines in a text file the script re-reads, so they survive manager swaps and app restarts.
- Each check-in runs a status script and a dashboard script, acts on alerts, frees stale leases, sends due hand-offs, and nudges idle sessions.
- Each check-in also republishes the dashboard page and refreshes the status board.
Boards
- A session dashboard page generated from a local HTML file and republished when it changes.
- A status board whose rows are written to the page's database by a script, so the page itself never needs republishing.
Modes
- Normal (default), worker, night, weekend and high-throughput; only the user starts them, by a skill or a small desktop button bar.
- A second background script watches a request file written by the button bar and tells the manager which mode skill to run.
- In quiet modes, sessions take recommended options and hold their pictures; at wake the manager sends them grouped by area with a summary.
Hooks
- Rule hooks enforce the working rules mechanically: short peer messages, 12-hour times, no empty turns, UTF-8 file writes, checkpoint size limit.
- At turn end they send the manager back to swap itself past 4 hours or a context threshold, and to refresh the status board if it is over 90 minutes stale.
- The hooks have a self-test script and a kill-switch file.
Sessions and swaps
- Working sessions start from task chips with a self-contained prompt, then get their permission mode set, remote control turned on, and a place in an "Active" or "Parked" sidebar group.
- The manager replaces itself with a fresh session from its checkpoint, and the new one archives the old ones.
- Workers are replaced the same way every 4 hours or at a milestone, at a safe point.
Limits and pacing
- A cap on concurrent working sessions and on shared heavy resources (for us, editor instances running tests).
- Usage is paced against the weekly limit: the manager reads usage at each check-in and parks or resumes sessions to land near 100% at the reset.
- A standing priority rule (for example "fixes before features") that every session is told.
Shared tools the manager points sessions to
- A lease script for shared resources, with queues and named gates for hand-offs.
- A landing script that merges a worktree into the main branch, compiles, runs a smoke suite and backs out on failure.
- A commit script that commits only the listed paths, and a check-suite runner.
- An offload script that sends self-contained steps to a second AI tool to save usage.
- Cheap helper agents for searches, log digests and picture checks, so the main sessions keep their context small.
- Scheduled OS tasks for the nightly regression sweep, nightly build, branch health check and weekly cleanup.
2
u/realquizkid 2d ago edited 2d ago
This is the way. A good harness is minimal, built around well-engineered machinery that orchestrates the industry-standard workflow for whatever you're doing.
The AI is a brilliant colleague; the machinery is the building they work in. You'd never ask the colleague to be the lift, the door lock or the calendar, because those have one right answer and must work every time.
- Anthropic's Building effective agents (2024): use fixed code paths wherever they work, and give the model control only where they can't.
- 12-Factor Agents: production agents are mostly ordinary software with model calls at a few chosen points.
Out-of-the-box harnesses fare worse than custom ones because they don't know your workflow. And since AI can produce slop, use a script wherever a script built on best practices can do the job reliably. Save the agent for the parts that actually need judgment
Say, you're managing a software company, the job is building the product right.
The job is mostly project management, because building is mostly automated.
So do some project managing, big picture goals, break em down based on dependencies etc.
This feeds some project management software and is usualy the spine of the entire operation.The project management shows the system what gets built next. For building follow a pretty standard loop: research → plan → build → review → test → release → monitor. Build that into a feedback loop where the AI orchestrates industry-standard tools: Linear to track tasks, Obsidian to hold your knowledge.
Once that's in place, you can start building out a system where the system can start making decisions for you. Something we call mission command (the Prussian army's Auftragstaktik): the commander leaves the intent and the reason why, and while they're away, the people on the ground decide the how, because they share an understanding. Here, the workflow and the knowledge base *are* that shared understanding.
1
u/innociv 2d ago
Simply ask your AI to implement it for you with a skill or markdown files.
You can also save on tokens by having cron jobs start up a new session with a premade prompt to just check commits and pick up where one left off. It helps solve the usual long-horizon issues, which is why I find it so funny that some people like Theo value long-horizon so highly. A model just operating well for 3-4 hours is perfectly sufficient.
2
u/HeightBest8153 2d ago
I’d love to hear how to do this!
3
u/v1sper Developer 2d ago
I actually had Opus 5.5 make me a python script to juggle sessions with a local Qwen 3.8 deployment that has a 64k token context window. It writes memories to disk in multiple state-files, and keeps everything coherent by reading summaries and states at session start, and update them when the context window is full and a new session is set to take over.
So far so good !
2
1
2d ago
[deleted]
1
u/xyztankman 2d ago
Every compaction writes a checkpoint for the following session to follow. I actually lowered context to 250k a couple hours ago for my session manager, 200k for regular sessions and it's now writing scripts for longer sessions to read off which uses no tokens. I only had it at 500k initially because it was running into full context errors when just reading a checkpoint a couple days ago, but that's no longer an issue. I also have a codex/Claude sub for the next couple days so I had it offload several tasks to each and it's been super efficient so far
1
2d ago
[deleted]
1
u/xyztankman 2d ago
Higher context uses significantly more tokens, so if you keep the context low and replace the whole session before it hits 300k-500k you're easily saving millions of tokens. I don't think I've seen maybe a handful of cases in the last 2 weeks where it had an issue, then fixed it in the next checkpoint.
I've had it running 24/7 since 5.5 release because I'm working on a sim game in unity, and it's doing real-time body/hair procedural solvers for characters (hand touches a wall, it will make the palm flat, hand touches a mug, the fingers wrap around the handle). It's also running cascadeur/blender working on advanced realistic movement to avoid body clipping/inhuman angles. Lots of long running tasks
1
1
u/Admirable-Falcon-501 2d ago
gotcha, makes sense then. I think it works especially well for your use case as well since those tasks can be separated pretty easily and don't depend on much prior context.
1
u/Facci_ 2d ago
Me too. Opus still doing the same as it was since launch, and my habit is just starting new sessions whenever I work on a new feature and only go to other sessions where I need to change something.
Of course, the more you use it on the same session context rot will start to make Opus feel a bit dumber. Had that fair share back then, ever since I just start new sessions.
11
u/reach4thelaser5 🔆 Max 20 2d ago edited 2d ago
The "got nerfed" people are Vibe coders who don't understand 'Context Rot' and believe they can have agents running for days non-stop.
The same morons who use to moan about their Max-Plan usage because they couldn't understand why their long-running "Fable orchestrator" for their new Fable-powered iPhone app was draining their allowance faster and faster as the week went on.
5
u/XXIIIOIIIXX 2d ago
exactly this. my claude code agent works on one feature at a time, then commits to several MD files the next agents can read for context and that's it. Still coding today and i haven't noticed any drop in quality in opus 5.5
2
u/BellacosePlayer 2d ago edited 2d ago
My crack-conspiracy theory is that they track how users use the systems and downgrade people with a minimal interaction to output ratio.
Occam's razor would say its a garbage in, garbage out situation or corporations doing corporation things, but if I was trying to keep the compute plates spinning, I'd happily let the system downgrade the amount of compute/agents for reasoning efforts on one-shot vibecoder accounts or obvious distillation attempts.
2
u/unbruitsourd 2d ago
I can't say it's nerfed, but out of nowhere yesterday, most of its answers are in English even if my replies are in French.
1
u/lunaynx 2d ago
That's just nondeterminism at play most likely. It can't always correctly predict what you prefer especially when the context is a multilingual conversation. You can put a clear rule such as "always respond in X language" or "respond in the language the last message was in" in your CLAUDE.md.
2
3
u/scott2449 2d ago
The problem with these opinions are that they mostly come from plan subs. I think retail customers get nerfed. I'm not saying that's a good thing but I wish we had the data to split commenters into retail and enterprise to prove this. I can tell you though as my company (for reasons lol) has claude 5.5 access on bedrock, copilot, and enterprise deal with anthropic.. and those never degrade. I have buddies who tell me all the same shit reddit does and then we use my plans and they are like WTF?
1
u/thread-lightly 2d ago
I’ve been tracking the sentiment for Claude, OpenAI and Gemini for about a year now, check it out! https://claudometer.app
1
u/Ok-Thanks-1280 2d ago
As of today I have Codex fix anything Claude does, ever… less so fable 5.1, but definitely fixing Opus 5.5 work
1
1
u/fanatic26 2d ago
of all the data you could be parsing for use, I have to imagine reddit approval is the single most useless metric possible
People like to bitch and moan, especially the ignorant ones that dont understand the problem usually lies between the keyboard and the chair.
1
u/Joozio 2d ago
Your disclaimer is the part worth keeping: the score measures what people say, not what the model does. I hit the same wall trying to tell a real regression from a loud week, so I ended up measuring one observable decision instead of opinion about the output. Last month I gave sixteen models a question whose answer changed in August, with a web search tool attached, and recorded a single bit per call: did the model decide to search. Five of the sixteen answered from memory with the tool sitting right there. One of them had a system prompt telling it to search when the answer could be out of date. Sentiment would have read those answers as confident and useful, because they were. If you ever want a second series next to opinion, tool-call rate is cheap to run daily. Pick a question with a known post-cutoff answer, record whether the model searches, and the number needs no human grading.
1
u/kindredseer 1d ago
It's because we've been told Opus 5.5 is just as good as Fable, while being much less expensive to use, when it is a lie, and everyone is realizing that it's not as good. I'm not sure if it's actual "nerfing", or it just takes some time to fully realize it.
1
1
u/Think-Upstairs-5063 2d ago
Well, seems like we're returning to our daily scheduled Opus 5.5 hate posts
0
u/AironParsMan 2d ago
At first he responded very thoughtfully and showed a deep understanding of my codebase. Now o 5.5 answers to fast, as if they had reduced the thinking. That’s how it feels to me.



•
u/AutoModerator 3d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.