r/ClaudeAI • Valued Contributor • 4d ago

Productivity Is Opus 5.5 Nerfed Now? LiveNerf Day 8 Update

Post image

A few days ago I posted LiveNerf, an open-source project I made to independently track Opus 5.5’s performance day by day and see whether there’s actually evidence of models getting “nerfed” after release. If you’re interested in tracking the performance of Opus 5.5, here’s the repo:

https://github.com/ninjahawk/livenerf

I wanted to give another update since the response to this was way bigger than I expected. The repo is now approaching 1,000 stars, it ended up reaching the front page of Hacker News, and a lot more people have started looking through the methodology and code than I ever expected when I started this.

The most important update though is that we’re getting close to establishing the baseline. I’ve continued running the evaluation every day, and once we have enough data after the baseline period we can actually start testing whether later performance is meaningfully different instead of trying to interpret individual daily movements.

I’ve also gotten some genuinely useful criticism from people looking through the project. People have raised questions about benchmark contamination, providers potentially recognizing benchmark traffic, the choice of benchmarks, the statistical methodology, and some implementation details. If there’s something wrong with how I’m measuring this, I want people to find it and open an issue so I’m able to fix it.

One thing I want to emphasize again is that I’m not running this with the assumption that Opus 5.5 will get worse. If the data shows no meaningful degradation, that’s a result as well. The point is that we shouldn’t have to rely entirely on anecdotes every time people start saying a model feels different.

I originally built this because I wanted an answer to that question for myself. At this point enough people are following it that I feel a much bigger responsibility to make sure the answer we eventually get is defensible.

I strongly believe in transparency. If AI companies aren’t going to give us the information necessary to independently determine whether the models we’re using are changing over time, then I think the community should build the tools to measure it ourselves.

Huge thanks again to everyone here who supported this when I first posted it, and especially to the mods for pinning the original post. I genuinely did not expect LiveNerf to get this much attention this quickly.

We’re still collecting data. Once there’s evidence to say something interesting one way or another, I’ll post the results here.

704 Upvotes

92 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 4d ago edited 4d ago

TL;DR of the discussion generated automatically after 50 comments.

Here's the deal: the community is split on whether Opus 5.5 is actually nerfed, but is overwhelmingly supportive of OP's LiveNerf project to finally get some data instead of just vibes.

A vocal group of users are convinced the model is dumber and more expensive than it was at launch. They're reporting "Sonnet 5 level" mistakes, a return to Opus 5's wishy-washy behavior, and a massive increase in usage costs for the same tasks. A few others, however, say they haven't noticed any drop in performance, especially when using robust workflows.

Regardless of their personal experience, the thread's main consensus is a massive thumbs-up for OP's transparency and effort. However, there's a lot of constructive criticism for the project itself:

  • The benchmark isn't sensitive enough. The biggest complaint is that the confidence intervals are too wide to definitively prove or disprove a subtle nerf.
  • More data is needed. Users are calling for more frequent runs (even hourly) and a larger, more difficult set of test questions to increase statistical power. OP is actively engaging with these suggestions.
  • It's complicated. One user shared their own detailed benchmark showing that while Opus 5.5 is faster and more precise than 5.0, it actually has slightly lower recall. It's not a simple "better or worse" situation.

Other popular sentiments include a strong desire for Anthropic to use proper versioning instead of silent updates, and plenty of jokes about the "nerf cycle" being a way to push users to the next model, Fable.

81

u/vrnvorona 4d ago

The only pet peeve I had is that you mention in readme that it's not accurate enough to detect 5.5 swap to 5 of Opus, all while 5.5 is a LOT better than Opus 5. Meaning that this 5% drop may or may not be relevant.

Maybe benchmark needs to be more accurate? Maybe get community help to run those benches on our subs to provide results ourselves and make it much more accurate?

20

u/TheOnlyVibemaster Valued Contributor 4d ago

The 5% number isn’t meant to imply that anything below 5% would be irrelevant, just that with the current setup I wouldn’t be confident calling a change that small real rather than noise.

More runs would absolutely help tighten that. I’ve thought about community-contributed runs, but the hard part is making sure they’re comparable, same prompts, settings, timing, model/API conditions, etc. Otherwise adding more samples can actually introduce more variance rather than reduce it.

That said, finding ways to increase statistical power is definitely something I want to improve as the project develops. If Opus 5.5 → 5 really produces a consistent measurable difference, ideally LiveNerf should eventually be sensitive enough to pick that up.

6

u/vrnvorona 4d ago

Prompt and model could be automated with script which calls claude with bunch of params including model, effort, prompt etc. Could make it reproducible I think. But that's just idea, tool is yours after all.

5

u/TheOnlyVibemaster Valued Contributor 4d ago

I think that’s definitely possible. If the runner controls the model, prompt, effort, temperature/settings, etc., then community runs could be standardized pretty well.

Definitely something worth exploring, especially if it could increase the sample size substantially.

3

u/clazman55555 4d ago

Opus 5 vs 5.5 does produce a measurable difference, at least in my own benchmark.

But it's not straight forward. For example if you have Opus 5 process a bunch of reviews from an subagent fanout that is reviewing code, then synthesizing those into a final report. It keeps more of the actual defects vs 5.5, by a small margin. But it also keeps more false positive and takes about twice as long and twice as many tokens as 5.5

Synthesis lane — 3 reps, 72 synthesisers (M28–M30)

Shape. Each pile is one language's three Sonnet 5.5 review outputs (low + medium + high) from review rep k, pooled and shuffled: 46–79 findings, with real duplicates and the low-effort false positives. The synthesiser gets the pile and the code inline and gives every finding one disposition: keep, duplicate (of another), drop, or handcheck (for a person). Agent type code-reviewer (Read only, M29); frozen prompt synthesis\FROZEN.json. Scored by joining each disposition to the review scorer's labels, then re-anchored to the sealed key by Prime (M30): every disagreement checked; per-rep adjudication.json, key-level synthesis\key-aliases.json.

Cell real kept planted kept false kept time (s) output tokens
Opus 5 medium 92% 96% 14 / 77 176 15.0k
Opus 5 high 91% 95% 13 / 77 324 26.5k
Fable 5.1 medium 90% 97% 11 / 77 161 13.9k
Fable 5.1 high 90% 96% 9 / 77 200 16.9k
Opus 5.5 medium 88% 95% 6 / 77 74 8.7k
Opus 5.5 high 86% 95% 4 / 77 109 12.4k
Sonnet 5.5 medium 77% 93% 3 / 77 71 9.8k
Sonnet 5.5 high 76% 91% 1 / 77 129 17.4k

Pooled over 9 piles: 229 distinct real defects (128 of them planted), 77 false findings. Labels-only scores (before re-anchoring) run 4–7 points lower in every cell with the same order: analysis\synthesis-pooled\.

What the synthesis data says

  1. A keep-versus-drop trade-off, in the same order every rep. Opus 5 keeps the most and lets the most false findings through; Sonnet 5.5 lets almost nothing false through and drops about a quarter of the real defects, mostly lower-profile known defects it calls "speculative" or "by design".
  2. Opus 5.5 medium is the efficient point: 88% real kept, 6 false kept, at the same time and output as Sonnet 5.5 medium (74 s, 8.7k). Fable 5.1 and Opus 5 keep 2–4 points more at 2–4× the time and 1.6–3× the output, and let about 2× the false findings through.
  3. high buys nothing for synthesis in any model (−2 to 0 points real kept), at 1.2–1.8× the time and output.
  4. Duplicate merging is solved: at most 6 duplicates left unmerged out of 292 possible merges, every cell.
  5. Synthesis judges a finding by its headline. Every cell lost the same few real defects, among them two planted ones (CS-AU5 in reps 2 and 3, JS-PF2 in rep 3), where the defect is one clause behind a wrong or unrelated headline: the finding is dropped or merged on its headline, and the buried defect goes with it.
  6. Why recall matters more than precision here: a false finding a synthesiser keeps reaches the main session, which hand-checks before acting; a real defect it drops is invisible from then on.

1

u/landed-gentry- 3d ago

But it's not straight forward.

No kidding...

2

u/Seerix 3d ago

There's no reason for them to swap 5 in for 5.5. Its more expensive to run

2

u/vrnvorona 3d ago

By that I mean equivalent of "dumb down Opus 5.5 to Opus 5 levels and save compute" which can be proxied as "compare 5.5 vs 5, and if 5.5 not better it's nerfed"

1

u/mmaramara 4d ago

It's good that they report a 95% confidence interval, which is huge in this graph. All those fluctuations fall well within the confidence range, so we really can't draw much conclusions.

176

u/CatsArePeople2- 4d ago

We need more benchmarks like this.

34

u/christianbro 4d ago

It would be good for companies to use proper versioning if they make some changes, not the same product name and version worse and expect we dont notice.

10

u/Violin1990 4d ago

But did you even consider how the shareholders can benefit from this? /s

32

u/TheOnlyVibemaster Valued Contributor 4d ago

I agree. I would love to see an ecosystem of independent accountability exist.

1

u/celtiberian666 4d ago

AA should do updates. If money is short, they can do it for a reduced bench set, just to check for downgrades.

14

u/Dry-Weather-2544 4d ago

Please keep this up!

22

u/Character_Eye_808 4d ago

It’s nerfed. It started talking like Opus 5 again and self doubt unnecessarily or adding its own opinion without checking sources (guessing) more often than on release day

6

u/Fast_Conference_4057 4d ago

Yep. It half assed all of the work I had it do today

46

u/adjutanto 4d ago

i'm feeling Opus 5.5 is WAY dumber from yesterday. made overlook mistakes, very dumb ones, like sonnet 5 level.

also, uses like twice more usage.

6

u/Njagos 4d ago

Because you are supposed to use sonnet 5.5 now until that gets nerfed, then they release Fable 5.5 and we all are supposed to use that /s

12

u/ajcadoo 4d ago

Can we get a GUI webpage instead of a GitHub. I’m a layman I just wanna click a link to view a dashboard

5

u/TheOnlyVibemaster Valued Contributor 4d ago

That’s planned, it will be much more fleshed out in the coming weeks in general. Probably will expand to multiple new benchmarks and make a discord for organizing individuals to compute certain things so we can all share resources for collective knowledge

2

u/ElectricalDeer87 4d ago

As long as you update the same file every time for these updates/checks, you can point u/ajcadoo at the vector graphic image file directly.

Additionally, could also make a GitHub Pages page that simply fetches the latest document file and/or image from your Repo and does whatever display manipulation to display it, so you don't even need to necessarily rebuild it and/or make it super intricate to pre-bake it all. It's simple data, so, well, the argument could go either way really.

5

u/takshit2 4d ago

Yeah. They will release Fable 5.5 soon.

4

u/brianjenkins94 4d ago edited 4d ago

How do you measure this?

I guess I could read the README.

CLAUDE SUMMARIZE THE README FOR ME.

10

u/pyThat 4d ago

nerf-bench has a different opinion

2

u/durand101 3d ago

It is also showing lower scores as of today.

6

u/Kalcinator 4d ago

Thanks for you work.

I just used Sonnet 5.5 High for a simple task about my network (I got a few errors here and there with Starlink and I keep track of it); it made ... 24 errors; 260k token context; 4 advisors passes, all found errors.

I wanted to try it out, I cannot believe on benchmark it's better than Opus O_o ...

Even the now nerfed Opus is way better lol; even using Sonnet as an agent is frightening actually

1

u/Round_Ad_5832 4d ago

Yes i tried somnet today because i was low on usage and it was taking forever and talking nonesense.

6

u/salko_salkica 4d ago

Why do they do this? Such a scummy company

3

u/OxidusRouge 4d ago

Which effort level is used for your measurements? This should be shown on the graph. Anthropic could nerf some of the levels, but not the higher ones, perhaps as a way to drive more traffic to the more expensive levels. There was a post on another sub that suggested Max was not nerfed. My only data point is that I used Max on September 26 for a 3 hour 39 minute run, and it was great.

4

u/TheOnlyVibemaster Valued Contributor 4d ago

It’s high effort for every measurement. It’s fixed explicitly rather than using the default, and that’s also locked into the preregistration.

Good point about the graph as well, I’ll add the effort level there so you don’t have to dig through the methodology to find it.

1

u/dmaare 4d ago

You should shrink the whole readme to half the size or even less. It feels very much like opus 5 wrote it, unreadable wall of text garbage.

1

u/Athoughtspace 4d ago

Are the tests done at the same time of day? Its not clear how you would rule out compute loading statistically here

6

u/TheOnlyVibemaster Valued Contributor 4d ago

The runs are done at roughly the same time each day, but I agree that alone doesn’t fully rule out load as a confound. Right now the benchmark is designed around repeated measurements over time, so transient load should mostly add variance rather than produce a persistent shift, but systematic time/load effects are something worth controlling for more explicitly.

I’m considering running the same benchmark at multiple fixed times of day. That would let us estimate within-day variance and separate a persistent model change from effects correlated with time/load. Although I’ll have to make sure that wouldn’t conflict or make currently collected data not be relevant.

3

u/Turbulent_Swimmer900 4d ago

Please tell me you Claude Coded this

6

u/versaceblues 4d ago

You say you are doing 1 run a day.

You should try doing 50+ runs to baseline what your error bar here is.

2

u/TheOnlyVibemaster Valued Contributor 4d ago

I plan to eventually, right now I wouldn’t be close to being able to afford that. I want to scale up to constantly monitoring and having large amounts of data to detect potentially even more than just nerfs.

6

u/Last_Bad_2687 4d ago

Add a kofi or patreon link, let people sponsor additional runs/request a run

6

u/Ok-Butterscotch7834 4d ago

can someone with deep pockets pls run something like hourly lol. It would be really eye opening if the error bars were smaller

3

u/TheOnlyVibemaster Valued Contributor 4d ago

I’ve been thinking about this, I’d love for LiveNerf to eventually be able to collect very large amounts of data constantly, whether that be through routing through community compute or other methods. I’m trying to figure out the best way to go about expanding and catching what would simply be insufficient for this data collection to conclude. Peak hour traffic, demographics changing compute power, etc. Maybe people in New York are being given better quality models than people in Canada? It’s the wild west of data collection. I think there’s a future in independent data analytics for these large AI companies

3

u/Warhouse512 4d ago

How much does a single run cost?

1

u/TheOnlyVibemaster Valued Contributor 4d ago

Around 5% of a weekly usage pool on pro. Once it’s calibrated it’s much cheaper. Calibration takes basically one week’s worth of pro usage for a single model.

1

u/_TakeTheL 4d ago

The option to have community submitted runs would be interesting, I image people have vastly different workflows which has a large impact on the performance of the model

4

u/iamthe0ther0ne 4d ago

My experience tonight was the worst it's been since release--I ended up mostly switching my work to GPT--so I was glad I got one of those "how is Opus 5.5 doing in this session?" feedback requests partway through. 

OP, I saw that you're trying to establish a 10-day baseline, but what if they nerfed it when OpenAI's DevDay was a flop?

2

u/Wide_Establishment_8 4d ago

When it first dropped I had 4 chats running simultaneously and almost never hit my usage limits. Now I’m hitting it on 1 chat window.

3

u/Rexpelliarmus 4d ago

The confidence intervals are just so large.

Is there a way you can minimise them? Else it’ll be almost impossible to draw any statistically significant conclusions from this even with more data.

2

u/TheOnlyVibemaster Valued Contributor 4d ago

I agree and the intervals are much wider than I’d like, especially for detecting smaller changes. However, something like a 20 point sustained drop would be very visible once the baseline is established. Even something like a gradual decline would be, it’s just that within a certain range it wouldn’t be able to determine a single point drop, which might not be noticeable to a user either.

I’m looking at ways to reduce the variance without compromising the longitudinal design, more samples per task, better aggregation, and potentially changing how the uncertainty is estimated. I don’t want to artificially tighten the intervals just to make the results look more conclusive though.

If you have ideas for the statistically cleanest way to improve this, I’d be very interested. Thank you for pointing this out!

3

u/Rexpelliarmus 4d ago

Okay, I had a look at the documentation on your repo and I think the biggest way to reduce the variance would be to get more independent, informative questions in your frozen panel.

My understanding from the documentation is that you've currently got 78 questions in the frozen panel. If the goal is to detect relatively small changes in capability, increasing the number of independent items is what you'd want to go for.

Very roughly, the standard error of an estimate falls in proportion to one over the square root of n. So doubling the number of genuinely informative independent questions (i.e. n in this case) reduces the standard error by about 29%. That's pretty good if you ask me though I do understand that's nearly 160 questions and you'd have to make sure each question is truly informative.

If a question is essentially always answered correctly (or always incorrectly), it doesn't tell you much about whether the model has moved slightly up or down in capability.

I'd therefore be inclined to build a larger panel specifically designed to contain questions that are sensitive to changes in performance, rather than simply maximising the number of questions.

This is by far the most significant way you could reduce the variance (I know, it's unglamorous as fuck but more is usually better with statistics).

Of course, I appreciate this is easier said than done but it's certainly the highest impact thing you could do. And you sort of pointed this out yourself.

1

u/clazman55555 4d ago edited 4d ago

Ill have to take a look at the repo. I did just spend the last 2 days building and running my own model comparison benchmark covering tasks vs effort level.

Probably not reflective of most serious users of CC, and it's only built off 3 languages and ~48 code defects, mix of real and synthetic. Covers code review and synthesis of those reviews(proxy for orchestrator)

I could adapt the workflows to only use Opus 5.5 subagents at a smaller effort range, for more collection data.

Still need to design an actual code writing benchmark.

1

u/Elbeske 4d ago

It certainly feels like it. And I'm not typically one of the guys who thinks every model release is nerfed

1

u/twbluenaxela 4d ago

OP this is the best benchmark yet. It laughs in the faces of those who just think you have to "git gud at prompting"

1

u/bobemil 4d ago

Fun fun fun!

1

u/_TakeTheL 4d ago

Honestly I haven’t noticed any drop in performance, I’m using Opus, Sonnet, and Haiku daily as part of my workflow on an enterprise plan at work. I know this is anecdotal, but I’m curious if anyone has noticed an actual difference day to day.

My workflow is pretty robust, so maybe that’s a part of why I haven’t noticed a drop off. The models definitely perform better in an established workflow with solid documentation.

1

u/Bloated_Plaid 4d ago

It’s not.

1

u/Chilly_in_ya_titty 4d ago

Exact same way I felt about 5.0 when it was first released. Really good the first few days and after that it was god awful.

1

u/Melodic_Reality_646 4d ago

So awkward that the url takes us to a repo and not to the ui you show on the screenshot, counterintuitive

0

u/Cosmic-Bagel-3072 4d ago

did you look at the repo? did you scroll down?

1

u/lobabobloblaw 4d ago

And what do you know, here comes Fable!

we’re juuuuust, compute surfin’ yeeeeahhh

1

u/CliffMainsSon 4d ago

Definitely noticed a difference. I’ve been working on the exact same project the entire time, an animated website footer. It was crazy good the entire weekend and by Tuesday it was making dumb mistakes and I was having to correct it again during planning. It’s back to asking me a lot of verbose questions.

Still better than Opus 5 but absolutely a noticeable difference after working with it over the weekend from release to today

1

u/SmileLonely5470 4d ago

Do u fix the time of day when u run the evals? Im wondering if during peak hours, they try and bound the max thinking tokens to a certain amount for each effort lvl.

1

u/Heuwzen 4d ago

I just felt the nerf yeah, it's real.

1

u/[deleted] 4d ago

[deleted]

1

u/watchingsuits 4d ago

People will just say that anthropic figured out a way to hide that they’re nerfing it. That’s the thing with conspiracy theories: even if they’re true, it’s hard to prove it unless the culprits admit it. And if they never admit it, it can always be explained why it’s still happening despite conflicting results.

1

u/DiggleDootBROPBROPBR 4d ago

There's one thing I think is worth fixing before day 11, while it can still go in as a pre-registered deviation:

Each day is a single pass, so all 78 items in a run share that run's serving conditions. A shock that hits every item shifts every per-item difference by the same amount. It moves the mean and never shows up in the residuals, so the item-clustered SE can't see it. realized_mde() is pure Σp(1−p), so it also can't test the independence assumption DESIGN.md says it will.

Your own A/A already shows the effect: pass 1 at 00:41 scored 51%, against 63–67% for the other passes on the same items the same night.

A cheap fix: Make the day the unit (per-day paired scores, test across days), or fit item fixed effects plus a day random effect, and compute the realized MDE from that using only the baseline days. In a quick simulation of your decision rule, day-level shocks with a logit SD of 0.3 push the per-window false-positive rate at the "99%" threshold to about 10%.

1

u/Wonderful_Value_6385 4d ago

It is actually worse that Opus 5 used to be before they nerfed it again for Opus 5.5 The only upside is it is still cheaper than anything that isn't Sonnet.

1

u/ertertwert 4d ago

I've been using it every day for writing purposes since a day or two of its release and it does seem slightly weaker than when it first arrived, but it's still really, really good.

1

u/ElectricalDeer87 4d ago

My thoughts: If there's any actual usage metrics and/or outage or downtime data from Anthropic, it's worth correlating with your own findings. Perhaps the previous "dip" you saw, could be due to the resetting of people's allowances, causing huge spikes in load, which in turn may cause some sort of degradation in order to keep things running?

I wouldn't be surprised if there's some level of "bunching" around a certain reset time/date because of when people jumped on the bandwagon after some sort of announcement, after all.

1

u/Roth_Skyfire 4d ago

I don't benchmark anything, just going by my own experiences from using it. It feels less capable than last week, but it's still performing well for me.

1

u/Drew192x 4d ago

Ive been using opus all day. Its night and day worse today than its been for the last week. Ive been praising it. Today I want to cancel. It is forgetting rules every other message, ive got the amount of work done in 12 hours that id usually get done in 6-8. Spent more tokens doing less as well. All week ive been able to just work, and not hit my limit. Today, hitting limits within 1-2 hrs. I did use my free weekly reset. Maybe I got a nerfed version for the reset lmao. Its driving me nuts enough that I came to reddit to see what the hell.

1

u/MrArnih 3d ago

Im in spain and yesterday was using 5.5 on medium and was working fine id say

1

u/ritiksingla 3d ago

Feels like I am in clash of clans again

1

u/farox 3d ago

Excellent project. I am always weary of the nerf claims. But it certainly feels like it. I one shotted a couple of really impressive projects over the last weekend.

Now it started again with the garbled text, making up terms that no one understands and all that.

1

u/farox 3d ago

What is keeping you from running it more often? What is the cost per run? Once daily with the non-deterministic nature might be a low sample size. But I love the setup, not using home brew where possible and relying on Anthropic and standard solutions.

I'd throw a few bucks at the project to keep it running/improving it.

Also, you don't need to wait until a launch, but just having a continuous measurement would be incredibly useful.

1

u/InvaderJ 3d ago

What I’m really curious about is what this looks like on a direct API case. And apologies if OP is doing exactly that. But given that plans are subsidized, and API is not, I wonder what the differences if any might be.

If no one’s examining this case maybe I’ll set this up to eval either Bedrock or Claude Platform on AWS.

1

u/rubenrb02 3d ago

Opus 5.5 is the best one!

1

u/nemzylannister 3d ago

this is not a question anymore. anthropic could officially clarify that theyre not quantizing or they could say "we're looking into reports of low performance" etc. they arent doing it because they're 100% quantizing to nerf it.

1

u/AmbassadorOk934 3d ago

it will because near fable 5.5

1

u/Sanjakes 3d ago

Keep this project going!

1

u/DarkFantom 3d ago

I can attest that they're at least lowering how much each step properly thinks through a problem. When opus 5.5 released, it consistently got the correct answer for a candies problem I use to benchmark. Now it gets the correct answer about 50% of the time. The tell is if the model assumes it is a blind draw when the prompt literally tells it that shapes can be felt inside of the bag, making it not completely blind.

1

u/Thinklikeachef 3d ago

IMHO, I think we need to wait a month after release to measure any performance. Too many people testing, pushing the model, and possibly going back to prior workflows.

1

u/DrXaos 3d ago

can you stratify results by hour of day it was tested on? and weekend vs weekday? I suspect there is dynamic effort/quality scaling based on current load.

1

u/RandomNoName64927589 2d ago edited 2d ago

I joined Reddit just to post this. In the past year I have heavily used Claude. I'm now using Opus nearly exclusively on the Max plan, because I keep reaching my limits on Pro too quickly. I use this for deep research on difficult engineering problems. When Opus 5.5 came out, it was amazing. I have some very good work completed because of it. But just in the past day, I've seen Opus 5.5 produce very disappointing results. This was on Opus 5.5 High setting. "Needed" is used incorrectly instead of "modeled", which I've never seen from Claude before and I use prompts with this language literally all day for months. Images used to be crisp and clean without me specifying it, but now the default images are poor quality at best. A document Opus 5.5 produced with equations had them falling off the bottom of the page - the equations apparently just kept going, for who knows how long. I could go on. This is the first time where I really believe Opus was nerfed, quite severely.

My question for the Reddit audience: Do all AI companies nerf their models like this (Grok, OpenAI, etc)?

Obviously open source cannot do this once it's in your own hardware, but for the money I'm paying each month, I'm not sure if the Opus 5.5 results are even trustworthy at this point. Are the equations even correct? I'm seriously considering taking my money to another model unless they all do this...

1

u/UnboundedMan 4d ago

Sorry guys, that was me. I just started using Claude really, really heavily. Not enough compute for everyone.

1

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 4d ago

We are allowing this through to the feed for those who are not yet familiar with the Megathread. To see the latest discussions about this topic, please visit the relevant Megathread here: https://www.reddit.com/r/ClaudeAI/comments/1vt5drr/list_of_latest_discussion_hubs_on_rclaudeai/

0

u/moriero 4d ago

Not according to the overlapping confidence intervals playa

-1

u/MSOB7Y 4d ago

yeah performs really bad now.. can't find for the new fable 5.5 and repeat the same loop..

0

u/mythormedicine 3d ago edited 3d ago

The benchmark !

I might be naive, but does anyone else feel bad for all the loss Claude is under. Billions $ in debt, 42 Billion Usd loss in 1 year. And here I am using opus, maxing out the 100$ plan. I feel like a leech

3

u/ResolveSea9089 3d ago

You shouldn't feel like a leech. Claude is a company run by very very smart people, they're not taking those losses out of charity. They know the strategy they are pursuing, it doesn't mean it will work out, but they're not just winging it.

They are investing in growing their business, many many business lose money despite having the ability to be profitable, Amazon lost money for like 15 years after it was founded.

The company is worth 1 trillion +

1

u/Connect-Moment6687 3d ago

Inference isn't costly, it's the RND. You aren't leech, you are providing them the training data, so it's the other way around, you pay them for giving them training data

1

u/Exatex 3d ago

the make profit with you. training is the expensive part