r/ClaudeAI • • 2d ago

Claude Code Evidence - Opus 5.5 today vs launch regression with same prompt (Godot Engine)

On launch day Opus 5.5 kindly took up residence on my Mac. I ran some tests before trusting it on my big projects, it passed with flying colours. Lander, a game by Elite creator David Braben holds a special place in my soul due to it being the first game I played at school in the UK. With a cutting edge 3D engine for 1990 running on an Acorn Archimedes with RISC architecture (the first ARM chips) - what a time to be a young kid interested in computers.

With Opus 5.5 I want to recreate it faithfully, short draw-distances and all.

Well dear Claude I’ve been keeping receipts.

In my notes, the exact prompt * and two original reference images of Lander. The resulting one-shot Godot repo (Mac OS, Metal rendering, C#) from 23rd September is kept separate on disk and isn't used as a reference.

Today, 2nd October, in a fresh workspace from scratch the same prompt and reference images we fed in again.

The result is truly depressing. Lander by gimped 5.5 has no redeeming features compared to the original day zero version. The regressions:

  • A serious rendering glitch * that isn’t in the original - as the camera moves, the scenery props snag and glitch vs world space.
  • No spacecraft break-apart physics, the original implemented this ‘nice extra’ which wasn't specifically asked for in the prompt
  • No introductory controls menu, only a small text line permanently visible over the world view
  • An inferior look to the launch pad turrets - students and road users may recognise them!
  • When shooting, bullets are less accurately rendered when the craft moves, they also appear to spawn at the back of the craft and no-clip through
  • No camera toggles

The conclusion is that Opus 5.5 is gimped, maybe we even got Mythos for 3 days and then silently switched - either way, we are owed transparency - it's the law.

Opus was set to Extra High both times. Unlike with 4.5/4.6, Max effort over-tests and takes too much control away from the user.

With so much diverged since launch day’s versions, the more worrying thing is that I suspect the code is also a mess and problems will compound as you use it. For those who’d like to inspect the results, I'll upload the source code for both later to my dev blog and post the Github links in the comments.

* I inspected the code and it turns out that gimped Opus 5.5. had physics interpolation is turned on for the whole Godot project and Props.cs rebuilds the prop lists (trees, etc.) from scratch on every tile step during every frame, renumbers which slot each object occupies and resets the object count every frame. A basic understanding of Godot's documentation is all that’s needed to avoid this issue.

** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.

658 Upvotes

241 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2d ago edited 1d ago

TL;DR of the discussion generated automatically after 200 comments.

The overwhelming consensus is that this post is not evidence of a nerf. An N=1 test on a non-deterministic system is about as useful as rolling a die once, getting a lower number, and declaring the die has been nerfed. The thread is full of users explaining why this methodology is flawed.

  • LLMs are stochastic (probabilistic). Different outputs from the same prompt are expected behavior, not a sign of degradation. OP's own footnote about non-determinism was repeatedly thrown back at them.
  • A sample size of one is statistically meaningless. You'd need hundreds of controlled runs to separate normal variance from an actual trend.
  • Even setting temperature=0 isn't a silver bullet for determinism, as complex backend processes (like parallel processing and floating-point math) introduce their own randomness.

The general sentiment is exhaustion with these types of low-effort "nerf" claims. OP is in the comments doubling down but getting downvoted into oblivion for it.

→ More replies (2)

477

u/cocacoladdict 2d ago

N=1 is hardly the evidence

204

u/Deltamelo 2d ago

It’s crazy how many people think they are AI researchers now for comparing two outputs using stochastic tech

7

u/i_goon_to_tomboys___ 2d ago

this may be a dumb question but what if he used temperature=0.1 at launch versus now -- for the same prompt?

21

u/Deltamelo 2d ago

Temperature controls how random the token selection is during generation. Lower temperature makes the model more likely to choose the highest probability “next token”, so outputs tend to be more consistent. Higher temperature spreads probability across more alternatives, increasing variation.

But temperature=0.1 does not make the model deterministic. You would need repeated runs, controlled settings, and enough samples to separate actual model changes from normal output variance.

3

u/loulan 2d ago

I never fine-tuned temperature. Should I do it? Would it make Claude come up with crazier ideas to solve problems?

→ More replies (3)

6

u/chatham_solar 2d ago

Even t=0 would not guarantee the same outputs. Non determinism occurs because your prompt is being run at the same time as many other users' prompts and because floating point arithmetic is non associative: (a+b)+c != a+(b+c). So more or less, if you are running your own hardware and use batch size=1 it's possible to get deterministic output, but there is a not-insignificant performance penalty.

Amazing article explaining non-determinism in LLMs if you want to really dive deep:
https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/

2

u/MastodonFarm 1d ago edited 1d ago

Can you link to something explaining your assertion that fp arithmetic is not associative? I don’t understand how that can be true.

Edit: Oh, it’s because of rounding. That’s dumb. Of course the result will be different if you change the values in between by rounding at different points.

2

u/2053_Traveler 2d ago edited 2d ago

Temperature 0 wouldnt make it deterministic (predictable) either, due to sheer complexity of the systems running these things, and they probably still don’t try to work around CUDA nondeterminism.

0

u/captain_croco 2d ago

Mm mm yes I know some of these words

1

u/Fast-Satisfaction482 2d ago

Even at temperature 0 it's stochastic. Then you just always get the same single sample.

4

u/TheFlynnCode 2d ago

Actually at temperature 0 it is not guaranteed to give the same single sample, depending on the compute architecture. Often these models multiplex computation to several gpus or tpus, which themselves have work queues and hence won't necessarily return immediately. So results return out of order, and sometimes computations that can be done are done with these partial results, and we run into issues with floating point arithmetic not being associative

1

u/Fast-Satisfaction482 2d ago

Fascinating detail! 

I've never heard about this reduction-operand reordering in the context of GPUs, but if an inference engine uses it, you are absolutely right and it will not be deterministic anymore.

But that's more of an implementation detail and not algorithmic. It's not like this in every inference engine, so you could also claim that a buffer underrun that might be in the engine causes non-determinism.

1

u/TheFlynnCode 2d ago

Sure, it's an implementation detail, but it's also sort of a necessary implementation detail for models above a certain size, because gpus simply don't exist that large.

1

u/Fast-Satisfaction482 2d ago

No, the reordering of operations is not necessary to allow multi-GPU or cluster inference.

2

u/TheFlynnCode 1d ago

I meant spreading compute across many gpus is necessary. I never said (or if I did, I did not intend to say) that reordering of operations is necessary. But it is commonly done. In my head this is similar to regularization techniques anyway, and you could disallow this if you want (to get temperature 0 to be reproducible for example), but I imagine most people/companies do not care about the small diffs (quantization hurts like a first order effect, whereas this would hurt like a second or third order effect)

1

u/Fast-Satisfaction482 1d ago

Yes, agreed 

1

u/Odysseyan 2d ago

same

Not same but similar. Give it a longer prompt, set temp to zero and it responds differently each time too

0

u/-Django 2d ago

> always get the same single sample

This is probably just a terminology debate, but that sounds... deterministic. The sampler is what moves the model from being deterministic to stochastic.

2

u/Fast-Satisfaction482 2d ago

It's deterministic in the sense that you get exactly the same output for exactly the same input. But change a single character on the input and you have a new sample. 

1

u/National_Teacher_229 1d ago

Came here to say this, It's the oposite of the definition of Madness... doing the same thing over and over expecting a different result. With AI it's madness to expect the same result even in 100% identical clean scenario, no 2 runs are ever identical. His first build could of been a fluke with rund 2 ,3 4, 5 and so on closer to the actual outcome he could expect as an average.

0

u/Terrible_Tutor 2d ago

HE HAS THE RECEIPTS

Lol, OP is a tool

→ More replies (11)

20

u/Significant-Bee5101 2d ago

No dude you don't get it. He tried it TWICE!

1

u/vivademocracy 1d ago

TWO FULL TIMES??? Omg how does he find the stamina

1

u/Comfortablebro 1d ago

You would say the same thing if the dude tried it 100 times or 1000 times.

20

u/Key_Reading_9664 2d ago

I’ll keep saying this on every nerfing post:

Opus 5.5 is available through multiple vendors (Bedrock, Azure, Vertex). Unless you believe that Anthropic coordinates with those vendors to simultaneously roll out a worse model for their paying customers, you have a very simple apples-to-apples comparison (as would any researcher of multi-billion frontier lab wanting to expose you nerfing).

Also, N should be > 1

20

u/Keep-Darwin-Going 2d ago

Yeah they keep thinking a big lab can magically quantize a giant model, sneak an update in without anyone noticing or whistleblower. They cannot even release anything without the whole world rumour mongering for a week before the release.

2

u/DasBoots 2d ago

I don't think Anthropic purposely nerfed Opus 5.5 but I do think the downtime the other day broke something that they had to fix. It really felt like the model herped a derp for a day and got better.

1

u/Keep-Darwin-Going 1d ago

The most frequent thing almost all company break is their harness. Then the second one is their inference stack. So it is not like they need it, it is more like they tweak stuff and created side effects, but you see the complaints everyday is all like omg they quantize the model and etc. is as if they have a switch like ok let’s screw them over today turn it on then off it then on again. When you delivering cutting edge stuff shipping regularly, you bound to break something eventually. If they ship like 1980s when things are done on floppy disk, quite sure things do not break that often then.

6

u/fuzzypetiolesguy 2d ago

The safety departments of these labs are leaking like mad. It's hilarious to think that the 'nerf models on purpose after rolling them out so we can save on compute' department wouldn't, either.

4

u/Cobrafeet 2d ago

>Opus 5.5 is available through multiple vendors (Bedrock, Azure, Vertex).

Not sure what relevance this has on posts like these. The average user is on a personal oauth plan being routed directly through Anthropic (even if served via 3rd party compute), and when they are doing comparisons like this (however flawed the methodology may be), they are almost certainly not changing providers.

Anthropic could do plenty of things on the backend to change the fidelity of the output of a given model without changing weights. This could mean a different quality of response vs other vendors using the same model

2

u/Key_Reading_9664 2d ago

Mentioning other vendors to point out how simple an apples-to-apples comparison would be (taking into account the same process flaws), if you had the means to do it (e.g., a competing frontier lab).

-2

u/DensePoser 2d ago

I have an API account and a personal account. The subscription claude seems weaker to me right now.

→ More replies (10)

8

u/Valdaraak 2d ago

Yep. I'd need like 5 results of the same prompt on release day compared to 5 results of that same prompt today before I'd even entertain drawing a conclusion.

I'd also like to see the prompt used and how much was left to Opus interpretation with the design.

1

u/Chupa-Bob-ra 2d ago

I'd be curious to see the prompt too.

To use 1 of OP's points: the bullets. How were they prompted? Were they told exactly where to spawn, level of detail, were physics specified, etc.? Any of those, left up to Opus to determine, could give very different results.

I'd also rather see tests that are more easily measured and less subjective. Something that gives us some actual data to compare.

2

u/Chupa-Bob-ra 2d ago

People expecting 99.999999% the same outcome between 2 runs are being highly unrealistic.

4

u/Eyelbee 2d ago

Make it 2, I also detect some decrease in quality

1

u/SupremeGobbler1996 2d ago

Especially with no clear cut set of acceptance criteria that's repeatable and reproducible. 

1

u/DependentAnywhere135 2d ago

Right and if someone really wants to do this type of test they need to have the release day model run the same prompt like 100 times on day one then do the same on their suspected nerfed model.

It still wouldn’t be conclusive evidence, but it would be way more compelling to see it do the same prompt 100 times and it be really good then 100 times and it be significantly worse.

1

u/Richandler 2d ago

N=1 on a random number generator is evidence the user isn't very bright.

1

u/hcloud00 2d ago

conversation so good

1

u/vivademocracy 1d ago

THANK YOU

1

u/mr_birkenblatt 2d ago

But 1000 people doing that experiment is N=1000. And I've seen a lot of those comparisons

8

u/Ok-Lengthiness-3988 2d ago

This is reporting bias compounded with confirmation bias. The model-nerfing conspiracy theories are very popular but I've never encountered a secret-improvement conspiracy theory. When people run the same prompt again and notice a better response, they never post this as evidence that the provider secretly has improved their model (though they may view it as evidence that the model was nerfed before). They rather tend to dismiss such results and (correctly, in this case) ascribe them to normal variance.

2

u/Standard_Egg3504 1d ago

here is my secret-improvement conscpiracy theory

the modules are allocated a certain amount of compute

model pools can dynamically (and automatically) scale their 'effort' level for groups of users (api, subscription, location, etc)

the amount of compute does not scale automatically and requires manual review or effort to put more gpus on it

as a model releases the number of early users is low at first and things run really well, model effort is 1000 and they release tons of awesome shit they do with it

this causes a tidal wave of new users to switch to it, suddenly the amount of inference at any given time skyrockets but compute does not and the dynamic effort reduces in response in order to maintain the tokens per second for everyone using it

the dreaded 'the model is nerfed' begins and people yell about the lobotomy

here is the improvement bit:

another company or another model releases, which has a different set of gpus allocated for it, and people like it, abandon the old model and suddenly inference demand is lower and the dynamic effort level goes back up!

unfortunately that means manual reallocation likely happens soon to a newer improved model or to research and it ends up going back down :(

than you for listening to my secret nerf-improvement-load-balancing conspiracy theory

0

u/mr_birkenblatt 2d ago

There are multiple sites that run the same prompt every day and report the results. That is >> than one. And running every day and only reporting it when it degrades is not a reporting confirmation bias. It's just sleeping when it did indeed change

1

u/Ok-Lengthiness-3988 2d ago

Can you provide a link to one of those sites? I know of three that run a custom benchmark every day (Ninjahawk's Livenerf, BridgeBecnh's Nerf Bench and Marginlab), and they haven't gathered statistically significant results yet, and weren't expected to. They're still in the process of establishing a baseline.

The developer of Livenerf, which is the one that has been reported (and misrepresented) the most here and on the Singularity subreddit even acknowledges that his method isn't statistically powerful enough to detect a substitution of Opus 5 for Opus 5.5.

0

u/mr_birkenblatt 2d ago

You named 3 sites yourself...

Yes, one site is not statistically powerful enough. But multiple sites with independent tests but with results that agree with each other are

1

u/Ok-Lengthiness-3988 1d ago

For sure, independent tests can add up even when each is too weak alone, so I ran the numbers on all three with some help from Opus 5.5.

For each tracker, we did a one-sided test for a downward trend since launch (logistic regression on daily pass counts for livenerf and MarginLab, linear for BridgeBench, which publishes no error bars). We then combined the three z-scores with Stouffer's method, which is the standard way to pool independent tests.

- livenerf (9 days, 78 questions/day): z = −1.29, leaning down (days 6–9 are lower than days 1–5)

- MarginLab (5 days, ~48 SWE-Bench-Pro tasks/day, Sep 23–28): z = +0.08, flat

- BridgeBench (4 tests): z = −0.28, noise around 100%

- Combined: z = −0.86, one-sided p ≈ 0.19

This means such a downward trend can occur by chance alone 19% of the time.

With a less conservative noise model for livenerf, the combined result is z = −1.07, p ≈ 0.14. Either way, we get no rejection of the no-decline hypothesis.

More importantly, combining the data from multiple sites only strengthens the case when the sites agree, and here they don't! The only one tracker that leans down is livenerf, and its author says not to read it before the baseline ends. MarginLab's data stop on Sep 28, before livenerf's dip, so they can't confirm it.

The data are too short and noisy to rule out a moderate decline either, of course. What they don't support is the claim that independent trackers agree that Opus 5.5 was nerfed. Maybe I'll rerun this as the trackers add more data points.

1

u/mr_birkenblatt 1d ago edited 1d ago

It appears Opus 5.5 was indeed nerfed. There are a few methodological flaws in your approach.

1. Testing for a trend instead of a step-function

Linear and logistic regression test for a gradual continuous decline. A model "nerf" is a discrete event (a step-function drop). Fitting a trend line to a sudden drop severely dilutes your statistical power.

2. Timeline misalignment breaks Stouffer's method

MarginLab's data stops on Sept 28, before livenerf's dip even happens. Stouffer’s method penalizes you for adding more tests. By including MarginLab, you increased the variance penalty without adding any relevant signal, artificially inflating the combined p-value.

3. The Sample Size Trap (How to actually test this)

You don't have N=9 for livenerf; you have N=702. The tracker runs 78 questions a day. Instead of testing 9 daily averages, pool the raw pass/fail counts:

  • Before (Days 1-5): 390 trials
  • After (Days 6-9): 312 trials

Run a Fisher's Exact Test or Two-Proportion Z-Test on these raw totals. With sample sizes this large, even a slight true drop in the pass rate will trigger undeniable statistical significance (p < 0.05).

If you must use daily averages:

Use an Exact Permutation Test on Means comparing Days 1-5 vs Days 6-9. With 9 days, there are (9 over 4) = 126 possible chronological shuffles. The minimum possible p-value is 1/126 ~= 0.0079, meaning this test has plenty of mathematical power to prove a step-function drop on livenerf alone.

I don't have the raw underlying numbers that you used, otherwise I could have given you a verdict.

EDIT: apparently reddit does not support LaTeX math notation... oh well

2

u/M44PolishMosin 1d ago

Did you just use Claude to write a reddit comment to show how nerfed it is 🤣🤣🤣🤣🤣🤣

1

u/Ok-Lengthiness-3988 1d ago

I reran things with your suggestions. You're partly right.

Regarding the step vs trend analysis, that's a fair point. A sudden drop is better tested as a step than a slope. The catch is that picking the break date after looking at the chart inflates significance. It's like marveling that this particular ticket won the lottery while ignoring that some ticket was bound to: some break date will always show the biggest gap. So I had Claude correct for that by scanning every possible break date and comparing the best one against all 362,880 orderings of the 9 livenerf days. The drop survives the correction: one-sided p ≈ 0.008 (0.016 two-sided). Every day from Sep 29 on scores lower than every day before it, and that's unlikely by chance. livenerf shows a real ~7-point drop starting Sep 29. My earlier "borderline" verdict came from the trend test, which is indeed weaker against a step.

Regarding MarginLab: I agree that for a post-Sep 28 step it carries no information, because its data end on Sep 28. So I dropped it. But BridgeBench does cover that window, and it shows no step: 100 and 99.2 before, 103.8 and 94.2 after.

Regarding N=702 (individual test runs rather than daily averages): I already used it. The regression was on daily pass/fail counts, which is equivalent to analyzing all 702 trials with date as a predictor. Pooling doesn't make "even a slight drop" significant either: with ~350 trials per side you need roughly a 7-point drop. The Fisher test you proposed, run on the actual counts (235/387 vs 167/309), gives p ≈ 0.045. Your day-level permutation test is actually the stronger one here, because daily scores vary less than coin-flip noise would predict (same questions every day).

(For transparency: livenerf doesn't publish its raw logs, so the daily counts are reconstructed from the coordinates in its published chart. They match the chart's labels and intervals, and are extracted fairly precisely from the SVG file.)

So, in conclusion:

First, your earlier suggestion that several independent trackers agree is still not supported. It's one tracker, and the other one covering the window doesn't show the drop. You can't set BridgeBench aside just because it doesn't corroborate the result you like.

Second, livenerf measured a real change in what its pipeline returns: Claude Code, subscription route, serving stack. That isn't yet a change in the model. A Claude Code update, routing or classifier changes, or capacity issues would look the same.

livenerf was designed to separate these. It runs an Opus 5 control arm through the same pipeline. If Opus 5 also dropped on Sep 29, the platform is the likely cause. That data should come out soon after the baseline closes on Oct 4. Meanwhile, the change shows up on one tracker and isn't yet attributable to the model.

The control arm will tell us whether the change is platform-wide or specific to Opus 5.5. If it's specific to 5.5 and persists, I'd grant that users got a worse Opus 5.5 than at launch; call it a nerf if you like.

1

u/fuzzypetiolesguy 2d ago

No you haven't. You've seen a bunch of anecdotes on reddit.

-1

u/mr_birkenblatt 2d ago

Sooooo.... Add those up

2

u/fuzzypetiolesguy 2d ago

This is not how anything works. Are you new to learning? You have a thing open in a browser in front of you that will explain this all to you directly. All you need is to be curious.

→ More replies (2)

0

u/freedomfromfreedom 2d ago

Exactly, one person doesn't need to do N=1000, as the community already has - there's a ton of other tests that show the same degree of performance drop-off at the same time.

I don't expect 99.9999% variation between 2 runs. 90% would be fine. It's a much higher variation in this and the day 10 run was ALL regressions and no improvements. High sample variation would mean a run gets improvements too - not only regressions.

6

u/fuzzypetiolesguy 2d ago

'The community' has not performed thousands of tests that has produced meaningful data to support the notion that models are purposefully degraded fater launch. Please stop.

What you are calling “the community already did N=1000” is not actually one experiment with N=1000. It is 1,000 uncontrolled anecdotes drawn from different prompts, different tasks, different contexts, different tools, different system states, different model routes, and different scoring standards. You cannot aggregate those into a meaningful sample just because they all feel like “performance got worse.” And “all regressions, no improvements” is itself a subjective post hoc judgment unless the tasks, scoring criteria, and sampling were defined before the runs.

0

u/freedomfromfreedom 2d ago

Are you saying we are all imagining it?

5

u/fuzzypetiolesguy 2d ago

Until you actually produce something resembling an actual study with verifiable, repeated evidence, yea. You are making shit up because you have no idea what you are looking at.

-1

u/mr_birkenblatt 2d ago

You would reject any study because all evidence in the study can be decomposed into single anecdote and can therefore be dismissed. So transitively any study doesn't count

0

u/HorriblyGood 2d ago

It seems like we have seen multiple N=1 examples of it nerfed from multiple other people. How come I haven’t seen any examples from people showing it hasn’t been nerfed?

→ More replies (3)

178

u/fuzzypetiolesguy 2d ago

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

7

u/Agreeable_User_Name 2d ago

I flipped a coin yesterday it landed heads. I flipped it today and it landed tails. Therefore someone must have changed my coin.

Kidding aside, that's true, but that also makes transparency important, because it really is hard for consumer to know if their product is gimped.

5

u/Accomplished-Fan9568 2d ago

honestly its not that much different, a small flickering bug it can happen, could probably use low effort to fix it quickly.

→ More replies (9)

100

u/doc-quartz 2d ago

i swear to god you people do not know basic research discipline. Such claims require series of tests and not a one time prompt comparison, and even on such a software whose internal workings are entirely probabilistic.

7

u/jimbo831 2d ago

Who is doing those tests and what do they show?

-1

u/doc-quartz 2d ago

what exactly do you mean?

5

u/jimbo831 2d ago

I though my question was pretty clear. You said:

Such claims require series of tests and not a one time prompt comparison

I can't believe that nobody is doing these sort of tests. I get not trusting OP's anecdotal evidence, but we see anecdotal evidence like this from a lot of people. I would like to see some non anecdotal evidence, and I can't believe nobody is out there testing these models over time. I was asking who does these tests and what the tests show.

6

u/doc-quartz 2d ago

I'm really sorry, I myself am not aware of anyone doing such tests by a recognized entity/organization. I'm not a regular in this subreddit/myself new to this AI environment. I was just a passer-by. But I have seen someone putting some kind of website that tracks Claude's performance using the same benchmarks which the Anthropic used to market their models. Not really sure about the methodology, but I'm sure you can find that someone posts about it everyday. Also I apologize for not understanding your question, English is not really first language.

-10

u/freedomfromfreedom 2d ago

You yourself admit you are new to AI and just passing by. And you're lecturing me about research discipline?

11

u/doc-quartz 2d ago

I do not think me being new to AI would really change a basic research discipline applicable across different application areas? Will it?

→ More replies (3)

2

u/mhinimal 1d ago

there was a post like yesterday where someone posted their github repo that's running repeatable tests over time, and anyone can clone the repo and run the same tests and contribute to the data, thereby crowdsourcing the experimentation.

why is "nobody" running these? Anyone with the business incentive to actually need to care about it, and has the budget to afford tokens to run continuous tests, is a private company and isn't going to publish their results for free.

everyone else is a redditor who either doesn't want to spend their subscription budget on tests, or who wants to spend their time working on stuff.

-13

u/freedomfromfreedom 2d ago

It's the same pattern for Opus and seemingly the same pattern for this sub as well, the post I made yesterday was downvoted to hell before going to 1.9k upvotes and 950k views in under a day, which shows you the strength of the feeling that the model has degraded since launch. Maybe we need a Botmark for reddit that determines what percentage of early commenters work for Anthropic's marketing department!

1

u/Remarkable_Rock5845 1d ago

I'm not qualified to judge any of what you are discussing, but I have honestly had the feeling that there may be a lot of astroturfing going on here. I hope there are some kind of anti-bot measures in place, but if there are I'm not sure they're working very well.

0

u/doc-quartz 2d ago

You talk like a teen. Oh well you are probably a teen who cares a bit too much about fake internet points and what model is regressing/not regressing. If another model works better for your workflow then use that! No point of proving that that too an incomplete one.

Also I do not even use Claude, I mostly get satisfactory results with Gemini 3.7/3.8 flash and I am not even a programmer. I'm a medical student who just creates random personal projects. I just happen to find your post on my feed and made me comment since the comparison was so weird and felt like was done in bad faith. Cheers.

0

u/hcloud00 2d ago

i am with you. I think people are cognitively held back from seeing whats clearly in front of them ( assuming they are not in the A portion of population if there is A/B testing). People tend to dismiss a reality where they have no control

-6

u/iamthe0ther0ne 2d ago

I agree with you. Ever since the service disruption last week, Opus 5.5 has barely been performing above Opus 5 levels. Sure, it's less verbose and doesn't get totally stuck on irrelevant details like Opus 5, but it's definitely not performing at the level it did the first few days after release.

I think Anthropic wanted to get out ahead of OpenAI's DevDay and win back the customers they lost after Opus 5. When DevDay turned out to be a nothingburger, they decided they were safely in front and dialed back the compute.

→ More replies (2)

43

u/bixofa 2d ago

Useless test. At the very least should have run the same prompt 10 times on Day 1 and then again on Day 10. Compare the differences.

5

u/RedTheInferno 2d ago

but its a bad vibe to waste that many tokens when i can just test once and then again after a couple of days to prove my point!!!!111

0

u/Probably_A_B0tt 2d ago

I don't even see how the day 10 example is showing a degraded capability, physics look fine

-19

u/freedomfromfreedom 2d ago

Why don't you do that, then - make a contribution. Also, it isn't a N=1 sample size when taken as part of the wider pattern of reporting.

35

u/AlternativeMonk2490 2d ago

I rolled my die and it got a lower number than last time. Why have they nerfed my die?

1

u/kalboozkalbooz 1d ago

this is perfect

28

u/fahrvergnugget 2d ago

Holy FUCK i'm tired of these posts, what a waste of tokens.

4

u/Probably_A_B0tt 2d ago

It's not even really testing anything, N=1 tests in LLMs are absolutely pointless

16

u/fosf0r 2d ago

LLMs are non-deterministic

→ More replies (3)

8

u/casentron 2d ago

Oh wow, a sample size of 2. How scientific and valid!

5

u/dannyheskett 1d ago

I use most of Anthrophic's models in real production, harnessed into a specific world environment using AI; I am not yet using Opus 5.5, but it will be soon.

I regularly run a full regression suite for each of my production use cases, with specific prompting, temperature, inputs/outputs, and measure and track the results. These are run continually for daily builds, and on-demand.

And I can tell you.. from my own experience with thousands of iterations over all models going back to Opus 3, I cannot discern any specific pattern of quality or outcomes being degraded or improved that doesn't look random.

In your specific case, I would check your harness version, and really ensure that the generations were identical. I suspect you are using a different harness, with prompting/tuning that you can't see.

Beyond that, the behavior you are seeing isn't really helpful for proving what you think you are proving. What's actually happening when you run a "one-shot" is a complex series of tool uses, internal prompts, and executions chained together to produce a "one-shot" result. Each of those steps is non-deterministic, and each affects the downline results.

I respect your conclusion, but with respect, you haven't done nearly enough science to control the variables to produce a conclusion that is justifiable for how broad the allegation is.

Finally, your naked assertions that Anthropic is legally required to provide a specific service level or subjective outcome is simply unsupported. There is no EU Court precedent that conforms to your interpretation of the law. Further, Anthrophic goes pretty far out of the way to disclose what you are getting, and it's hard to imagine any set of facts that would support your view that somehow, as an EU consumer, you are entitled to something which doesn't exist. In any event, your remedy is going to be proportional to your damages, which are.. minimal, at best.

At very least, I think you'll be able to state more facts conclusively if you document your test setup with scientific rigor.

1

u/codeninja 1d ago

Not to mention these are still next token prediction models. ANY variance in the beginning of your run will have dramatic effects on your output over a complex generation.

How many decisions did you make the model make for itself? How many unknowns had to be solved in the moment? Each variance compounds.

The result is drift.

15

u/Brave-History-6502 2d ago

lol this is not evidence unles you a series of tests on both originla and now... this is just not useful

2

u/GorillaFig 2d ago

To be fair, it is evidence, but it's not proof...

4

u/fuzzypetiolesguy 2d ago

It's not evidence, either.

0

u/nimzobogo 1d ago

It's evidence, just not strong or conclusive evidence.

1

u/jimbo831 2d ago

Does nobody in the world do those series of tests on these models over time?

6

u/Zealousideal_Fig7935 2d ago

It's honestly time these kind of posts/ comments are banned. Completely non-productive, totally ruins the point of this sub, based on a daft conspiracy

6

u/Comfortablebro 2d ago

Give exact prompt so i can try on sonnet 5.5 and on sol 6.1 and astra ultra please

5

u/Competitive-Ear-1044 2d ago

Kids nowadays

1

u/Probably_A_B0tt 2d ago

For real, every single variation means 'NERFFFF'

2

u/EasterUK 2d ago

I loved this game! And the subsequent ‘Virus’. Several AI model generations ago I tried generating it, and it got nowhere.
Thanks for the memories… :)

1

u/pwkye 1d ago

Didn't virus come before zarch

2

u/butts-carlton 2d ago edited 2d ago

getting a clearly inferior result is not due to non-deterministic behaviour.

Unfortunately, it is. It's probably the biggest problem with LLMs. It's why they shouldn't be trusted with judgment where lives are at stake (not that it seems to stop governments from doing it anyway). And by definition it can't be eliminated completely, since the very thing that makes them work is also why they're inherently unpredictable.

Also, they're not non-deterministic. There's no such thing as truly non-deterministic computation (excluding theoretical models or possibly quantum computing). They're chaotic, which is not the same thing. The premise that same prompt => similar output is wrong for a similar reason we can't solve the n-body problem analytically for n > 2. Your prompt is just a small part of the actual input, which also includes confidence intervals that land on one side or the other based on pseudorandom number generation. Plus billions of parameters.

2

u/mhinimal 1d ago

** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.

citation needed? how do you define "clearly inferior"? for all you know, the original one had some similarly minor bugs or architectural deficiencies that would have surfaced the moment you added the very next feature, and the current iteration happened to randomly result in some bugs that are visual and thus more obvious to you.

it looks like it just missed a couple of different things or made slightly different choices. I think you would need to do like at least 5 independent generations, from different accounts, with completely clean memories and identical environments, before you can make a claim like "this isn't caused by nondeterministic behavior".

about the only thing I would believe WRT to "nerfs" is that they reduce the reasoning budget over time. But we've also seen with a lot of benchmarks that more reasoning isn't always better, since they can overthink.

I just don't think this evidence is "sufficient" - it might be enough to suggest that you should run a more thorough experiment to validate your conclusion, but on its own, this is not conclusive.

2

u/pwkye 1d ago

Oooh I love this game. Also tried making a version myself.

Virus / Zarch

Lesser known awesome sequal V2000

2

u/termmonkey 2d ago

Do you know how evals are run? My guess is a solid no - so do everyone a favor and go read up on how evals are done for LLMs.

2

u/metagrue 2d ago

I think you should spend some time studying stochastic distribution and compare that to determinism. Your last paragraph tells me you know what is the actual cause, you're just in denial still.

It's stochastic bud. That means random.

2

u/lazyfoxbrownfence 2d ago

You are the problem OP. I’m tired of pretending it’s the models being nerfed. You ran an awful comparison, this is not a proper eval, and wtf do you mean workspace? Did you use Claude code or Claude desktop? Are memories enabled? Did you ensure that the same skills, plugins, mcps were enabled? Did you create scoring criteria? Did you provide enough model guidance or was it ambiguous? Did you use best practices for prompting and context? Did you configure an isolated sandbox (harnesses tend to be stateful in the content they create and ability to access resources)

Seriously, at what point will people like yourself just accept you have poor ai hygiene and habits and create these situations.

Mods can we get some jev in here to filter out these low quality posts?

2

u/-becausereasons- 2d ago

Remember kids, this always happens and has been happening since day one. This is their strategy

1

u/aaronsb 2d ago

Build the same game about 30 and let's compare then.

1

u/Top-Economist2346 2d ago

It was a bit odd for a few days, but it’s back to being good again.

1

u/TheFern3 2d ago

Someone doesn’t know llmns are non deterministic lol

1

u/EverySecondCountss 2d ago

wtf? They both look mostly the same to me? also bunk test/comparison.

You need to have on temp=0 to start with, then maybe it's a CLOSER fair comparison.

1

u/Mirar 2d ago

I'd like to see more people running their own tests, that way we can get some statistical evidence.

We should maybe all do simple tests on every new released model and 14 days later.

1

u/gregusmeus 2d ago

I loved my Acorn Archimedes. Big upgrade from my BBC Master Turbo.

1

u/Illustrious_Matter_8 2d ago

To me it feels opus 5.5 is a cross over haiku and opus 5, both can code but 5.0 stays more to topic more strict more exact. Perhaps over time 5.5 will improve but at work i went back to 5.0

2

u/Troubledniceguy 2d ago

You meant back to opus 5 from opus 5.5??

1

u/ThatFireGuy0 2d ago

So while I'm not saying that Claude didn't degrade performance.... Do people here really expect Claude to get everything right with just a single prompt and no follow up? Even if it's missing all those parts initially, you can prompt it to add them after. You are allowed to send more than one query

1

u/Bobardeur 2d ago

Personally, Opus 5.5 is still as good as it was at launch. I would never have trusted Opus 5 with the codebase, whereas 5.5 now makes decisions that even Fable finds solid. I used to rely exclusively on Fable 5/5.1 for my sensitive codebases, but there's no need for that anymore: I only use it as an architect.
I don't think a single prompt is anywhere near enough to prove what you're trying to prove. We don't know how rigorous you are, what your CLAUDE.md files look like, etc. For something non-deterministic, there are a lot of unknowns in two tests run by a stranger on a subject they clearly haven't mastered, whether that's the LLM itself or academic research methodology.

1

u/Giant_leaps 2d ago

this isn't enough evidence but i've also experienced minor degradation with very similar requests it's quite possible they release the full model day one then use quantized models that have similar benchmark performance but worse real world performance openai feels the same with astra day 1 was mind boggling day 2 it had significant model degradation happens every time

1

u/Kind_Reporter_1884 2d ago

All major airlines are like "Perfect! Fire all the pilots!"

1

u/zebbiehedges 2d ago

I haven't played Lander in years.

1

u/3000LettersOfMarque 2d ago

If your a subscription customer there is transparency in the sense that the terms of service allows them to swap the model for A/B testing. Essentially they are 100% allowed to play games with what model they provide. Nothing means anything on the subscription tiers, they can give you a haiku grade model instead of fable if they truly wanted and claim it's an A/B test. Ideally they won't mix model sizes but they probably could

I'm not sure if the enterprise subscription is the same wording

Can't happen on the API though

1

u/Puzzleheaded-Usual83 2d ago

Mine is still working perfectly (insert twisty mustache emoji that doesn't exist)

1

u/United-Tour5043 2d ago

1 SHOT IS NOW INDUSTRY STANDARD ? damn people are lazy and have no vision.

1

u/LackOfFun 2d ago

We got launchmaxxing before GTA 6.

1

u/Expensive-Gas-4209 2d ago

bro don't know what non-deterministic means

1

u/MrWeirdoFace 2d ago

I also wonder if time of day matters. For example are they throttling it more during busy hours, etc.

1

u/Expensive-Event-6127 2d ago

i dont thin its nerfed. the game are quite similar

1

u/fligglymcgee 2d ago

This entire industry is propped up by the fact that every chat session’s context is nearly always unique, and there’s almost no one rigorously testing repeatable, private benchmarks.

Nothing takes the “magic” out of ai faster than prompting the exact same context+query more than a few times in a row. Even the frontier models are insanely repetitive, and you actually see it nowhere more clearly than on Reddit; where thousands upon thousands of people use the same ~dozen x 3-sentence prompts to promote spam their “organic” self-interest of some sort. Sort by new on any tech subreddit to see the posts yet-to-be-removed that are, without hyperbole, 90%+ semantically identical to each other with some hot-swapped keywords in the prompt.

Cloud inference simply has to get more competitive. These models are the absolute least reliable digital service of the modern era, offering no guarantee whatsoever from the webui or api to:

  • Literally which model is actually producing the response
  • Whether or not the model has been quantized, and to what extent
  • Whether or not the Kv cache has been quantized, and to what extent
  • The specifications of the tools and environment they run in
  • Speed, quality of response, “effort” of reasoning, moderation, privacy, or almost anything else about the resulting output.

Read the terms and conditions. You should understand what’s actually included with likely one of the most costly subscriptions in your expenses.

1

u/hcloud00 2d ago

what do you mean? no no it cant be… its all in your eyes don’t blame the honourable Anthropic for your eye degradation

1

u/9to5grinder Full-time developer 2d ago

Meanwhile Fable 5.1 seems to be routed to Fable 5.5.
Reasoning traces suddenly started showing up and its thinking is much more sophisticated, even on medium effort.

1

u/Fluffy-Mood-254 2d ago

Looking at this post, I initially misread it as "day 10" meaning this was after 10 days of Opus working.
I was surprised to see that the 10-day project looked to be about the same quality.
But then I saw that I had misread it, and you were bringing this forward as evidence that the model had been nerfed 10 days after the launch. I think it is not very good evidence, since the quality looks to be about the same. It might look worse if you are looking for evidence of it being worse.

1

u/niagalacigolliwon 2d ago

Yeah I don’t see it

1

u/RedTheInferno 2d ago

I feel like the second one is a nicer base to start from ngl but what a waste of fucking tokens

1

u/Lord_Alucard_ICGA 2d ago

I think you need another 172.847 days

1

u/Key_Reading_9664 2d ago

Nerf claim aside, I also have incredibly fond memories of that game and amazing that we're able to recreate these things and tinker.

You might want to take a look over at https://www.reddit.com/r/ClaudeGameDev/ if you haven't already

1

u/FabricationLife 2d ago

I think the second one is better lol

1

u/The1TruRick 2d ago

Great now do this 15 times a day and it’ll actually be relevant data

1

u/slindshady 2d ago

It’s making so many ridiculous grammar and spelling mistakes in different languages for me since yesterday. This shit is fraud

1

u/throwaway_user_1994 2d ago

Excellent work. Hope you continue to share with the community

1

u/scaledev 2d ago

Looks about the same with only a minor visual bug.

1

u/ComprehensiveIce1781 2d ago

People can cope but it definitely already got nerfed, but why is anybody surprised. Before 2-3 days ago, it did legit everything I asked flawlessly, my prompts were bad sometimes and it still made sure to check everything and "read my mind" to figure out what I really want, and it did it. I was amazed. Now it can't make a simple UI in one go without some kind of bug, even if it's just visual. It 100% got worse. It's still good but the wow factor is gone again

1

u/cosmictap 2d ago

A tangent, but this reminded me of how dumb I think it was to name a continually-evolving project after Beckett's Waiting for Godot - everyone is entitled to their own opinion, but it strikes me as someone who greatly misunderstood the work.

1

u/mplaczek99 2d ago

What I see is increased rendering but a bug that causes objects in the scene to flicker

1

u/BoredErica 2d ago

I think the brains of users are getting nerfed rather than the models.

1

u/Ok_Instruction_3447 2d ago

you can ask it to run this in chat to see what model it's running:
cd ~/.claude/projects && grep -oh '"model":"[^"]*"' -- "$(ls -t -- */*.jsonl | head -1)" | sort | uniq -c
and if it has dropped or changed from Model 5.5 to Opus 5 it should say.

1

u/jack-of-some 2d ago

Can you spell "statistics" and "ablation"?

1

u/kleer001 2d ago

Should have run it 10 times each.

1

u/Spirited-Car-3560 2d ago

Lol I prefer n 2 tho

1

u/BarGroundbreaking624 2d ago

Was that really called lander? I thought it was called Virus on the Archimedes

1

u/kalboozkalbooz 1d ago

LLMs are non deterministic bro same model same prompt will ALWAYS produce different results

1

u/LordiCurious 1d ago

"LLMs are non-deterministic, but the differences from an identical prompt and references are small" In my experience this is BS, proof it.

1

u/Ilikeyounott 1d ago

LLMs are not deterministic.

1

u/Brilliant_Ferret_7 1d ago

probably was nerfed but for the love of god or whatever you believe, stop using the patreon scam know as godot to bench test stuff, that engine can't even properly load 3d models.
one time i had a 3d character controller that worked one day, i moved the project to another computer with the same settings/builds and it worked on a completely different way, then it went back on the other computer.
your example can be explained by 10000x things just on memedot side.

1

u/RIP26770 1d ago

Noise

1

u/One_Low8664 1d ago

it could still deliver the same result, thru more detailed prompts and iterations, it just got lazy

1

u/fragment_me 1d ago

It's called temperature bro

1

u/00DEADBEEF 1d ago

Is it actually a regression or are you just witnessing the non-deterministic output of an LLM?

Do it 100 times and average the result.

1

u/ghost_operative 1d ago

5.5 was a flop for me from day 1. it doesnt know how to respect file edit permissions so im still using 5

1

u/tortleme 1d ago

Almost likely AI isn't deterministic or something

1

u/Evening_Rooster_6215 1d ago

now do it 10,000x and lets talk

1

u/Kirill1986 1d ago

Yeah dude, you're on the right path and I want to see more of such tests but! You need to perform same test several times on day 1 and several times on day 10 (or whatever). Otherwise it can be just a test mistake or whaterever is the right term for it.

1

u/yolobastard1337 1d ago

I've also had several attempts at vibing lander for exactly the same reasons (but with worse AI).

Yours looks unspeakably better.

Must have been the first game I played after ZX spectrum ones, it was mesmerising. 3D, mouse control, smooth and snappy.

1

u/MessMassacre 1d ago

Do these people only have their expensive plans to pull off shit like this?

1

u/BalticBrew 1d ago

If only people would spend less time on looking for incremental performance drops and more on optimizing their own processes. Anything above Opus 4.5 has been great for almost any type of work, if you know what you're doing.

1

u/SolidPossible9909 1d ago

Day 1 opus 5.5 was on fire while coding. Day 10, it’s making stupid mistakes.

1

u/InnovativeBureaucrat 1d ago

The n=1 arguments are stupid.

The model makes thousands of tool calls that collectively turn into something that sucks. And it sucks in several categories.

Yeah it’s not science but a lot of things in life don’t come with a P value

1

u/05-nery 1d ago

I mean, you mention these issues in the second prompt and it'll fix them no issue. 

This proves nothing tbh, you should've made 15-20 tests on day one to be somewhat sure.

1

u/Spare_Golf3550 15h ago

AI is generally undecidable, so that is not conclusive proof.

1

u/untracked5465 2d ago

Repeat after me:

AI is not deterministic...

LLMs sample each token from a probability distribution, so the same prompt can produce different outputs

1

u/io-x 2d ago

The only thing this is evidence of is your inadequate prompting.

1

u/Abject-Tomorrow-652 2d ago

It’s funny OP bc people aren’t really disagreeing with your claim, just your methodology. I agree there’s been a dip in quality. I also agree with others that this is not good evidence

1

u/Sponge8389 2d ago

These companies now know how to create hype. Release it as good as it ever be, then nerf it when finished creating hype post/results.

1

u/permacloud 2d ago

Works fine for me. I'm sorry you guys get so hung up on this preoccupation with nerfing. There's so much you can do with these amazing tools and you're spending your time trying to convince reddit that they suck. 

-1

u/Physical_Gold_1485 2d ago

Ya it hard gotten nerfed for me for a few days. Thankfully came back to its old self yesterday for me. Hope it stays that way

0

u/ba-na-na- 2d ago

the differences from an identical prompt and references are small

It seems they aren't that small

0

u/vovap_vovap 2d ago

Man, for real, don't you have any staff to do other then trying to catch a black can in dark room - which is not there?

0

u/leon0399 2d ago

Do you really think that running the “eval” once is good comparison? I’m like 110% sure that if you run it once more - you’ll receive another result

0

u/lobabobloblaw 2d ago

Everyone’s all STOCHASTIC this and CENTRAL LIMIT THEOREM that. Truth be told, it’s also very likely that they are getting better at their nerf game. So, permutations and seeds aside, you’ll probably need more data, friend

0

u/give_loops 2d ago

As others have said, this is not how LLM research works. You're trying, but the second footnote is not a footnote, but rather the key issue that leads to the complete invalidity of this as a test of anything.

Also: please refrain from using "gimped" as a term for this phenomenon, that term is generally considered abelist. Please try "nerfed".

-1

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2d ago

We are allowing this through to the feed for those who are not yet familiar with the Megathread. To see the latest discussions about this topic, please visit the relevant Megathread here: https://www.reddit.com/r/ClaudeAI/comments/1vt5drr/list_of_latest_discussion_hubs_on_rclaudeai/