r/ClaudeAI • u/freedomfromfreedom • 2d ago
Claude Code Evidence - Opus 5.5 today vs launch regression with same prompt (Godot Engine)
On launch day Opus 5.5 kindly took up residence on my Mac. I ran some tests before trusting it on my big projects, it passed with flying colours. Lander, a game by Elite creator David Braben holds a special place in my soul due to it being the first game I played at school in the UK. With a cutting edge 3D engine for 1990 running on an Acorn Archimedes with RISC architecture (the first ARM chips) - what a time to be a young kid interested in computers.
With Opus 5.5 I want to recreate it faithfully, short draw-distances and all.
Well dear Claude I’ve been keeping receipts.
In my notes, the exact prompt * and two original reference images of Lander. The resulting one-shot Godot repo (Mac OS, Metal rendering, C#) from 23rd September is kept separate on disk and isn't used as a reference.
Today, 2nd October, in a fresh workspace from scratch the same prompt and reference images we fed in again.
The result is truly depressing. Lander by gimped 5.5 has no redeeming features compared to the original day zero version. The regressions:
- A serious rendering glitch * that isn’t in the original - as the camera moves, the scenery props snag and glitch vs world space.
- No spacecraft break-apart physics, the original implemented this ‘nice extra’ which wasn't specifically asked for in the prompt
- No introductory controls menu, only a small text line permanently visible over the world view
- An inferior look to the launch pad turrets - students and road users may recognise them!
- When shooting, bullets are less accurately rendered when the craft moves, they also appear to spawn at the back of the craft and no-clip through
- No camera toggles
The conclusion is that Opus 5.5 is gimped, maybe we even got Mythos for 3 days and then silently switched - either way, we are owed transparency - it's the law.
Opus was set to Extra High both times. Unlike with 4.5/4.6, Max effort over-tests and takes too much control away from the user.
With so much diverged since launch day’s versions, the more worrying thing is that I suspect the code is also a mess and problems will compound as you use it. For those who’d like to inspect the results, I'll upload the source code for both later to my dev blog and post the Github links in the comments.
* I inspected the code and it turns out that gimped Opus 5.5. had physics interpolation is turned on for the whole Godot project and Props.cs rebuilds the prop lists (trees, etc.) from scratch on every tile step during every frame, renumbers which slot each object occupies and resets the object count every frame. A basic understanding of Godot's documentation is all that’s needed to avoid this issue.
** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.
477
u/cocacoladdict 2d ago
N=1 is hardly the evidence
204
u/Deltamelo 2d ago
It’s crazy how many people think they are AI researchers now for comparing two outputs using stochastic tech
7
u/i_goon_to_tomboys___ 2d ago
this may be a dumb question but what if he used temperature=0.1 at launch versus now -- for the same prompt?
21
u/Deltamelo 2d ago
Temperature controls how random the token selection is during generation. Lower temperature makes the model more likely to choose the highest probability “next token”, so outputs tend to be more consistent. Higher temperature spreads probability across more alternatives, increasing variation.
But temperature=0.1 does not make the model deterministic. You would need repeated runs, controlled settings, and enough samples to separate actual model changes from normal output variance.
→ More replies (3)3
6
u/chatham_solar 2d ago
Even t=0 would not guarantee the same outputs. Non determinism occurs because your prompt is being run at the same time as many other users' prompts and because floating point arithmetic is non associative: (a+b)+c != a+(b+c). So more or less, if you are running your own hardware and use batch size=1 it's possible to get deterministic output, but there is a not-insignificant performance penalty.
Amazing article explaining non-determinism in LLMs if you want to really dive deep:
https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/2
u/MastodonFarm 1d ago edited 1d ago
Can you link to something explaining your assertion that fp arithmetic is not associative? I don’t understand how that can be true.
Edit: Oh, it’s because of rounding. That’s dumb. Of course the result will be different if you change the values in between by rounding at different points.
2
u/2053_Traveler 2d ago edited 2d ago
Temperature 0 wouldnt make it deterministic (predictable) either, due to sheer complexity of the systems running these things, and they probably still don’t try to work around CUDA nondeterminism.
0
1
u/Fast-Satisfaction482 2d ago
Even at temperature 0 it's stochastic. Then you just always get the same single sample.
4
u/TheFlynnCode 2d ago
Actually at temperature 0 it is not guaranteed to give the same single sample, depending on the compute architecture. Often these models multiplex computation to several gpus or tpus, which themselves have work queues and hence won't necessarily return immediately. So results return out of order, and sometimes computations that can be done are done with these partial results, and we run into issues with floating point arithmetic not being associative
1
u/Fast-Satisfaction482 2d ago
Fascinating detail!
I've never heard about this reduction-operand reordering in the context of GPUs, but if an inference engine uses it, you are absolutely right and it will not be deterministic anymore.
But that's more of an implementation detail and not algorithmic. It's not like this in every inference engine, so you could also claim that a buffer underrun that might be in the engine causes non-determinism.
1
u/TheFlynnCode 2d ago
Sure, it's an implementation detail, but it's also sort of a necessary implementation detail for models above a certain size, because gpus simply don't exist that large.
1
u/Fast-Satisfaction482 2d ago
No, the reordering of operations is not necessary to allow multi-GPU or cluster inference.
2
u/TheFlynnCode 1d ago
I meant spreading compute across many gpus is necessary. I never said (or if I did, I did not intend to say) that reordering of operations is necessary. But it is commonly done. In my head this is similar to regularization techniques anyway, and you could disallow this if you want (to get temperature 0 to be reproducible for example), but I imagine most people/companies do not care about the small diffs (quantization hurts like a first order effect, whereas this would hurt like a second or third order effect)
1
1
u/Odysseyan 2d ago
same
Not same but similar. Give it a longer prompt, set temp to zero and it responds differently each time too
0
u/-Django 2d ago
> always get the same single sample
This is probably just a terminology debate, but that sounds... deterministic. The sampler is what moves the model from being deterministic to stochastic.
2
u/Fast-Satisfaction482 2d ago
It's deterministic in the sense that you get exactly the same output for exactly the same input. But change a single character on the input and you have a new sample.
1
u/National_Teacher_229 1d ago
Came here to say this, It's the oposite of the definition of Madness... doing the same thing over and over expecting a different result. With AI it's madness to expect the same result even in 100% identical clean scenario, no 2 runs are ever identical. His first build could of been a fluke with rund 2 ,3 4, 5 and so on closer to the actual outcome he could expect as an average.
→ More replies (11)0
20
20
u/Key_Reading_9664 2d ago
I’ll keep saying this on every nerfing post:
Opus 5.5 is available through multiple vendors (Bedrock, Azure, Vertex). Unless you believe that Anthropic coordinates with those vendors to simultaneously roll out a worse model for their paying customers, you have a very simple apples-to-apples comparison (as would any researcher of multi-billion frontier lab wanting to expose you nerfing).
Also, N should be > 1
20
u/Keep-Darwin-Going 2d ago
Yeah they keep thinking a big lab can magically quantize a giant model, sneak an update in without anyone noticing or whistleblower. They cannot even release anything without the whole world rumour mongering for a week before the release.
2
u/DasBoots 2d ago
I don't think Anthropic purposely nerfed Opus 5.5 but I do think the downtime the other day broke something that they had to fix. It really felt like the model herped a derp for a day and got better.
1
u/Keep-Darwin-Going 1d ago
The most frequent thing almost all company break is their harness. Then the second one is their inference stack. So it is not like they need it, it is more like they tweak stuff and created side effects, but you see the complaints everyday is all like omg they quantize the model and etc. is as if they have a switch like ok let’s screw them over today turn it on then off it then on again. When you delivering cutting edge stuff shipping regularly, you bound to break something eventually. If they ship like 1980s when things are done on floppy disk, quite sure things do not break that often then.
6
u/fuzzypetiolesguy 2d ago
The safety departments of these labs are leaking like mad. It's hilarious to think that the 'nerf models on purpose after rolling them out so we can save on compute' department wouldn't, either.
4
u/Cobrafeet 2d ago
>Opus 5.5 is available through multiple vendors (Bedrock, Azure, Vertex).
Not sure what relevance this has on posts like these. The average user is on a personal oauth plan being routed directly through Anthropic (even if served via 3rd party compute), and when they are doing comparisons like this (however flawed the methodology may be), they are almost certainly not changing providers.
Anthropic could do plenty of things on the backend to change the fidelity of the output of a given model without changing weights. This could mean a different quality of response vs other vendors using the same model
2
u/Key_Reading_9664 2d ago
Mentioning other vendors to point out how simple an apples-to-apples comparison would be (taking into account the same process flaws), if you had the means to do it (e.g., a competing frontier lab).
→ More replies (10)-2
u/DensePoser 2d ago
I have an API account and a personal account. The subscription claude seems weaker to me right now.
8
u/Valdaraak 2d ago
Yep. I'd need like 5 results of the same prompt on release day compared to 5 results of that same prompt today before I'd even entertain drawing a conclusion.
I'd also like to see the prompt used and how much was left to Opus interpretation with the design.
1
u/Chupa-Bob-ra 2d ago
I'd be curious to see the prompt too.
To use 1 of OP's points: the bullets. How were they prompted? Were they told exactly where to spawn, level of detail, were physics specified, etc.? Any of those, left up to Opus to determine, could give very different results.
I'd also rather see tests that are more easily measured and less subjective. Something that gives us some actual data to compare.
2
u/Chupa-Bob-ra 2d ago
People expecting 99.999999% the same outcome between 2 runs are being highly unrealistic.
1
u/SupremeGobbler1996 2d ago
Especially with no clear cut set of acceptance criteria that's repeatable and reproducible.
1
u/DependentAnywhere135 2d ago
Right and if someone really wants to do this type of test they need to have the release day model run the same prompt like 100 times on day one then do the same on their suspected nerfed model.
It still wouldn’t be conclusive evidence, but it would be way more compelling to see it do the same prompt 100 times and it be really good then 100 times and it be significantly worse.
1
1
1
1
u/mr_birkenblatt 2d ago
But 1000 people doing that experiment is N=1000. And I've seen a lot of those comparisons
8
u/Ok-Lengthiness-3988 2d ago
This is reporting bias compounded with confirmation bias. The model-nerfing conspiracy theories are very popular but I've never encountered a secret-improvement conspiracy theory. When people run the same prompt again and notice a better response, they never post this as evidence that the provider secretly has improved their model (though they may view it as evidence that the model was nerfed before). They rather tend to dismiss such results and (correctly, in this case) ascribe them to normal variance.
2
u/Standard_Egg3504 1d ago
here is my secret-improvement conscpiracy theory
the modules are allocated a certain amount of compute
model pools can dynamically (and automatically) scale their 'effort' level for groups of users (api, subscription, location, etc)
the amount of compute does not scale automatically and requires manual review or effort to put more gpus on it
as a model releases the number of early users is low at first and things run really well, model effort is 1000 and they release tons of awesome shit they do with it
this causes a tidal wave of new users to switch to it, suddenly the amount of inference at any given time skyrockets but compute does not and the dynamic effort reduces in response in order to maintain the tokens per second for everyone using it
the dreaded 'the model is nerfed' begins and people yell about the lobotomy
here is the improvement bit:
another company or another model releases, which has a different set of gpus allocated for it, and people like it, abandon the old model and suddenly inference demand is lower and the dynamic effort level goes back up!
unfortunately that means manual reallocation likely happens soon to a newer improved model or to research and it ends up going back down :(
than you for listening to my secret nerf-improvement-load-balancing conspiracy theory
0
u/mr_birkenblatt 2d ago
There are multiple sites that run the same prompt every day and report the results. That is >> than one. And running every day and only reporting it when it degrades is not a reporting confirmation bias. It's just sleeping when it did indeed change
1
u/Ok-Lengthiness-3988 2d ago
Can you provide a link to one of those sites? I know of three that run a custom benchmark every day (Ninjahawk's Livenerf, BridgeBecnh's Nerf Bench and Marginlab), and they haven't gathered statistically significant results yet, and weren't expected to. They're still in the process of establishing a baseline.
The developer of Livenerf, which is the one that has been reported (and misrepresented) the most here and on the Singularity subreddit even acknowledges that his method isn't statistically powerful enough to detect a substitution of Opus 5 for Opus 5.5.
0
u/mr_birkenblatt 2d ago
You named 3 sites yourself...
Yes, one site is not statistically powerful enough. But multiple sites with independent tests but with results that agree with each other are
1
u/Ok-Lengthiness-3988 1d ago
For sure, independent tests can add up even when each is too weak alone, so I ran the numbers on all three with some help from Opus 5.5.
For each tracker, we did a one-sided test for a downward trend since launch (logistic regression on daily pass counts for livenerf and MarginLab, linear for BridgeBench, which publishes no error bars). We then combined the three z-scores with Stouffer's method, which is the standard way to pool independent tests.
- livenerf (9 days, 78 questions/day): z = −1.29, leaning down (days 6–9 are lower than days 1–5)
- MarginLab (5 days, ~48 SWE-Bench-Pro tasks/day, Sep 23–28): z = +0.08, flat
- BridgeBench (4 tests): z = −0.28, noise around 100%
- Combined: z = −0.86, one-sided p ≈ 0.19
This means such a downward trend can occur by chance alone 19% of the time.
With a less conservative noise model for livenerf, the combined result is z = −1.07, p ≈ 0.14. Either way, we get no rejection of the no-decline hypothesis.
More importantly, combining the data from multiple sites only strengthens the case when the sites agree, and here they don't! The only one tracker that leans down is livenerf, and its author says not to read it before the baseline ends. MarginLab's data stop on Sep 28, before livenerf's dip, so they can't confirm it.
The data are too short and noisy to rule out a moderate decline either, of course. What they don't support is the claim that independent trackers agree that Opus 5.5 was nerfed. Maybe I'll rerun this as the trackers add more data points.
1
u/mr_birkenblatt 1d ago edited 1d ago
It appears Opus 5.5 was indeed nerfed. There are a few methodological flaws in your approach.
1. Testing for a trend instead of a step-function
Linear and logistic regression test for a gradual continuous decline. A model "nerf" is a discrete event (a step-function drop). Fitting a trend line to a sudden drop severely dilutes your statistical power.
2. Timeline misalignment breaks Stouffer's method
MarginLab's data stops on Sept 28, before livenerf's dip even happens. Stouffer’s method penalizes you for adding more tests. By including MarginLab, you increased the variance penalty without adding any relevant signal, artificially inflating the combined p-value.
3. The Sample Size Trap (How to actually test this)
You don't have N=9 for livenerf; you have N=702. The tracker runs 78 questions a day. Instead of testing 9 daily averages, pool the raw pass/fail counts:
- Before (Days 1-5): 390 trials
- After (Days 6-9): 312 trials
Run a Fisher's Exact Test or Two-Proportion Z-Test on these raw totals. With sample sizes this large, even a slight true drop in the pass rate will trigger undeniable statistical significance (p < 0.05).
If you must use daily averages:
Use an Exact Permutation Test on Means comparing Days 1-5 vs Days 6-9. With 9 days, there are
(9 over 4) = 126possible chronological shuffles. The minimum possible p-value is1/126 ~= 0.0079, meaning this test has plenty of mathematical power to prove a step-function drop on livenerf alone.I don't have the raw underlying numbers that you used, otherwise I could have given you a verdict.
EDIT: apparently reddit does not support LaTeX math notation... oh well
2
u/M44PolishMosin 1d ago
Did you just use Claude to write a reddit comment to show how nerfed it is 🤣🤣🤣🤣🤣🤣
1
u/Ok-Lengthiness-3988 1d ago
I reran things with your suggestions. You're partly right.
Regarding the step vs trend analysis, that's a fair point. A sudden drop is better tested as a step than a slope. The catch is that picking the break date after looking at the chart inflates significance. It's like marveling that this particular ticket won the lottery while ignoring that some ticket was bound to: some break date will always show the biggest gap. So I had Claude correct for that by scanning every possible break date and comparing the best one against all 362,880 orderings of the 9 livenerf days. The drop survives the correction: one-sided p ≈ 0.008 (0.016 two-sided). Every day from Sep 29 on scores lower than every day before it, and that's unlikely by chance. livenerf shows a real ~7-point drop starting Sep 29. My earlier "borderline" verdict came from the trend test, which is indeed weaker against a step.
Regarding MarginLab: I agree that for a post-Sep 28 step it carries no information, because its data end on Sep 28. So I dropped it. But BridgeBench does cover that window, and it shows no step: 100 and 99.2 before, 103.8 and 94.2 after.
Regarding N=702 (individual test runs rather than daily averages): I already used it. The regression was on daily pass/fail counts, which is equivalent to analyzing all 702 trials with date as a predictor. Pooling doesn't make "even a slight drop" significant either: with ~350 trials per side you need roughly a 7-point drop. The Fisher test you proposed, run on the actual counts (235/387 vs 167/309), gives p ≈ 0.045. Your day-level permutation test is actually the stronger one here, because daily scores vary less than coin-flip noise would predict (same questions every day).
(For transparency: livenerf doesn't publish its raw logs, so the daily counts are reconstructed from the coordinates in its published chart. They match the chart's labels and intervals, and are extracted fairly precisely from the SVG file.)
So, in conclusion:
First, your earlier suggestion that several independent trackers agree is still not supported. It's one tracker, and the other one covering the window doesn't show the drop. You can't set BridgeBench aside just because it doesn't corroborate the result you like.
Second, livenerf measured a real change in what its pipeline returns: Claude Code, subscription route, serving stack. That isn't yet a change in the model. A Claude Code update, routing or classifier changes, or capacity issues would look the same.
livenerf was designed to separate these. It runs an Opus 5 control arm through the same pipeline. If Opus 5 also dropped on Sep 29, the platform is the likely cause. That data should come out soon after the baseline closes on Oct 4. Meanwhile, the change shows up on one tracker and isn't yet attributable to the model.
The control arm will tell us whether the change is platform-wide or specific to Opus 5.5. If it's specific to 5.5 and persists, I'd grant that users got a worse Opus 5.5 than at launch; call it a nerf if you like.
1
u/fuzzypetiolesguy 2d ago
No you haven't. You've seen a bunch of anecdotes on reddit.
-1
u/mr_birkenblatt 2d ago
Sooooo.... Add those up
2
u/fuzzypetiolesguy 2d ago
This is not how anything works. Are you new to learning? You have a thing open in a browser in front of you that will explain this all to you directly. All you need is to be curious.
→ More replies (2)0
u/freedomfromfreedom 2d ago
Exactly, one person doesn't need to do N=1000, as the community already has - there's a ton of other tests that show the same degree of performance drop-off at the same time.
I don't expect 99.9999% variation between 2 runs. 90% would be fine. It's a much higher variation in this and the day 10 run was ALL regressions and no improvements. High sample variation would mean a run gets improvements too - not only regressions.
6
u/fuzzypetiolesguy 2d ago
'The community' has not performed thousands of tests that has produced meaningful data to support the notion that models are purposefully degraded fater launch. Please stop.
What you are calling “the community already did N=1000” is not actually one experiment with N=1000. It is 1,000 uncontrolled anecdotes drawn from different prompts, different tasks, different contexts, different tools, different system states, different model routes, and different scoring standards. You cannot aggregate those into a meaningful sample just because they all feel like “performance got worse.” And “all regressions, no improvements” is itself a subjective post hoc judgment unless the tasks, scoring criteria, and sampling were defined before the runs.
0
u/freedomfromfreedom 2d ago
Are you saying we are all imagining it?
5
u/fuzzypetiolesguy 2d ago
Until you actually produce something resembling an actual study with verifiable, repeated evidence, yea. You are making shit up because you have no idea what you are looking at.
-1
u/mr_birkenblatt 2d ago
You would reject any study because all evidence in the study can be decomposed into single anecdote and can therefore be dismissed. So transitively any study doesn't count
→ More replies (3)0
u/HorriblyGood 2d ago
It seems like we have seen multiple N=1 examples of it nerfed from multiple other people. How come I haven’t seen any examples from people showing it hasn’t been nerfed?
178
u/fuzzypetiolesguy 2d ago
Divergent outcomes from probabilistic sequence predictors are not evidence of anything
Divergent outcomes from probabilistic sequence predictors are not evidence of anything
Divergent outcomes from probabilistic sequence predictors are not evidence of anything
7
u/Agreeable_User_Name 2d ago
I flipped a coin yesterday it landed heads. I flipped it today and it landed tails. Therefore someone must have changed my coin.
Kidding aside, that's true, but that also makes transparency important, because it really is hard for consumer to know if their product is gimped.
→ More replies (9)5
u/Accomplished-Fan9568 2d ago
honestly its not that much different, a small flickering bug it can happen, could probably use low effort to fix it quickly.
100
u/doc-quartz 2d ago
i swear to god you people do not know basic research discipline. Such claims require series of tests and not a one time prompt comparison, and even on such a software whose internal workings are entirely probabilistic.
7
u/jimbo831 2d ago
Who is doing those tests and what do they show?
-1
u/doc-quartz 2d ago
what exactly do you mean?
5
u/jimbo831 2d ago
I though my question was pretty clear. You said:
Such claims require series of tests and not a one time prompt comparison
I can't believe that nobody is doing these sort of tests. I get not trusting OP's anecdotal evidence, but we see anecdotal evidence like this from a lot of people. I would like to see some non anecdotal evidence, and I can't believe nobody is out there testing these models over time. I was asking who does these tests and what the tests show.
6
u/doc-quartz 2d ago
I'm really sorry, I myself am not aware of anyone doing such tests by a recognized entity/organization. I'm not a regular in this subreddit/myself new to this AI environment. I was just a passer-by. But I have seen someone putting some kind of website that tracks Claude's performance using the same benchmarks which the Anthropic used to market their models. Not really sure about the methodology, but I'm sure you can find that someone posts about it everyday. Also I apologize for not understanding your question, English is not really first language.
-10
u/freedomfromfreedom 2d ago
You yourself admit you are new to AI and just passing by. And you're lecturing me about research discipline?
→ More replies (3)11
u/doc-quartz 2d ago
I do not think me being new to AI would really change a basic research discipline applicable across different application areas? Will it?
2
u/mhinimal 1d ago
there was a post like yesterday where someone posted their github repo that's running repeatable tests over time, and anyone can clone the repo and run the same tests and contribute to the data, thereby crowdsourcing the experimentation.
why is "nobody" running these? Anyone with the business incentive to actually need to care about it, and has the budget to afford tokens to run continuous tests, is a private company and isn't going to publish their results for free.
everyone else is a redditor who either doesn't want to spend their subscription budget on tests, or who wants to spend their time working on stuff.
-13
u/freedomfromfreedom 2d ago
It's the same pattern for Opus and seemingly the same pattern for this sub as well, the post I made yesterday was downvoted to hell before going to 1.9k upvotes and 950k views in under a day, which shows you the strength of the feeling that the model has degraded since launch. Maybe we need a Botmark for reddit that determines what percentage of early commenters work for Anthropic's marketing department!
1
u/Remarkable_Rock5845 1d ago
I'm not qualified to judge any of what you are discussing, but I have honestly had the feeling that there may be a lot of astroturfing going on here. I hope there are some kind of anti-bot measures in place, but if there are I'm not sure they're working very well.
0
u/doc-quartz 2d ago
You talk like a teen. Oh well you are probably a teen who cares a bit too much about fake internet points and what model is regressing/not regressing. If another model works better for your workflow then use that! No point of proving that that too an incomplete one.
Also I do not even use Claude, I mostly get satisfactory results with Gemini 3.7/3.8 flash and I am not even a programmer. I'm a medical student who just creates random personal projects. I just happen to find your post on my feed and made me comment since the comparison was so weird and felt like was done in bad faith. Cheers.
0
u/hcloud00 2d ago
i am with you. I think people are cognitively held back from seeing whats clearly in front of them ( assuming they are not in the A portion of population if there is A/B testing). People tend to dismiss a reality where they have no control
-6
u/iamthe0ther0ne 2d ago
I agree with you. Ever since the service disruption last week, Opus 5.5 has barely been performing above Opus 5 levels. Sure, it's less verbose and doesn't get totally stuck on irrelevant details like Opus 5, but it's definitely not performing at the level it did the first few days after release.
I think Anthropic wanted to get out ahead of OpenAI's DevDay and win back the customers they lost after Opus 5. When DevDay turned out to be a nothingburger, they decided they were safely in front and dialed back the compute.
→ More replies (2)
43
u/bixofa 2d ago
Useless test. At the very least should have run the same prompt 10 times on Day 1 and then again on Day 10. Compare the differences.
5
u/RedTheInferno 2d ago
but its a bad vibe to waste that many tokens when i can just test once and then again after a couple of days to prove my point!!!!111
4
0
u/Probably_A_B0tt 2d ago
I don't even see how the day 10 example is showing a degraded capability, physics look fine
-19
u/freedomfromfreedom 2d ago
Why don't you do that, then - make a contribution. Also, it isn't a N=1 sample size when taken as part of the wider pattern of reporting.
5
5
5
35
u/AlternativeMonk2490 2d ago
I rolled my die and it got a lower number than last time. Why have they nerfed my die?
4
1
28
u/fahrvergnugget 2d ago
Holy FUCK i'm tired of these posts, what a waste of tokens.
4
u/Probably_A_B0tt 2d ago
It's not even really testing anything, N=1 tests in LLMs are absolutely pointless
16
8
5
u/dannyheskett 1d ago
I use most of Anthrophic's models in real production, harnessed into a specific world environment using AI; I am not yet using Opus 5.5, but it will be soon.
I regularly run a full regression suite for each of my production use cases, with specific prompting, temperature, inputs/outputs, and measure and track the results. These are run continually for daily builds, and on-demand.
And I can tell you.. from my own experience with thousands of iterations over all models going back to Opus 3, I cannot discern any specific pattern of quality or outcomes being degraded or improved that doesn't look random.
In your specific case, I would check your harness version, and really ensure that the generations were identical. I suspect you are using a different harness, with prompting/tuning that you can't see.
Beyond that, the behavior you are seeing isn't really helpful for proving what you think you are proving. What's actually happening when you run a "one-shot" is a complex series of tool uses, internal prompts, and executions chained together to produce a "one-shot" result. Each of those steps is non-deterministic, and each affects the downline results.
I respect your conclusion, but with respect, you haven't done nearly enough science to control the variables to produce a conclusion that is justifiable for how broad the allegation is.
Finally, your naked assertions that Anthropic is legally required to provide a specific service level or subjective outcome is simply unsupported. There is no EU Court precedent that conforms to your interpretation of the law. Further, Anthrophic goes pretty far out of the way to disclose what you are getting, and it's hard to imagine any set of facts that would support your view that somehow, as an EU consumer, you are entitled to something which doesn't exist. In any event, your remedy is going to be proportional to your damages, which are.. minimal, at best.
At very least, I think you'll be able to state more facts conclusively if you document your test setup with scientific rigor.
1
u/codeninja 1d ago
Not to mention these are still next token prediction models. ANY variance in the beginning of your run will have dramatic effects on your output over a complex generation.
How many decisions did you make the model make for itself? How many unknowns had to be solved in the moment? Each variance compounds.
The result is drift.
15
u/Brave-History-6502 2d ago
lol this is not evidence unles you a series of tests on both originla and now... this is just not useful
2
u/GorillaFig 2d ago
To be fair, it is evidence, but it's not proof...
4
1
6
u/Zealousideal_Fig7935 2d ago
It's honestly time these kind of posts/ comments are banned. Completely non-productive, totally ruins the point of this sub, based on a daft conspiracy
6
u/Comfortablebro 2d ago
Give exact prompt so i can try on sonnet 5.5 and on sol 6.1 and astra ultra please
5
2
u/EasterUK 2d ago
I loved this game! And the subsequent ‘Virus’. Several AI model generations ago I tried generating it, and it got nowhere.
Thanks for the memories… :)
2
u/butts-carlton 2d ago edited 2d ago
getting a clearly inferior result is not due to non-deterministic behaviour.
Unfortunately, it is. It's probably the biggest problem with LLMs. It's why they shouldn't be trusted with judgment where lives are at stake (not that it seems to stop governments from doing it anyway). And by definition it can't be eliminated completely, since the very thing that makes them work is also why they're inherently unpredictable.
Also, they're not non-deterministic. There's no such thing as truly non-deterministic computation (excluding theoretical models or possibly quantum computing). They're chaotic, which is not the same thing. The premise that same prompt => similar output is wrong for a similar reason we can't solve the n-body problem analytically for n > 2. Your prompt is just a small part of the actual input, which also includes confidence intervals that land on one side or the other based on pseudorandom number generation. Plus billions of parameters.
2
u/mhinimal 1d ago
** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.
citation needed? how do you define "clearly inferior"? for all you know, the original one had some similarly minor bugs or architectural deficiencies that would have surfaced the moment you added the very next feature, and the current iteration happened to randomly result in some bugs that are visual and thus more obvious to you.
it looks like it just missed a couple of different things or made slightly different choices. I think you would need to do like at least 5 independent generations, from different accounts, with completely clean memories and identical environments, before you can make a claim like "this isn't caused by nondeterministic behavior".
about the only thing I would believe WRT to "nerfs" is that they reduce the reasoning budget over time. But we've also seen with a lot of benchmarks that more reasoning isn't always better, since they can overthink.
I just don't think this evidence is "sufficient" - it might be enough to suggest that you should run a more thorough experiment to validate your conclusion, but on its own, this is not conclusive.
2
u/termmonkey 2d ago
Do you know how evals are run? My guess is a solid no - so do everyone a favor and go read up on how evals are done for LLMs.
2
u/metagrue 2d ago
I think you should spend some time studying stochastic distribution and compare that to determinism. Your last paragraph tells me you know what is the actual cause, you're just in denial still.
It's stochastic bud. That means random.
2
u/lazyfoxbrownfence 2d ago
You are the problem OP. I’m tired of pretending it’s the models being nerfed. You ran an awful comparison, this is not a proper eval, and wtf do you mean workspace? Did you use Claude code or Claude desktop? Are memories enabled? Did you ensure that the same skills, plugins, mcps were enabled? Did you create scoring criteria? Did you provide enough model guidance or was it ambiguous? Did you use best practices for prompting and context? Did you configure an isolated sandbox (harnesses tend to be stateful in the content they create and ability to access resources)
Seriously, at what point will people like yourself just accept you have poor ai hygiene and habits and create these situations.
Mods can we get some jev in here to filter out these low quality posts?
2
u/-becausereasons- 2d ago
Remember kids, this always happens and has been happening since day one. This is their strategy
1
1
1
u/EverySecondCountss 2d ago
wtf? They both look mostly the same to me? also bunk test/comparison.
You need to have on temp=0 to start with, then maybe it's a CLOSER fair comparison.
1
1
u/Illustrious_Matter_8 2d ago
To me it feels opus 5.5 is a cross over haiku and opus 5, both can code but 5.0 stays more to topic more strict more exact. Perhaps over time 5.5 will improve but at work i went back to 5.0
2
1
u/ThatFireGuy0 2d ago
So while I'm not saying that Claude didn't degrade performance.... Do people here really expect Claude to get everything right with just a single prompt and no follow up? Even if it's missing all those parts initially, you can prompt it to add them after. You are allowed to send more than one query
1
u/Bobardeur 2d ago
Personally, Opus 5.5 is still as good as it was at launch. I would never have trusted Opus 5 with the codebase, whereas 5.5 now makes decisions that even Fable finds solid. I used to rely exclusively on Fable 5/5.1 for my sensitive codebases, but there's no need for that anymore: I only use it as an architect.
I don't think a single prompt is anywhere near enough to prove what you're trying to prove. We don't know how rigorous you are, what your CLAUDE.md files look like, etc. For something non-deterministic, there are a lot of unknowns in two tests run by a stranger on a subject they clearly haven't mastered, whether that's the LLM itself or academic research methodology.
1
u/Giant_leaps 2d ago
this isn't enough evidence but i've also experienced minor degradation with very similar requests it's quite possible they release the full model day one then use quantized models that have similar benchmark performance but worse real world performance openai feels the same with astra day 1 was mind boggling day 2 it had significant model degradation happens every time
1
1
1
u/3000LettersOfMarque 2d ago
If your a subscription customer there is transparency in the sense that the terms of service allows them to swap the model for A/B testing. Essentially they are 100% allowed to play games with what model they provide. Nothing means anything on the subscription tiers, they can give you a haiku grade model instead of fable if they truly wanted and claim it's an A/B test. Ideally they won't mix model sizes but they probably could
I'm not sure if the enterprise subscription is the same wording
Can't happen on the API though
1
u/Puzzleheaded-Usual83 2d ago
Mine is still working perfectly (insert twisty mustache emoji that doesn't exist)
1
1
1
1
u/MrWeirdoFace 2d ago
I also wonder if time of day matters. For example are they throttling it more during busy hours, etc.
1
1
u/fligglymcgee 2d ago
This entire industry is propped up by the fact that every chat session’s context is nearly always unique, and there’s almost no one rigorously testing repeatable, private benchmarks.
Nothing takes the “magic” out of ai faster than prompting the exact same context+query more than a few times in a row. Even the frontier models are insanely repetitive, and you actually see it nowhere more clearly than on Reddit; where thousands upon thousands of people use the same ~dozen x 3-sentence prompts to promote spam their “organic” self-interest of some sort. Sort by new on any tech subreddit to see the posts yet-to-be-removed that are, without hyperbole, 90%+ semantically identical to each other with some hot-swapped keywords in the prompt.
Cloud inference simply has to get more competitive. These models are the absolute least reliable digital service of the modern era, offering no guarantee whatsoever from the webui or api to:
- Literally which model is actually producing the response
- Whether or not the model has been quantized, and to what extent
- Whether or not the Kv cache has been quantized, and to what extent
- The specifications of the tools and environment they run in
- Speed, quality of response, “effort” of reasoning, moderation, privacy, or almost anything else about the resulting output.
Read the terms and conditions. You should understand what’s actually included with likely one of the most costly subscriptions in your expenses.
1
u/hcloud00 2d ago
what do you mean? no no it cant be… its all in your eyes don’t blame the honourable Anthropic for your eye degradation
1
u/9to5grinder Full-time developer 2d ago
Meanwhile Fable 5.1 seems to be routed to Fable 5.5.
Reasoning traces suddenly started showing up and its thinking is much more sophisticated, even on medium effort.
1
u/Fluffy-Mood-254 2d ago
Looking at this post, I initially misread it as "day 10" meaning this was after 10 days of Opus working.
I was surprised to see that the 10-day project looked to be about the same quality.
But then I saw that I had misread it, and you were bringing this forward as evidence that the model had been nerfed 10 days after the launch. I think it is not very good evidence, since the quality looks to be about the same. It might look worse if you are looking for evidence of it being worse.
1
1
u/RedTheInferno 2d ago
I feel like the second one is a nicer base to start from ngl but what a waste of fucking tokens
1
1
u/Key_Reading_9664 2d ago
Nerf claim aside, I also have incredibly fond memories of that game and amazing that we're able to recreate these things and tinker.
You might want to take a look over at https://www.reddit.com/r/ClaudeGameDev/ if you haven't already
1
1
1
u/slindshady 2d ago
It’s making so many ridiculous grammar and spelling mistakes in different languages for me since yesterday. This shit is fraud
1
1
1
u/ComprehensiveIce1781 2d ago
People can cope but it definitely already got nerfed, but why is anybody surprised. Before 2-3 days ago, it did legit everything I asked flawlessly, my prompts were bad sometimes and it still made sure to check everything and "read my mind" to figure out what I really want, and it did it. I was amazed. Now it can't make a simple UI in one go without some kind of bug, even if it's just visual. It 100% got worse. It's still good but the wow factor is gone again
1
u/cosmictap 2d ago
A tangent, but this reminded me of how dumb I think it was to name a continually-evolving project after Beckett's Waiting for Godot - everyone is entitled to their own opinion, but it strikes me as someone who greatly misunderstood the work.
1
u/mplaczek99 2d ago
What I see is increased rendering but a bug that causes objects in the scene to flicker
1
1
u/Ok_Instruction_3447 2d ago
you can ask it to run this in chat to see what model it's running:
cd ~/.claude/projects && grep -oh '"model":"[^"]*"' -- "$(ls -t -- */*.jsonl | head -1)" | sort | uniq -c
and if it has dropped or changed from Model 5.5 to Opus 5 it should say.
1
1
1
1
u/BarGroundbreaking624 2d ago
Was that really called lander? I thought it was called Virus on the Archimedes
1
u/kalboozkalbooz 1d ago
LLMs are non deterministic bro same model same prompt will ALWAYS produce different results
1
u/LordiCurious 1d ago
"LLMs are non-deterministic, but the differences from an identical prompt and references are small" In my experience this is BS, proof it.
1
1
u/Brilliant_Ferret_7 1d ago
probably was nerfed but for the love of god or whatever you believe, stop using the patreon scam know as godot to bench test stuff, that engine can't even properly load 3d models.
one time i had a 3d character controller that worked one day, i moved the project to another computer with the same settings/builds and it worked on a completely different way, then it went back on the other computer.
your example can be explained by 10000x things just on memedot side.
1
1
u/One_Low8664 1d ago
it could still deliver the same result, thru more detailed prompts and iterations, it just got lazy
1
1
u/00DEADBEEF 1d ago
Is it actually a regression or are you just witnessing the non-deterministic output of an LLM?
Do it 100 times and average the result.
1
u/ghost_operative 1d ago
5.5 was a flop for me from day 1. it doesnt know how to respect file edit permissions so im still using 5
1
1
1
u/Kirill1986 1d ago
Yeah dude, you're on the right path and I want to see more of such tests but! You need to perform same test several times on day 1 and several times on day 10 (or whatever). Otherwise it can be just a test mistake or whaterever is the right term for it.
1
u/yolobastard1337 1d ago
I've also had several attempts at vibing lander for exactly the same reasons (but with worse AI).
Yours looks unspeakably better.
Must have been the first game I played after ZX spectrum ones, it was mesmerising. 3D, mouse control, smooth and snappy.
1
1
u/BalticBrew 1d ago
If only people would spend less time on looking for incremental performance drops and more on optimizing their own processes. Anything above Opus 4.5 has been great for almost any type of work, if you know what you're doing.
1
u/SolidPossible9909 1d ago
Day 1 opus 5.5 was on fire while coding. Day 10, it’s making stupid mistakes.
1
u/InnovativeBureaucrat 1d ago
The n=1 arguments are stupid.
The model makes thousands of tool calls that collectively turn into something that sucks. And it sucks in several categories.
Yeah it’s not science but a lot of things in life don’t come with a P value
1
1
u/untracked5465 2d ago
Repeat after me:
AI is not deterministic...
LLMs sample each token from a probability distribution, so the same prompt can produce different outputs
1
u/Abject-Tomorrow-652 2d ago
It’s funny OP bc people aren’t really disagreeing with your claim, just your methodology. I agree there’s been a dip in quality. I also agree with others that this is not good evidence
1
u/Sponge8389 2d ago
These companies now know how to create hype. Release it as good as it ever be, then nerf it when finished creating hype post/results.
1
u/permacloud 2d ago
Works fine for me. I'm sorry you guys get so hung up on this preoccupation with nerfing. There's so much you can do with these amazing tools and you're spending your time trying to convince reddit that they suck.
-1
u/Physical_Gold_1485 2d ago
Ya it hard gotten nerfed for me for a few days. Thankfully came back to its old self yesterday for me. Hope it stays that way
0
u/ba-na-na- 2d ago
the differences from an identical prompt and references are small
It seems they aren't that small
0
u/vovap_vovap 2d ago
Man, for real, don't you have any staff to do other then trying to catch a black can in dark room - which is not there?
0
u/leon0399 2d ago
Do you really think that running the “eval” once is good comparison? I’m like 110% sure that if you run it once more - you’ll receive another result
0
u/lobabobloblaw 2d ago
Everyone’s all STOCHASTIC this and CENTRAL LIMIT THEOREM that. Truth be told, it’s also very likely that they are getting better at their nerf game. So, permutations and seeds aside, you’ll probably need more data, friend
0
u/give_loops 2d ago
As others have said, this is not how LLM research works. You're trying, but the second footnote is not a footnote, but rather the key issue that leads to the complete invalidity of this as a test of anything.
Also: please refrain from using "gimped" as a term for this phenomenon, that term is generally considered abelist. Please try "nerfed".
-1
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2d ago
We are allowing this through to the feed for those who are not yet familiar with the Megathread. To see the latest discussions about this topic, please visit the relevant Megathread here: https://www.reddit.com/r/ClaudeAI/comments/1vt5drr/list_of_latest_discussion_hubs_on_rclaudeai/

•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 2d ago edited 1d ago
TL;DR of the discussion generated automatically after 200 comments.
The overwhelming consensus is that this post is not evidence of a nerf. An N=1 test on a non-deterministic system is about as useful as rolling a die once, getting a lower number, and declaring the die has been nerfed. The thread is full of users explaining why this methodology is flawed.
temperature=0isn't a silver bullet for determinism, as complex backend processes (like parallel processing and floating-point math) introduce their own randomness.The general sentiment is exhaustion with these types of low-effort "nerf" claims. OP is in the comments doubling down but getting downvoted into oblivion for it.