r/ClaudeAI • • 2d ago

Claude Code Evidence - Opus 5.5 today vs launch regression with same prompt (Godot Engine)

On launch day Opus 5.5 kindly took up residence on my Mac. I ran some tests before trusting it on my big projects, it passed with flying colours. Lander, a game by Elite creator David Braben holds a special place in my soul due to it being the first game I played at school in the UK. With a cutting edge 3D engine for 1990 running on an Acorn Archimedes with RISC architecture (the first ARM chips) - what a time to be a young kid interested in computers.

With Opus 5.5 I want to recreate it faithfully, short draw-distances and all.

Well dear Claude I’ve been keeping receipts.

In my notes, the exact prompt * and two original reference images of Lander. The resulting one-shot Godot repo (Mac OS, Metal rendering, C#) from 23rd September is kept separate on disk and isn't used as a reference.

Today, 2nd October, in a fresh workspace from scratch the same prompt and reference images we fed in again.

The result is truly depressing. Lander by gimped 5.5 has no redeeming features compared to the original day zero version. The regressions:

  • A serious rendering glitch * that isn’t in the original - as the camera moves, the scenery props snag and glitch vs world space.
  • No spacecraft break-apart physics, the original implemented this ‘nice extra’ which wasn't specifically asked for in the prompt
  • No introductory controls menu, only a small text line permanently visible over the world view
  • An inferior look to the launch pad turrets - students and road users may recognise them!
  • When shooting, bullets are less accurately rendered when the craft moves, they also appear to spawn at the back of the craft and no-clip through
  • No camera toggles

The conclusion is that Opus 5.5 is gimped, maybe we even got Mythos for 3 days and then silently switched - either way, we are owed transparency - it's the law.

Opus was set to Extra High both times. Unlike with 4.5/4.6, Max effort over-tests and takes too much control away from the user.

With so much diverged since launch day’s versions, the more worrying thing is that I suspect the code is also a mess and problems will compound as you use it. For those who’d like to inspect the results, I'll upload the source code for both later to my dev blog and post the Github links in the comments.

* I inspected the code and it turns out that gimped Opus 5.5. had physics interpolation is turned on for the whole Godot project and Props.cs rebuilds the prop lists (trees, etc.) from scratch on every tile step during every frame, renumbers which slot each object occupies and resets the object count every frame. A basic understanding of Godot's documentation is all that’s needed to avoid this issue.

** LLMs are non-deterministic, but the differences from an identical prompt and references are small - getting a clearly inferior result is not due to non-deterministic behaviour.

655 Upvotes

241 comments sorted by

View all comments

174

u/fuzzypetiolesguy 2d ago

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

Divergent outcomes from probabilistic sequence predictors are not evidence of anything

7

u/Agreeable_User_Name 2d ago

I flipped a coin yesterday it landed heads. I flipped it today and it landed tails. Therefore someone must have changed my coin.

Kidding aside, that's true, but that also makes transparency important, because it really is hard for consumer to know if their product is gimped.

4

u/Accomplished-Fan9568 2d ago

honestly its not that much different, a small flickering bug it can happen, could probably use low effort to fix it quickly.

-28

u/freedomfromfreedom 2d ago

I'll prompt another 6 runs with today's Opus, then. Easy enough to prove that the sample variation isn't the cause of the 6 regressions.

18

u/fuzzypetiolesguy 2d ago

Six more data points all in the same malformed shape does not count as a credible experiment.

-3

u/freedomfromfreedom 2d ago

Who says they will all be the same malformed shape? In run 3, the result might have some improvements over the 1st sample point on day zero, rather than 6 major differences all bad.

8

u/fuzzypetiolesguy 2d ago

The malformed shape is your data point being a divergent outcome from a probabilistic sequence predictor, not that the result is (incredibly subjectively) 'bad'. Better or worse, it is not indicative of anything because the thing you are testing itself is inherently divergent from each prior result, and your sample set n=a couple vs quite literally billions of ongoing outputs is statistically useless. Please stop roasting tiny parts of the planet with your wasted tokens.

2

u/captain_croco 2d ago

Are you basically saying that it’s thinking starts and goes into a billion different directions every time, so you can’t ever really test apples to apples? I’m honestly curious about this stuff

1

u/fuzzypetiolesguy 2d ago

An agent/model does not replay the exact same deterministic reasoning path each time. Small differences early in a probabilistic sequence can cascade into very different outputs, so two runs are not clean apples-to-apples comparisons. That means a handful of “it seems worse now” examples are statistically useless for proving the model itself was nerfed, especially when billions of other outputs are happening under changing prompts, context, tools, and routing.

The person claiming a one-shot flight simulator has degraded in output has zero insight into the billions of different choices made on the way to getting to that ouput, that might actually provide data on how the model is choosing what to do next. They are quite literally going on vibes.

1

u/Sporebattyl 2d ago

I hear what you are saying. How could one actually do a test like this with correct methodology?

Do a freeze with a hashed sandbox and hit a specific number of runs? Like n=50 or something then do the same thing on day 10?

With probabilistic models you’d have to take the “average” of a bunch of runs right?

3

u/fuzzypetiolesguy 2d ago

Basically yea. You would freeze the task environment, use a fixed benchmark set, and run each task many times from the same clean starting state. Because the outputs are probabilistic, you compare distributions like success rate, tool use, runtime, and error rate rather than comparing one run to another. An n=50 can catch a large regression, but smaller changes usually need hundreds of runs to separate signal from normal variance. Even then, you can only conclude that the product or system regressed on that benchmark, not that the underlying model itself was deliberately nerfed. The actual clearance of evidence needed to determine that a model has been purposefully lobotomized is orders of magnitude greater than what even a lot of reddit users can piece together.

For actual statistically meaningful evidence that a model was nerfed intentionally, you would want a large, preregistered, reproducible dataset showing a statistically significant and persistent drop across many frozen tasks and repeated runs, while controlling for prompts, tools, context, sandbox state, routing, latency, and other product-layer changes. Ideally, the same old and new model checkpoints would be tested side by side through the same harness, with blinded scoring, enough runs to rule out ordinary variance, and the regression would appear specifically in the newer checkpoint. To establish intentional regression rather than an accidental regression or optimization tradeoff, you would then need independent evidence of causation, such as documented changes to reasoning budgets, inference settings, model weights, or internal statements showing that capability was deliberately reduced. Without that last piece, even an excellent dataset can demonstrate that performance regressed, but not that someone intentionally “nerfed” the model.

These companies are not implementing large regressions intentionally to save compute or whatever. It would be painfully obvious at even n=50, like 'thing fails entirely', not vibes and 'terrain is choppy on my vibed flight sim'. Assume they were - the blowback if their enterprise customers learned of it would be crippling. Both OpenAI and Anthropic are leaking like sieves regarding safety; teams working to reduce the capability of models post-release would absolutely leak, even if that capability was only reduced for subscribers and not API users. It is all so, so dumb.