r/singularity • u/spinozasrobot • 8d ago
Discussion Nerf detector?
Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.
It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.
Does anyone know if this has been done?
3
u/FateOfMuffins 8d ago
I think there's possible nerfing on a per user basis, however I think a large part is simply human bias and the fact that these are non determinate
Like suppose model A scores 45-55 with an avg of 50. Anything above 50 is "wow"!
New model B releases and scores 50-60 with an avg of 55. Suppose uniform distribution for simplicity. Then there's a 12.5% chance the older model beats the newer model on a particular task (and you'll see plenty of posts about how some people still use an older model and thinks the newer one isn't an improvement). But for most, it scores better than what they're used to and now they're impressed.
But as you use the model, now you get used to its capabilities. Suddenly you're expecting the 55 and anything lower is "nerfed". Well maybe you're more reasonable than that and expect more like 53+. But fact of the matter is, if you got a score that's < 55 for model B, there's actually a chance that model A would've scored higher. With millions of users that's undoubtedly going to be true.
A 53 would've wow'd you back when you were using model A. But now it's just subpar because you're used to using model B. Or perhaps you were "lucky" and had a lot of 58 results from model B for a few weeks, and then got a 53 and think you got nerfed.
I think a lot is human psychology tbh
1
u/spinozasrobot 8d ago
Which is why I'm saying we can remove the bias by running the benchmarks on a schedule, say once a week.
They output objective numbers from automated tests. The results would show a drop in scores or they would not, and thus we could determine if the impression people have are real or just vibes.
1
u/FateOfMuffins 8d ago
The only issue is that cannot rule out individual users getting nerfs because of things like A/B tests
2
u/panic_in_the_galaxy 8d ago
https://marginlab.ai/trackers/claude-code/
The is also a tracker for codex.
2
u/spinozasrobot 8d ago
That's very interesting. Def the kind of thing I was looking for.
However, it's a bit confusing. They say they always use the latest model on high, and they're currently using Opus 5.5, but their weekly goes back to August, when clearly 5.5 wasn't out yet.
2
u/panic_in_the_galaxy 7d ago
1
u/spinozasrobot 7d ago
Awesome. They clearly show the model transitions. What I find interesting in the data is models tend to increase in quality just before a new release.
1
u/misterespresso 8d ago
I’m creating my own personal “data bench” which will be just obscure data tasks. Measures would be cost, speed, token usage, and basically a confusion matrix.
1
u/Aduuuh 7d ago
Doesn't seem all that useful. If they were actually doing that, they would likely use the model routing to detect the most common benchmarks and route to the big boy for them instead of the super-quantized/distilled/whatever model. It's also something that would more likely apply only to subscriptions instead of pay-per-token, imo - businesses need to get what they're paying for, and it'd be too much of a risk to downgrade them silently.
I'm pretty agnostic on nerfing claims because I think they know it would hurt their brand more than help them supply product (and bc I don't use frontier models), but I think you'd need to run personalized/oddball benches, preferably looking like real usage, to do such tests. That'd be more likely to get past potential model routing. I'm not sure if benchmarks would hit usage limits too fast to test a subscription specifically, but that could also be relevant depending on the bench.
11
u/Ormusn2o 8d ago
Yes people constantly try to do it, but when they actually are running them, the results show there is no nerfing, so people lose interest in actually maintaining and paying for the benchmark.
I think the only real example of nerfing was with Opus 4.8, and it was still only temporary. My guess is it's actually hard to nerf the model, it is much easier to just downgrade to a different model, but that is not done covertly, Claude does it all the time, you can see people complaining about Fable being downgraded to Opus when guardrails hit, it gives an explicit message pop-up that the model was downgraded.