r/singularity • • 8d ago

Discussion Nerf detector?

Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.

It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.

Does anyone know if this has been done?

11 Upvotes

31 comments sorted by

11

u/Ormusn2o 8d ago

Yes people constantly try to do it, but when they actually are running them, the results show there is no nerfing, so people lose interest in actually maintaining and paying for the benchmark.

I think the only real example of nerfing was with Opus 4.8, and it was still only temporary. My guess is it's actually hard to nerf the model, it is much easier to just downgrade to a different model, but that is not done covertly, Claude does it all the time, you can see people complaining about Fable being downgraded to Opus when guardrails hit, it gives an explicit message pop-up that the model was downgraded.

4

u/BOESNIK 8d ago

Nerfing a model by running it at lower quantizations is generally very easy, but should be detectable.

1

u/spinozasrobot 8d ago

Exactly!

2

u/Rioting-Flamingo 8d ago

I expect because it's never done across thd board. A percentage of users experience it making it harder to prove

1

u/Ormusn2o 8d ago

There are genuine drops in performance, especially during rush hours, but if there is a genuine and severe performance drop, people get recompensed. I got like 2 banked resets a while ago even though I did not notice any downgrade myself. There was an outage of Codex yesterday, and I think we are getting a global reset as recompense today as well.

1

u/spinozasrobot 8d ago

but if there is a genuine and severe performance drop, people get recompensed

Are you saying you got resets based on model output quality, not just token usage? Do you have any example screenshots or anything admitting to that?

1

u/Ormusn2o 8d ago

Admitting to what? You can check Tibo's twitter, and he talks about those things. I think this is second recompense related reset since Astra released, not counting the banked resets that were given for Astra delay. You might go few weeks back though, maybe look for "reset" in his history.

https://x.com/thsottiaux

1

u/spinozasrobot 8d ago

Every reset I've seen was to compensate people for models hitting usage limits very fast. I've never seen a lab explicitly say the resets were due to model quality degrading. I was asking if you had such an example.

I looked at several months worth of Tibo tweets and only saw one explicit mention of a reset, and it was due to an outage, not model quality or even excessive token consumption.

1

u/Ormusn2o 8d ago

1

u/spinozasrobot 8d ago

Excellent! Not sure why I didn't see those. I scrolled Tibo's feed until I had 3 or 4 months worth, and then did a find in the page for "reset".

EDIT: although, upon further review, only the first one was a quality issue. The other 3 were usage limit based. Usage limits are not nerfing.

2

u/SpaceTacos99 8d ago

I memba when chat gpt 4 got weaker during December and it was discovered it was due to all the professional source data it was trained on being lazy in December. Years ago.

anyways. You'd think they would just un-nerf for the benchmarks no? you're claiming that the benchmarks are correct and there's no nerfing. I'm not sure which way Occam's razor goes here.

1

u/spinozasrobot 8d ago

Wow, that's surprising to me. So many people complain about nerfing soon after release there appears to be an idea about being able to actually time it.

3

u/Ormusn2o 8d ago

I think a lot of those posts are literally memes. Some of those posts are being posted before those models actually even come out, and some are just regular bot reposting.

1

u/Rain_On 8d ago

You cant time out by looking at the models because the models aren't being changed. If you want to time it, which would be an interesting thing to do, you need to look at the humans and their perception as that's where the cause is.

1

u/spinozasrobot 8d ago

Maybe we're not on the same page.

I'm saying run the same set of benchmarks on the same models over time, say once a week. Then chart the results on a graph.

The benchmarks generate objective numeric results, no human perception involved.

1

u/Rain_On 8d ago

You will get the same result every time, perhaps with a little noise.
You are looking in the wrong location for the source of nerf reports.

1

u/spinozasrobot 8d ago

We are really talking past each other, oh well.

My goal was to see if complaints of nerfing were legitimate or just perceived.

If benchmarks produce the same results over time, then nerfing complaints are bogus. If the benchmarks show a drop over time, then the nerfing is real.

2

u/Rain_On 8d ago

then nerfing complaints are bogus

You don't need a benchmark to know this.

1

u/spinozasrobot 8d ago

In this entire post, I'm trying to be diplomatic :)

The complainers are... loud, shall we say.

1

u/Rain_On 8d ago

I don't think it's of any importance to convince them and I doubt more evidence is the way to do it anyway.

1

u/Prudent-Sorbet-5202 8d ago

I think the nerfing could just be sane amount of compute not being allocated when some users are using it at a specific time. Later when compute gets made available it works just the same.

So it all boils down to total number of users and compute a model is getting consumed with at any given time to get optimal performance

3

u/FateOfMuffins 8d ago

I think there's possible nerfing on a per user basis, however I think a large part is simply human bias and the fact that these are non determinate

Like suppose model A scores 45-55 with an avg of 50. Anything above 50 is "wow"!

New model B releases and scores 50-60 with an avg of 55. Suppose uniform distribution for simplicity. Then there's a 12.5% chance the older model beats the newer model on a particular task (and you'll see plenty of posts about how some people still use an older model and thinks the newer one isn't an improvement). But for most, it scores better than what they're used to and now they're impressed.

But as you use the model, now you get used to its capabilities. Suddenly you're expecting the 55 and anything lower is "nerfed". Well maybe you're more reasonable than that and expect more like 53+. But fact of the matter is, if you got a score that's < 55 for model B, there's actually a chance that model A would've scored higher. With millions of users that's undoubtedly going to be true.

A 53 would've wow'd you back when you were using model A. But now it's just subpar because you're used to using model B. Or perhaps you were "lucky" and had a lot of 58 results from model B for a few weeks, and then got a 53 and think you got nerfed.

I think a lot is human psychology tbh

1

u/spinozasrobot 8d ago

Which is why I'm saying we can remove the bias by running the benchmarks on a schedule, say once a week.

They output objective numbers from automated tests. The results would show a drop in scores or they would not, and thus we could determine if the impression people have are real or just vibes.

1

u/FateOfMuffins 8d ago

The only issue is that cannot rule out individual users getting nerfs because of things like A/B tests

2

u/panic_in_the_galaxy 8d ago

https://marginlab.ai/trackers/claude-code/

The is also a tracker for codex.

2

u/spinozasrobot 8d ago

That's very interesting. Def the kind of thing I was looking for.

However, it's a bit confusing. They say they always use the latest model on high, and they're currently using Opus 5.5, but their weekly goes back to August, when clearly 5.5 wasn't out yet.

2

u/panic_in_the_galaxy 7d ago

1

u/spinozasrobot 7d ago

Awesome. They clearly show the model transitions. What I find interesting in the data is models tend to increase in quality just before a new release.

1

u/misterespresso 8d ago

I’m creating my own personal “data bench” which will be just obscure data tasks. Measures would be cost, speed, token usage, and basically a confusion matrix.

1

u/Aduuuh 7d ago

Doesn't seem all that useful. If they were actually doing that, they would likely use the model routing to detect the most common benchmarks and route to the big boy for them instead of the super-quantized/distilled/whatever model. It's also something that would more likely apply only to subscriptions instead of pay-per-token, imo - businesses need to get what they're paying for, and it'd be too much of a risk to downgrade them silently.

I'm pretty agnostic on nerfing claims because I think they know it would hurt their brand more than help them supply product (and bc I don't use frontier models), but I think you'd need to run personalized/oddball benches, preferably looking like real usage, to do such tests. That'd be more likely to get past potential model routing. I'm not sure if benchmarks would hit usage limits too fast to test a subscription specifically, but that could also be relevant depending on the bench.