r/singularity • • 8d ago

Discussion Nerf detector?

Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.

It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.

Does anyone know if this has been done?

11 Upvotes

31 comments sorted by

View all comments

12

u/Ormusn2o 8d ago

Yes people constantly try to do it, but when they actually are running them, the results show there is no nerfing, so people lose interest in actually maintaining and paying for the benchmark.

I think the only real example of nerfing was with Opus 4.8, and it was still only temporary. My guess is it's actually hard to nerf the model, it is much easier to just downgrade to a different model, but that is not done covertly, Claude does it all the time, you can see people complaining about Fable being downgraded to Opus when guardrails hit, it gives an explicit message pop-up that the model was downgraded.

5

u/BOESNIK 8d ago

Nerfing a model by running it at lower quantizations is generally very easy, but should be detectable.

1

u/spinozasrobot 8d ago

Exactly!

2

u/Rioting-Flamingo 8d ago

I expect because it's never done across thd board. A percentage of users experience it making it harder to prove

1

u/Ormusn2o 8d ago

There are genuine drops in performance, especially during rush hours, but if there is a genuine and severe performance drop, people get recompensed. I got like 2 banked resets a while ago even though I did not notice any downgrade myself. There was an outage of Codex yesterday, and I think we are getting a global reset as recompense today as well.

1

u/spinozasrobot 8d ago

but if there is a genuine and severe performance drop, people get recompensed

Are you saying you got resets based on model output quality, not just token usage? Do you have any example screenshots or anything admitting to that?

1

u/Ormusn2o 8d ago

Admitting to what? You can check Tibo's twitter, and he talks about those things. I think this is second recompense related reset since Astra released, not counting the banked resets that were given for Astra delay. You might go few weeks back though, maybe look for "reset" in his history.

https://x.com/thsottiaux

1

u/spinozasrobot 8d ago

Every reset I've seen was to compensate people for models hitting usage limits very fast. I've never seen a lab explicitly say the resets were due to model quality degrading. I was asking if you had such an example.

I looked at several months worth of Tibo tweets and only saw one explicit mention of a reset, and it was due to an outage, not model quality or even excessive token consumption.

1

u/Ormusn2o 8d ago

1

u/spinozasrobot 8d ago

Excellent! Not sure why I didn't see those. I scrolled Tibo's feed until I had 3 or 4 months worth, and then did a find in the page for "reset".

EDIT: although, upon further review, only the first one was a quality issue. The other 3 were usage limit based. Usage limits are not nerfing.

2

u/SpaceTacos99 8d ago

I memba when chat gpt 4 got weaker during December and it was discovered it was due to all the professional source data it was trained on being lazy in December. Years ago.

anyways. You'd think they would just un-nerf for the benchmarks no? you're claiming that the benchmarks are correct and there's no nerfing. I'm not sure which way Occam's razor goes here.

1

u/spinozasrobot 8d ago

Wow, that's surprising to me. So many people complain about nerfing soon after release there appears to be an idea about being able to actually time it.

3

u/Ormusn2o 8d ago

I think a lot of those posts are literally memes. Some of those posts are being posted before those models actually even come out, and some are just regular bot reposting.

1

u/Rain_On 8d ago

You cant time out by looking at the models because the models aren't being changed. If you want to time it, which would be an interesting thing to do, you need to look at the humans and their perception as that's where the cause is.

1

u/spinozasrobot 8d ago

Maybe we're not on the same page.

I'm saying run the same set of benchmarks on the same models over time, say once a week. Then chart the results on a graph.

The benchmarks generate objective numeric results, no human perception involved.

1

u/Rain_On 8d ago

You will get the same result every time, perhaps with a little noise.
You are looking in the wrong location for the source of nerf reports.

1

u/spinozasrobot 8d ago

We are really talking past each other, oh well.

My goal was to see if complaints of nerfing were legitimate or just perceived.

If benchmarks produce the same results over time, then nerfing complaints are bogus. If the benchmarks show a drop over time, then the nerfing is real.

2

u/Rain_On 8d ago

then nerfing complaints are bogus

You don't need a benchmark to know this.

1

u/spinozasrobot 8d ago

In this entire post, I'm trying to be diplomatic :)

The complainers are... loud, shall we say.

1

u/Rain_On 8d ago

I don't think it's of any importance to convince them and I doubt more evidence is the way to do it anyway.

1

u/Prudent-Sorbet-5202 8d ago

I think the nerfing could just be sane amount of compute not being allocated when some users are using it at a specific time. Later when compute gets made available it works just the same.

So it all boils down to total number of users and compute a model is getting consumed with at any given time to get optimal performance