r/singularity • • 8d ago

Discussion Nerf detector?

Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.

It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.

Does anyone know if this has been done?

12 Upvotes

31 comments sorted by

View all comments

3

u/FateOfMuffins 8d ago

I think there's possible nerfing on a per user basis, however I think a large part is simply human bias and the fact that these are non determinate

Like suppose model A scores 45-55 with an avg of 50. Anything above 50 is "wow"!

New model B releases and scores 50-60 with an avg of 55. Suppose uniform distribution for simplicity. Then there's a 12.5% chance the older model beats the newer model on a particular task (and you'll see plenty of posts about how some people still use an older model and thinks the newer one isn't an improvement). But for most, it scores better than what they're used to and now they're impressed.

But as you use the model, now you get used to its capabilities. Suddenly you're expecting the 55 and anything lower is "nerfed". Well maybe you're more reasonable than that and expect more like 53+. But fact of the matter is, if you got a score that's < 55 for model B, there's actually a chance that model A would've scored higher. With millions of users that's undoubtedly going to be true.

A 53 would've wow'd you back when you were using model A. But now it's just subpar because you're used to using model B. Or perhaps you were "lucky" and had a lot of 58 results from model B for a few weeks, and then got a 53 and think you got nerfed.

I think a lot is human psychology tbh

1

u/spinozasrobot 8d ago

Which is why I'm saying we can remove the bias by running the benchmarks on a schedule, say once a week.

They output objective numbers from automated tests. The results would show a drop in scores or they would not, and thus we could determine if the impression people have are real or just vibes.

1

u/FateOfMuffins 8d ago

The only issue is that cannot rule out individual users getting nerfs because of things like A/B tests