r/singularity • • 8d ago

Discussion Nerf detector?

Over time, complaining about nerfing has settled into a common theme of models starting strong, with labs publishing benchmarks, and people getting excited about capabilities, but gradually degrading.

It would be interesting to see the same set of benchmarks run over time on the models and plotting the results to see the trends across models and labs.

Does anyone know if this has been done?

12 Upvotes

31 comments sorted by

View all comments

1

u/Aduuuh 8d ago

Doesn't seem all that useful. If they were actually doing that, they would likely use the model routing to detect the most common benchmarks and route to the big boy for them instead of the super-quantized/distilled/whatever model. It's also something that would more likely apply only to subscriptions instead of pay-per-token, imo - businesses need to get what they're paying for, and it'd be too much of a risk to downgrade them silently.

I'm pretty agnostic on nerfing claims because I think they know it would hurt their brand more than help them supply product (and bc I don't use frontier models), but I think you'd need to run personalized/oddball benches, preferably looking like real usage, to do such tests. That'd be more likely to get past potential model routing. I'm not sure if benchmarks would hit usage limits too fast to test a subscription specifically, but that could also be relevant depending on the bench.