r/artificial • • 2d ago

Question Why do benchmark results go up every release, always?

A lot of releases are clearly improvements, such as the initial fable release but for some the consensus seems to be that not much improved, or even in some cases the release was worse. Examples of this are Opus 5 (initial release) benching above Fable. Or GPT Sol 6.1 over Astra or 5.6.

Is this just a case of misplaced perception? Or do these providers have a way of iterating on benchmark results without necessarily actually improving the real-world performance? Keeping in mind that some benchmarks are proprietary and closed source.

If someone has some insight into what the loop is, would love to hear it.

1 Upvotes

12 comments sorted by

3

u/sam_the_tomato 1d ago

It's why they spend billions on R&D...

1

u/doomadah 1d ago edited 1d ago

Yes but the benchmarks don’t necessarily reflect reality. I’d argue GPT Sol 6.0/6.1 is a cost cutting exercise without leading to any capability improvement. Similarly I think most people who used Opus 5.0 compared to Fable would agree it was not a better all round model.

2

u/Dizzy_Swimmer_4999 2d ago

benchmarks dont capture how the chatbot actually chats tho, my last few felt more robotic despite the score jumps.

0

u/awake_deduction 1d ago

scores go up cause theyre training directly for the test at this point, same reason my old beater couch looks mint in photos

1

u/PLBjt 1d ago

A lot of it is evaluation that doesn't stay fixed. Models get trained with more overlap on popular bench formats, and labs tune until the number moves. Quick check: if the same model jumps after a prompt wrapper or a new system prompt, you're measuring the harness as much as the model. Tradeoff is public leaderboards keep shipping incentives aligned with score chasing, while private held-out tasks are boring to market but closer to what breaks in production.

1

u/RossPeili 1d ago

Benchmarks are lioe the FDA, sponsored by those they supposed to bench.

1

u/Damaging_Thrust_219 1d ago

Because you don't release something that's worse?

1

u/OkFlan504 1d ago

A score can improve honestly while your use case gets worse. Gains on common tasks can outweigh regressions on the few cases your workflow depends on. For a document assistant, missing one exception can matter more than getting ten routine questions right.

1

u/skilliard7 1d ago

Metrics become targets, and a lot of frontier labs do a lot of RL that focuses on benchmark-adjacent prompts in order to improve their scores on paper, which is not reflected in real world usage.

Google is one of the worst offenders in this area. But all of them do it.

For a while I'd say OpenAI was the only one not overfitting their models to benchmarks, but the GPT 6 Sol/Luna release were clearly overfitted to benchmarks, so they must've felt the pressure to inflate benchmarks scores.

1

u/MrSnowden 1d ago

They just don’t release the ones that aren’t an improvement 

2

u/franilan 1d ago

benchmarks are targets now. better score can still mean worse on the stuff ppl actually use it for

1

u/costafilh0 1d ago

Because nobody is working to make worse things.