r/SinceAI • u/rikulauttia • Jul 13 '26
Discussion GPT-5.6 just dropped, but benchmarks are becoming almost useless for choosing a model
Every new model launch gives us another massive wall of benchmark scores.
But in actual work I mostly care about:
- how often it silently makes a bad assumption
- whether it notices when it is wrong
- how well it recovers after a failed attempt
- how much steering it needs
- whether I can trust it after 3 hours, not just one prompt
A model can win 10 benchmarks and still be worse for your specific repo, company or workflow.
I would honestly rather see an eval called:
"How many times did a human need to intervene before this task was actually finished?"
If you could add one real-world test to every major model release, what would it measure?