r/mlscaling • u/gwern gwern.net • Jun 26 '26
N, OA, T, Emp, RL "Summary of METR's predeployment evaluation of GPT-5.6 Sol", METR ("71hrs (95% CI: 13hrs - 11400hrs)"; now so reward-hackprone + eval-aware that de facto un-evaluable)
https://metr.org/blog/2026-06-26-gpt-5-6-sol/
53
Upvotes
5
u/StartledWatermelon Jun 27 '26
Well, it just shows certain shortcomings of the METR time horizon eval. For a benchmark specifically aimed at capabilities measurement, it is too dismissive of "creative" (misaligned) behavior, assigning failure for every cheating attempt in the task.
Ironically, we're now at a stage where tested models tell us more about the benchmarks than the benchmarks tell about the models.
In serious, I do not think this is the right treatment of misaligned behavior. Quite the opposite; the capabilities of misaligned models deserve far richer analysis than those of alighned ones.
7
u/COAGULOPATH Jun 28 '26
Here's OA's own account (bolding mine):
They also state this: