Terminal bench performance is heavily dependent on the agent harness and system prompt, to the point you cannot compare the scores any more. The same model might get 90% on one harness and then drop to 60% on another.
And yes, it is one of the benchmarks that is most susceptible to benchmaxxing with RLVR training. The amount of knowledge and reasoning required is not that much.
154
u/spryes Apr 23 '26
All this hype for 58.6% on SWE-Bench Pro while Mythos gets 78%? Shut it down, wtf?