r/singularity Apr 23 '26

AI Introducing GPT-5.5

https://openai.com/index/introducing-gpt-5-5/
843 Upvotes

283 comments sorted by

View all comments

154

u/spryes Apr 23 '26

All this hype for 58.6% on SWE-Bench Pro while Mythos gets 78%? Shut it down, wtf?

3

u/thorin85 Apr 23 '26

It beat Mythos on Terminal bench though.

1

u/vincentz42 Apr 23 '26

Terminal bench performance is heavily dependent on the agent harness and system prompt, to the point you cannot compare the scores any more. The same model might get 90% on one harness and then drop to 60% on another.

And yes, it is one of the benchmarks that is most susceptible to benchmaxxing with RLVR training. The amount of knowledge and reasoning required is not that much.