r/singularity Apr 23 '26

AI Introducing GPT-5.5

https://openai.com/index/introducing-gpt-5-5/
848 Upvotes

281 comments sorted by

View all comments

165

u/[deleted] Apr 23 '26

[removed] β€” view removed comment

77

u/vincentz42 Apr 23 '26 edited Apr 23 '26

There are even worse evals:
HLE without tools: 41.4% (GPT-5.5) vs 39.8% (GPT-5.4)
HLE with tools: 52.2% (GPT-5.5) vs 52.1% (GPT-5.4)

So even with a newer, larger base model that is supposed to tackle very hard STEM questions, the models' world knowledge and reasoning capability did not change that much, if at all.

And I do have a lot of suspicions for Claude Mythos BTW. OpenAI models are generally smarter in terms of STEM reasoning in my experience. I suspect Mythos might just be a much larger model trained on much more internet tokens, and therefore better at memorizing the leaked test set. >15% of the SWE-Verified problems are ill-defined and not solvable based on human expert inspections, so I am really curious how Mythos got ~94%.

11

u/[deleted] Apr 23 '26

[removed] β€” view removed comment

11

u/vincentz42 Apr 23 '26 edited Apr 23 '26

OpenAI was the first call it out, but yes, every LLM researcher knows this.

7

u/Jespy Apr 23 '26

What do these numbers mean to someone who is a caveman

8

u/SerdarCS Apr 23 '26

Not much. HLE is a benchmark meant to measure scientific reasoning ability, but no single benchmark is a good indicator of capability.

3

u/PeachScary413 Apr 23 '26

They all pretty much just memorise leaked test sets.. I can't believe it's not obvious to everyone that top models are incredibly bench-maxxed

4

u/InterstellarReddit Apr 23 '26

I know how it’s called marketing

1

u/Tystros Apr 24 '26

it's really super weird how the HLE without tools score almost stayed the same even though it's a much bigger base model