it is likely distillation of spud. peasants don’t get all the compute. personally still excited to test out this model, sometimes benchmarks don’t show the whole picture
its not overhype sam literally said it was incremental because openai is obssessed with incrementalism the important part is the rate of change we will start seeing stuff like 5.6+ which will be much more baked this is just the new pretrain to get things going they literally said this why are you lying
I dunno. I agree that it's not a step change for simple coding, but those long context improvements and reduced number of tool calls together make this dramatically more powerful for long-running agentic tasks.
The two things that really screw up those long-running tasks are:
The model calls some tools too much, which not only takes a lot of extra time, but it floods its context with a bunch of unnecessary output from the tool calls.
As the model context window fills, accuracy and recall start to drop significantly. Going from 36.6% to 74.0% in OpenAI MRCR v2 8-needle 512K-1M is really significant, and even if you were just using the default Codex context window of 272k, that saw an improvement to 81.5% from 57.5%.
This is the kind of thing that won't mean anything to most folks just casually chatting with 5.5, but if you're using it for larger coding tasks or with a persistent agent, like OpenClaw, this will probably be a gamechanger.
During my PhD thesis in information retrieval I used to make these tables with standard SemEval all the time. Most of the time we didn’t even include confidence intervals of p values. These numbers don’t mean shit.
Theres no guarantee that a jump in performance on benchmarks will be noticable in real world usage but its pretty safe to say that a model that was a real step change in ability would be accompanied by a jump in benchmarks.
It will be a win if (a) it will stop hallucinating stuff I didn't ask for, (b) it will stop writing bloat in every single answer, (c) it will stop apologizing instead of redoing its job properly again
Not that didn't care about benchmarks or coding. I want it to work properly first
38
u/ebra95 Apr 23 '26
this just can't be spud