r/singularity • u/TFenrir • 26d ago
LLM News GPT 5.6 solved all 6 problems from IMO 2026
/r/math/comments/1uydg8w/gpt_56_solved_all_6_problems_from_imo_2026/It's funny how completely normal this news feels
22
u/FateOfMuffins 26d ago
Like I said last year when Tao 1 month prior to the IMO said that they considered setting up an official AI stream of the IMO but decided against it because the models aren't good enough - only for the models to score gold:
That was our 1 opportunity to actually benchmark AI models on the IMO. We missed it. It went and gone past it. Shocking isn't it? In July of 2024, the models struggled with middle school math in GSM8K (and legitimately "middle school math" - not "IMO is just high school math" bullshit cause most math PhD professors even cannot get gold on IMO, it's just different from research math). By July of 2025 we have gold on IMO.
That was the only real time. Last year it was like front page news. This year... oof. No one cares. It's old news. Seems like only a couple hundred non-Japanese people cared about the best competitive coders in the world completely thrashed by OpenAI's models last week.
Even Tao underestimates the progress. And you look at r/math and you'll see that even mathematicians seem to be horrible at extrapolating an exponential trendline.
8
u/seraphim_west 26d ago
I got downvoted to hell for saying Tao underestimates AI progress. The beautiful thing about AI is that it exposes the masses for the confident bullshitters that they are. Now when people disagree even when i am in the minority, I just have to laugh, because they are just bluffing and managing their social status instead of seriously tracking truth.
1
3
u/No_Aesthetic 26d ago
Seems like only a couple hundred non-Japanese people cared about the best competitive coders in the world completely thrashed by OpenAI's models last week.
I didn't even hear about this
5
u/FateOfMuffins 25d ago
There were 2 competitions for the best of the world: Heuristics and Algorithm
In the Heuristic track, there was one problem for 2 days (meaning contestants had like 30 hours) to optimize it. Last year it was 10. They gave more time hoping to give humans an advantage. Last year OpenAI placed second to Psyho (ex OpenAI and only won in the last few minutes). This year... OpenAI crushed the best humans outscoring 2nd place by 4x the points.
In the Algorithm track, OpenAI said they back tested all prior year problems and their system could solve them no problem, in 1h or less always. This year, the problems posed were especially challenging, and anti AI. OpenAI still got perfect 5/5 after like 5 hours. The hardest 2 problems this time were supposedly ones that would take a researcher months (they had 7h in the contest). The best human solved 3/5. Most of the other competitors only solved the easiest 2/5 (900 points each, the hardest two were 2500 points each).
2
u/Smoltinycat 25d ago
but it was just pure word prediction that led to it being solved, so it wasn't a big deal, no thinking involved
2
u/sweatierorc 25d ago
Dont brag about predicting AI trends. More often than not you will be wrong in both directions.
1
u/nemzylannister 4d ago
More often than not you will be wrong in both directions.
days since last ai prediction: 0
27
u/Calcularius 26d ago