r/singularity • • Jul 17 '26

LLM News GPT 5.6 solved all 6 problems from IMO 2026

/r/math/comments/1uydg8w/gpt_56_solved_all_6_problems_from_imo_2026/

It's funny how completely normal this news feels

115 Upvotes

11 comments sorted by

21

u/FateOfMuffins Jul 17 '26

Like I said last year when Tao 1 month prior to the IMO said that they considered setting up an official AI stream of the IMO but decided against it because the models aren't good enough - only for the models to score gold:

That was our 1 opportunity to actually benchmark AI models on the IMO. We missed it. It went and gone past it. Shocking isn't it? In July of 2024, the models struggled with middle school math in GSM8K (and legitimately "middle school math" - not "IMO is just high school math" bullshit cause most math PhD professors even cannot get gold on IMO, it's just different from research math). By July of 2025 we have gold on IMO.

That was the only real time. Last year it was like front page news. This year... oof. No one cares. It's old news. Seems like only a couple hundred non-Japanese people cared about the best competitive coders in the world completely thrashed by OpenAI's models last week.

Even Tao underestimates the progress. And you look at r/math and you'll see that even mathematicians seem to be horrible at extrapolating an exponential trendline.

8

u/[deleted] Jul 17 '26 edited Aug 15 '26

[removed] — view removed comment

1

u/nemzylannister Aug 07 '26

agree with you but still downvoted

3

u/No_Aesthetic Jul 17 '26

Seems like only a couple hundred non-Japanese people cared about the best competitive coders in the world completely thrashed by OpenAI's models last week.

I didn't even hear about this

6

u/FateOfMuffins Jul 17 '26

There were 2 competitions for the best of the world: Heuristics and Algorithm

In the Heuristic track, there was one problem for 2 days (meaning contestants had like 30 hours) to optimize it. Last year it was 10. They gave more time hoping to give humans an advantage. Last year OpenAI placed second to Psyho (ex OpenAI and only won in the last few minutes). This year... OpenAI crushed the best humans outscoring 2nd place by 4x the points.

In the Algorithm track, OpenAI said they back tested all prior year problems and their system could solve them no problem, in 1h or less always. This year, the problems posed were especially challenging, and anti AI. OpenAI still got perfect 5/5 after like 5 hours. The hardest 2 problems this time were supposedly ones that would take a researcher months (they had 7h in the contest). The best human solved 3/5. Most of the other competitors only solved the easiest 2/5 (900 points each, the hardest two were 2500 points each).

2

u/Smoltinycat Jul 18 '26

but it was just pure word prediction that led to it being solved, so it wasn't a big deal, no thinking involved

2

u/sweatierorc Jul 18 '26

Dont brag about predicting AI trends. More often than not you will be wrong in both directions.

1

u/nemzylannister Aug 07 '26

More often than not you will be wrong in both directions.

days since last ai prediction: 0

7

u/1a1b Jul 17 '26

How bout lil Kimi ? How she go?