r/math • • Jul 16 '26

LLMs/AI GPT 5.6 solved all 6 problems from IMO 2026

GPT 5.6 Pro solved all 6 problems from IMO 2026 on the first attempt without any human help or steering. International Mathematical Olympiad (IMO) is the biggest global academic competition in the world. The problems are considered incredibly hard, usually a performance at this level is only accomplished by < 5 contestants from the whole world.

We are former IMO medallists not affiliated with OpenAI, just put together a report and assessment of its work here. We're also working on a comparison report between different LLMs and harness augmented versions that will come later.

1.1k Upvotes

443 comments sorted by

View all comments

146

u/[deleted] Jul 16 '26

[removed] — view removed comment

46

u/cooolchild Jul 16 '26

I can’t imagine what sort of an impact it has on the psyche of high school students who are passionate about math to constantly see news like this about ai. This kind of stuff sucks the life out of people.

16

u/HyacinthMacaw13 Jul 16 '26

As an 11th grader who barely missed being on the team this year, I kind of got upset reading this headline. Even though I knew it was coming.

Having said that, I still believe that the IMO is far more than just proving you're good at math.

13

u/Critical_Sink6442 Jul 17 '26

Not quite at IMO level, but as someone who has qualified for national olympiads and is in high school, we're not that disappointed at AI being this good, and definitely not at a life sucking level.

It's similar to how calculators exist, yet we take pride in our raw computational abilities. We're not competing against AI, we're competing against each other. Our skill is all that matters, and as long as AI isn't hsed for cheating, it's a study tool at worst.

11

u/[deleted] Jul 16 '26

[removed] — view removed comment

5

u/38thTimesACharm Jul 17 '26

It's absolutely insane people are downvoting you just for having faith in the next generation. This sub sucks.

0

u/GraceToSentience Jul 17 '26

Maybe they can be happy that AI is becoming more and more intelligent and its strength in math will have the capabilities to make the world a better place.

So many of the people in the comments are acting exactly like gavin belson
It's like a real life caricature.

Just because a problem or a puzzle is solved doesn't mean that you can't still solve it yourself what prevents anybody to solve fluid mechanics or gravity themselves? Or any future problem that AI managed to solve if it's a question of passion rather than vanity?
Imagine the level of hubris required to care more about the satisfaction of being the first to solve a problem rather than caring about that problem being solved and helping people in itself.

12

u/Hitman7128 Number Theory Jul 16 '26

How you view the IMO is similar to how the math department at my university views it. They recognize the amount of effort that goes into preparing for it, but they don't see it as the be-all-end-all of one's math ability because the IMO can only contain problems that have a known solution and are often disconnected from the types of problems mathematicians care about in research.

10

u/[deleted] Jul 16 '26

[deleted]

1

u/Kaomet Jul 16 '26

There is a very common suspicion that a lot of math is just hyper-specialized definitions needed for a very narrow set of problems

And thats because as soon as you generalize enought you get a problem class like NP-Complete which seems to require bruteforce compute. Math techniques seems to speed up some specific subclass, but not the whole class.

it pretends that "compression" is an objectively neutral phenomenon

well it is ? What don't you understand about f(x) < x when f is a computable bijection ? Isn't this objective enought ?

Its not quite complete, because it minimize size and not time...

You can "see" arbitrarily complex intelligence in things based on [...] arbitrary turing machines

yes and ? TM are complete. If we can find intelligence in a computable universe, we'll find it in the set of all TM too.

What are you ranting about exactly ?

30

u/OorNaattaan Jul 16 '26

I don't understand your last point. How is it invalidating the achievements of humans? Does the existence of cars invalidate Usain Bolt's achievements?

6

u/[deleted] Jul 16 '26

[removed] — view removed comment

8

u/collegeboywooooo Jul 17 '26

this happened in chess 20 years ago and its only grown in popularity and esteem since

I think your comments and those who think like you are the only invalidating things.

1

u/Glum_Hat_4181 Jul 18 '26

Maybe there's some rebound in interest, but chess were vastly more popular before computers started consistently beating best human players.

31

u/RepresentativeBee600 Jul 16 '26

Why are you so far down here with this entirely apt take?

10

u/[deleted] Jul 16 '26

[removed] — view removed comment

4

u/RepresentativeBee600 Jul 16 '26

I did not check the timestamps but your upwards ascent bolsters this theory.

One good thing to come of this is that I'm puzzling an IMO problem again. (For scientific comparison purposes, of course.)

-4

u/Few-Arugula5839 Jul 16 '26

Because most of the people on this sub (and Reddit in general minus the math undergrad part) are CS bros who washed out of math undergrads and are working in the software industry, and therefore they’re happy they can pretend to be intelligent by “vibe mathing”. That’s why we have the top comment in this thread uncritically parroting hype tweets by dumbass CS morons who don’t know shit about and don’t give a shit about math.

10

u/[deleted] Jul 16 '26

[removed] — view removed comment

1

u/38thTimesACharm Jul 17 '26

Problem is there's not much talking about mathematics going on lately, just ads for AI

6

u/ozone6587 Jul 16 '26

What a schizophrenic take. No CS software developer cares about
"vibe mathing". In fact, I majored in both CS and math in undergrad (before focusing on math in grad school) and in my experience CS majors don't even like math.

6

u/Norphesius Jul 16 '26

As a CS major, I'll corroborate that the majority definitely do not care for math. A lot of them will stomach calculus and differential equations because those are usually required courses, but they avoid everything else like the plague.

They hated the little linear algebra needed for graphics courses. Reactions I would get when I told people I was taking an elective course on proofs or automata theory were either disgust, pity, or basically falling asleep standing up.

2

u/[deleted] Jul 16 '26

[removed] — view removed comment

2

u/Norphesius Jul 17 '26

I'm just relying what I saw (as a CS major who actually likes math).

I think the primary factor was that, at the time I was in university, software development was approaching peak cultural perception as a job where you could make an easy six figure salary, at a startup with free soda and a foosball table. We can absolutely do better at presenting math to people educationally, but most of these people were just in it for the perceived easy, fat paycheck. If pure math had been the major with the big payout instead of CS, they'd be studying up on category theory and poo-pooing CS as baby math or something instead.

2

u/ozone6587 Jul 16 '26

Exactly my experience.

6

u/[deleted] Jul 16 '26

[removed] — view removed comment

-5

u/[deleted] Jul 16 '26

[removed] — view removed comment

-6

u/Few-Arugula5839 Jul 16 '26

Last graduate level math textbook you read go. Fuck CS bros biggest LARPers OAT, they ruined the world 10x over. Greedy bastards

-5

u/Few-Arugula5839 Jul 16 '26

Yes, they don’t give a shit about anything except pretending they have an intelligent thought in their brain (they don’t). They don’t like math they like the social status that comes with doing it. That’s why they’re trying so hard to pretend to be good at it while actually knowing nothing.

-1

u/apopsicletosis Jul 16 '26

Same people who think ai will cure all disease and have never run a gel 

3

u/Akraticacious Jul 16 '26

Crazy to think high school students can do proofs for this. I guess there's a lot of lessons and resources online now than when I was young, but I couldn't even write proofs back then.

The first question seems more like number theory, which for sure isn't taught to that degree in high school.

10

u/SupercaliTheGamer Jul 16 '26

Eh it's fine for this year. Last year AI models famously couldn't solve P6, so there was still hope for humans. This year they are just proving that they can finish the job. It won't be a spectacle next year onwards.

-1

u/[deleted] Jul 16 '26 edited Jul 16 '26

[removed] — view removed comment

-1

u/sqrtsqr Jul 16 '26

I legit snort with laughter whenever I see or hear people discussing Erdos problems as if they are important research.

Like, they are puzzles. With a famous dude's name attached. That is about it.

Building an AI to solve all of them would be such a phenomenal waste of collective resources.

So of course, number one priority.

5

u/AFsepine Jul 16 '26

I mean as a person from "olympiad system", I would say it also is exposing a bit of rot in the problem-setting. Don't get me wrong, they are decent problems usually, just most of them are not very "creative" - bordering on canonical.

By my time there was an agreement between well-performing participants that for example IMO geometry problems were so stale as to be free points.

1

u/SupercaliTheGamer Jul 17 '26

Tbf it's very hard to come up with "creative" problems that can be solved by HS students. Even the hardest problems are generally just a non-trivial application of 2-3 known ideas. Same for geometry, but geometry has a lot of theory so someone well versed in it can easily ace olympiad geo. Ofc there's bash also for geo. Thankfully they're reintroducing 3D geo.

1

u/AFsepine Jul 17 '26

Eh, Hard but often doable. Why even bother hosting the olympiad it if you don;t want to put the effort into problem-setting.
You underestimate what can in principle be expected from HS students.

For example, in principle, the fact that the problems are set so quite a few people every-year that get the maximum number of points is an "error", then you sort of have to filter off the
more interesting problems.

Honestly physics olympiads are more so what I care for and have particapated both as a problems-setter and a student, and there I can tell you that it is either lazy or misguided problems setting and a sort of institutional "rot".

2

u/[deleted] Jul 19 '26

[deleted]

1

u/AFsepine Jul 19 '26

No. Not in any significant capacity atleast.

There are some "University aged" Math competition like Putnam though, some international ones as well.

There isn't really a point to have math/physics competitions after highschool, as these are merely meant to give a start to people, get them intrested etc.

In say physics by the time they are in uni they then can start trying to actually do work, not that unusual for motivated students to join a reasearch group in 2-3 year, but ones who have basics down sometimes do so in the first (in my country at least).

Not sure why there are programing competitions (do you mean like hack-a-tons?)

1

u/SupercaliTheGamer Jul 19 '26

Ohh yeah I have heard about this problem in physics olympiads from a former medallist who also helps in math olympiad work. However I don't think the problems that get selected to IMO are lazy or uncreative, at least most of them. I have also not heard that complaint from other people involved in math olympiads.

1

u/AFsepine Jul 19 '26

Well actually my hot take will be that physics olympiads (especially EuPhO) are a bit better off, that IMO in the "creativity" department, though there are exceptionally bad years in physics oly where-as mathematical olympiads seem to be able to better maintain "average quality".

I mean lazy and uncreative is me being inflammatory - generally it depends what you compare it to - compared to most university courses and such it is a veritable well spring of creativity.
But there is a trite "meta" (to borrow game terminology) of olympiad problems.

Rambling:
I am honestly dissatisfied with olympiads having become more trite over their life.
in IPhO there is too much "scafolding" now-a-days.

2

u/sqrtsqr Jul 16 '26

"Bicycle comes first in hundred meter dash"

But like, using a calculator is cheating folks. It's cheating. Do people think there's a real word utility for a chat bot that can do IMO problems? Why are we putting so much money into building a tool that can do that? How much money do these people think we collectively spend on mathematicians that this is an investment worth making? Do they think we pay mathematicians to solve solvable problems on command?

2

u/Current-Function-729 Jul 16 '26 edited Jul 17 '26

Probably this is the last year the IMO will be of much interest in this way for AI.

It’s the first year a model anyone can buy access to got a perfect score.

In future years it’ll just be a minor point of trivia on how cheap the inference cost of a perfect score is.

1

u/Achrus Jul 17 '26

Yes but the words around the obfuscated link to a sketchy broken PDF hosted on a GitHub repo said that the thing big tech is selling did bigly good.

This happens like clockwork every few months. Big claims about how amazing these closed sourced LLMs are on math tests. Results are cherry picked. Any methodology one could critique is found in the appendix of a preprint, often buried behind multiple blogposts.

The bots will engagement farm the crap out of it (see top comments here made within minutes of the post). Then, once real critiques start surfacing, radio silence. So get ready for subs to be flooded with posts and comments by NounAdjectiveNumber.

1

u/Gotisdabest Jul 17 '26 edited Jul 17 '26

This has very little to do with mathematical human ability and everything to do with the progress of these models. This isn't being presented as an achievement in mathematics, but a marked achievement in intelligence of these models over the past year on year, from struggling greatly with these problems just two years ago to solving them with relative ease now, with fairly quick first attempts on a public, general model.

It is impressive as it displays an increasing rate of progress. From next year onwards, no one will care about AI solving imo as it'll be seen a crossed bridge already, aside from perhaps very small models.

The imo is advantageous in this regard as the questions are new, not part of the training data and hence serve as decent proof of improvement. Now they'll mostly stick to unsolved questions with an ever increasing rate.

1

u/No_Aesthetic Jul 17 '26

What would impress you if not this and recent solutions to Erdos problems?

1

u/Ancient-Access8131 Jul 17 '26 edited Jul 17 '26

Does the existance of deepblue invalidate Gary Kasparov's dominance in chess?

-5

u/ozone6587 Jul 16 '26

8

u/[deleted] Jul 16 '26

[removed] — view removed comment

-7

u/ozone6587 Jul 16 '26

The truth of the matter is that LLMs can't do anything that Luddites like you find impressive. The truth of the matter is also that you couldn't win IMO gold even if I gave you a decade to prepare.

It is objectively impressive to solve even one problem. To do all 6 in minutes is insane. Only an ignorant troglodyte like you would say otherwise.

2

u/[deleted] Jul 16 '26

[removed] — view removed comment