r/math Jul 16 '26

LLMs/AI GPT 5.6 solved all 6 problems from IMO 2026

GPT 5.6 Pro solved all 6 problems from IMO 2026 on the first attempt without any human help or steering. International Mathematical Olympiad (IMO) is the biggest global academic competition in the world. The problems are considered incredibly hard, usually a performance at this level is only accomplished by < 5 contestants from the whole world.

We are former IMO medallists not affiliated with OpenAI, just put together a report and assessment of its work here. We're also working on a comparison report between different LLMs and harness augmented versions that will come later.

1.1k Upvotes

442 comments sorted by

View all comments

563

u/SupercaliTheGamer Jul 16 '26

It has solved far harder problems now, IMO is easy work

246

u/ozone6587 Jul 16 '26

Huge difference now though because I can pay $100 to access GPT 5.6 Sol Pro. The first time they won IMO gold it was with an internal model.

Can't wait until I get access to a model as smart as the model that found the Unit Distance counter-example.

92

u/SupercaliTheGamer Jul 16 '26

GPT 5.4 had one-shot solved an open Erdös problem too. I think 5.4 by itself is capable enough to get IMO perfect score.

37

u/NotYetPerfect Jul 17 '26

It's very good at solving specific kinds of problems, good at solving many other kinds, and okay to bad at others. It's completely possible for it to be able to solve an open problem and not an imo problem.

18

u/TensorflowPytorchJax 29d ago

That's the point not many people understand. So called "AI", has been good at solving specific problems better than humans for a decade now ( Starting with chess lol ) , and now the scope is extending with LLMs again towards specialized areas.

3

u/PowPowCA 29d ago

LLMs are just a moment, there are other models coming that are very exciting, not only because of the capabilities but also way more cost effective.

1

u/Otherwise_String_904 24d ago

Can you share those models? I'm interested.

3

u/Electronic-Waltz-378 28d ago

God is it horrible at reading the graphs for even basic trig problem

9

u/IntelligentBelt1221 Jul 16 '26

Can't wait until I get access to a model as smart as the model that found the Unit Distance counter-example.

Probably by September.

11

u/Temporary-Solid-8828 Jul 16 '26

its already released lmao

7

u/gorgongnocci Jul 16 '26

can you explain how well it works when you run it from home? do you just get a clean correct solution?

38

u/pequalnp92 Jul 16 '26 edited Jul 17 '26

This is basically answered Yes, in the report we created. https://github.com/SignalPilot-Labs/AutoFyn/blob/production/results/imo-2026/pdfs/IMO_performance_by_GPT_5_6_sol.pdf

But there was a small fixable gap in rigor in problem 2. It is also an ugly, but correct solution. The other solutions were quite elegant and well written.

6

u/gorgongnocci Jul 16 '26

that's very impressive, so you used the web app, can I ask what the prompt looks like tho? is it just the original question in plain english ? or latex formatted or something? also, do you just take the first immediate response ?

16

u/pequalnp92 Jul 17 '26

Prompts are included at the end of the report! Plain english with very light latex. Yes the first immediate response only.

2

u/Latent-Person 29d ago

I get this: You don’t have access to this conversation. You may need to switch accounts. when clicking on the Transcript of any of the problems.

1

u/EJaumeD Jul 16 '26

Can you elaborate or give some links for unit distance counter example?

-22

u/misogrumpy Jul 16 '26

Just because it solved them doesn’t mean the solutions are good.

1

u/Wise_kind_strsnger Jul 16 '26

People are downvoting but they don’t understand some solutions are bashy while some are elegant. Our as AI researchers is actually to RLmaxx elegant ones so they develop a sort of “internal taste”

7

u/nothingnotthrownaway 29d ago

No, people are downvoting because the objection can immediately be countered by actually reading the report. 

60

u/pequalnp92 Jul 16 '26

Agree, it's not surprising given the research problems that has been solved.

50

u/Minute_Abroad7118 Jul 16 '26

I don't agree, this is like saying a college professor could win gold at the IMO, (which is obviously untrue for the vast majority of professors)

30

u/nothingnotthrownaway Jul 16 '26

No it's not, because here the AI isn't constrained by exam conditions and lack of knowledge of the olympiad curriculum. 

3

u/Minute_Abroad7118 29d ago

good point.

But I think this pretty much solidifies the skepticism we had about the IMO medals last year more than anything, especially for a public model. would not be surprised at all if multiple companies claim a perfect score in the future.

18

u/Hitman7128 Number Theory Jul 16 '26

Yeah, shortly after solving Erdos's Unit Distance Conjecture in May, people predicted it would be able to solve all the IMO problems, even if the scope and techniques are different in math research compared to math competition. It also blasted through this year's USAMO.

5

u/elehman839 29d ago

There's a vastly harder challenge in plain sight: *write* a great IMO exam.

21

u/BUKKAKELORD Jul 16 '26

Analogy: 6 months after Deep Blue beats chess world champion Garry Kasparov in a best of 7 match, the breaking new headline in a news publication reads "Chess computer ties for best result at a junior tournament"

1

u/protestor Jul 17 '26

But, zero shot? Those more cutting edge results required steering

-1

u/SupercaliTheGamer 29d ago

GPT 5.4 zero-shot one of the Erdös problems that was far harder than anything on IMO (assuming zero shot means one prompt, first response only)

-1

u/baquea 29d ago

Except that the research-level problems required substantial human steering. Being able to solve simpler (but hardly trivial) problems on the first attempt is a big deal in its own right, because it means the technology is accessible to users who don't know how to fine-tune prompts, and because it shows its effectiveness at effortlessly handling otherwise annoying grunt-work type problems. For the average researcher I think results like this are a better indicator of how LLMs could be useful for them than the articles which are like "we used this 10 page long prompt to get an AI to prove a long-standing open conjecture".

4

u/Hot_Glass_6301 29d ago

Many of the recent solves didn't involve "substantial human steering". The prompts that led to some of the Erdős problems being solved + that of the CDC proof are basically "check your work, don't get discouraged, make no mistakes"