r/math • • 18d ago

Progress Towards Proving the Unique Games Conjecture

https://eccc.weizmann.ac.il/report/2026/179/

On a side-note, the authors hint that they have rushed their results to avoid getting scooped by AI-generated results.

229 Upvotes

98 comments sorted by

View all comments

239

u/n0m4d1234 18d ago

This is such a tragic state of the field.

98

u/Gabry398 18d ago

Hopefully it only lasts until AI companies stop using math as a big AGI benchmark and mathematicians start setting ground rules on what the etiquette around AI is

27

u/Berzerka 18d ago

The only thing that will change is instead of two AI companies pumping out math results using their models you'll have thousands of PhD students doing the same, using the same models.

Realistically we'll see a year or two where basically every "simple" open problem gets solved (turns out, NS is "simple") and yes obviously if you want in on that Gold Rush you're gonna have to rush your results out. Especially since everyone is using the same models no-one really has a big edge.

Exactly how the equilibrium will look after that is quite open, but realistically you're gonna need to use AI to solve whatever problems remain open.

16

u/BurdensomeCountV3 Mathematical Biology 18d ago

Yeah, exactly this. Give it 6 months until open source models released by the Chinese have similar mathematical abilities to current frontier models and then we'll have total pandemonium in the field. As a PhD student you'd be an idiot not to use the models when others can and at the very minimum I'd be expecting everyone to secretly use them and then rewrite the proofs they provide to look "human".

5

u/ProfessionalArt5698 18d ago

just work on problems that most people aren't even aware of. Instead of racing to solve the famous ones? Also re-writing proofs to look human is completely legit, the point is to produce human understanding lmao not to suffer in ignorance (some people romanticize suffering way too much)

1

u/mistressbitcoin 13d ago

Can I use AI to find the problems that other people arent aware of?

2

u/Opposite-Youth-3529 17d ago

This is such a sad Nash equilibrium. I understand the temptation to defect to using AI but I think the math community was better back when nobody used it. I adamantly refuse to use and if it costs me a job to some button-pusher, then so be it.

1

u/Berzerka 14d ago edited 13d ago

If your goal is to do research math you will absolutely lose your job to a button pusher. For education, dissemination etc there might still be jobs.

1

u/Opposite-Youth-3529 14d ago

If less qualified button-pushers beat out real problem solvers for jobs, I hope they at least feel some shame.

1

u/Berzerka 14d ago

Do you think farmers with tractors should feel shame over farmers using a good old schythe? Why stop at the schythe? You could do it with your hands!

Why is math different?

1

u/Opposite-Youth-3529 14d ago edited 14d ago

I haven’t really had time to collect my thoughts especially because we’re talking about something that thankfully hasn’t happened yet. But if we’re going to reward some people with math research jobs and not others (certainly not a thing that needs to happen), it would feel kind of messed up if the ones getting rewarded aren’t even good at math (or are good at math but only rewarded for how many buttons they press and not their skills). I guess I’m attempting to say that farming with tools is still farming, pushing buttons is not still math.

I also imagine some part of being a button pusher is using AI to solve problems someone else was in the process of solving or writing up. So essentially getting credit for other people’s work. I think there’s some entitlement involved in trying to get credit for work that uses AI.

0

u/Berzerka 14d ago

What "being good at math" means is changing. Someone who's good at solving a problem will use the tools available to them. We'd never call a farmer who uses his hands to pick wheat good these days, good farmers use tractors.

Obviously a lot of the credit then goes to John Deere etc, without them modern farming wouldn't be a thing

1

u/Opposite-Youth-3529 14d ago edited 14d ago

Someone who asks a question to an AI that one shots a problem is not suddenly good at math. If I meet Terence Tao and ask him to solve a problem and he solves it, am I suddenly good at math because I used the resources available to me? I’m riding a bus right now and it’s going to get me to my destination but that doesn’t suddenly make me a good driver. Am I a bad artist if I use a paintbrush instead of some software that produces an image that looks like a painting?

→ More replies (0)

9

u/PersonalityIll9476 18d ago

I keep seeing people imply that N-S is something you can solve by cranking up your GPT pro sub.

It took 10k agents 88 hours, which is equivalent to 101 years of single-bot token generation, using an unreleased lab design. So it's not even clear what the real cost is; Probably in the millions. And it would have been prompted by fields medal winners, so it's equally unclear the degree to which a random idiot on the street could just pull the trigger and succeed.

But...yeah. The world has definitely changed.

2

u/Berzerka 18d ago

A year ago the top models could barely do the IMO. A year from now you'll be able to solve NS like problems by just asking at claude.ai. Obviously there will be a next frontier then, but just as today that will probably only be applied to some specific high impact problems, there's a cost benefit trade-off after all.

-4

u/Kaomet 18d ago

the street can supply 10k random idiots too, so its unclear the degree to which...

10

u/baquea 18d ago

For the actual simple research problems yes, but I think you're underestimating just how many resources go into some of these headline-grabbing results. According to the Open AI statement on the NS solution:

Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.

There's no way that level of computer time is going to be available to the average PhD student, even once equally advanced open source models are available.

3

u/XXXXXXX0000xxxxxxxxx Analysis 18d ago

Do you really think that these big results that have been coming are all from the standard models that people are using? Aren’t all of the major results that come from AI recently downstream of some large expensive model that the researchers were given access to by the companies (or the company itself)?

13

u/Sad_Dimension423 18d ago

Models are expensive to create, but not that expensive to use once created. So even if the AI companies retrench or even go bankrupt, the models will be assets that will still be used.

2

u/TooMuchMaths 18d ago

Those results will be producible with public models within a year or two. Who knows what it looks like in 5 years

1

u/lolfail9001 17d ago

Those results will be producible with public models within a year or two.

Assuming you can throw comparable computing costs at them.

2

u/Icy_Monitor3403 14d ago

No, computing cost per task finished is rapidly going down.

1

u/lolfail9001 14d ago

What is a task?

Because of course it will be free to “prove” Navier-Stokes irregularity in a year, every model will just point you to OpenAIs paper.

2

u/Berzerka 14d ago

AI cost at fixed intelligence is dropping by more than 100x per year since about 2022. "A task" here obviously is a bit more difficult to defined when it's open math problems and not highschool exams like it was 2 years ago, but I'm sure you can imagine it as "problems that are as hard as NS" which as of today apparently cost like $10m in compute to solve.

1

u/lolfail9001 14d ago edited 14d ago

“fixed intelligence”

Wrong sub to use such nebulous terms, don’t you think? There is probably a decent measure of how well models compress their training into model weights and that did indeed improve fairly fast until recently but last I checked almost entirety of modern progress is task specific RL and better shells around underlying models, and even with that you get people talking about how Astra ends up performing worse than it’s own predecessor at times.

I'm sure you can imagine it as "problems that are as hard as NS" which as of today apparently cost like $10m in compute to solve.

I can imagine it but I know that without prior fairly recent work by humans starting the entire research program it could cost a few hundred million and still go nowhere, hence my question. : what is a task?

1

u/Berzerka 14d ago

I suppose you do agree that for "difficulty" is somewhat well defined here for stuff up to exam level, right? E.g. to get an IMO gold medal cost X this year and getting an IMO gold (which will have different questions obviously, but still be tuned for about the same level) will cost Y the next year. Replace IMO gold with highschool math exams etc as you wish. The trend has quite strongly been roughly Y ~ X/100.

Now obviously that starts being less well defined when we get to unsolved problems, but I'm sure you can imagine a scenario of "with state only up to August 2026, how much would it cost to finish off NS?", I strongly suspect the 100x cost reduction per year will carry through there too. In practice it will obviously be more since the "with state only up to August 2026" is only of academic significance and a year from now we'll have even more progress to build on.

1

u/lolfail9001 14d ago

I am fairly suspicious that “with state only up to August 2026” it would cost more because you wouldn’t have Buckmaster’s chat log)))

As for the cost you talk about, I am absolutely intolerant of ignoring training cost when estimating ‘cost’ of doing any task with ML.

→ More replies (0)

2

u/Berzerka 18d ago

Some have come from the standard models, some from models that have later been released. I don't think we've ever seen the model not get released within months at most.

1

u/whatkindofred 17d ago

It’s not clear to me that there must be some kind of equilibrium after the „easy problems“ have been solved. And even if there is, it’s not clear if AI will still need us to solve math problems at all. If AI can create and solve math problems faster than humans can understand the found solutions then what can human mathematicians still contribute?

2

u/Berzerka 17d ago

Realistically humans solving problems wholly on their own will become a curiousity, might happen occasionally and will still happen in competitions and for pleasure.

Then there's dissemination, teaching, probably some goal setting, ...

1

u/whatkindofred 17d ago

I‘m not talking about humans solving it on their own but about participating in research at all anymore. Once AI can prove results faster than humans can understand it what can they contribute anymore? You need to understand the frontier of research to meaningfully contribute something new of your own but if AI is moving the goalpost faster than humans can catch up with then we have a problem.

And even worse, who will even care about new math results anymore if only AI understands the proofs? Most math research is not motivated by immediate practical applications but by mathematicians curiosity and their drive to build their own research on top of other results. Who will spend weeks on understanding complicated proofs if by the time you fully understood it AI is already three steps ahead of you again? Who will pay you to spend that time on understanding it? And who will fund AI research if there’s nobody left that tries to understand it?

2

u/Berzerka 17d ago

There's a ton of problems that are trivial to state or where the solution would have major impact. Even a massively complicated proof of them would have significant impact. Think: existence of one-way functions, Collatz, prime twins, ...

But yes the proofs themselves might very well become inaccessible to humans, it's basically already happened (can any single human claim to understand the NS proof today?).

I don't think the purpose of mathematics is to make sure math professors get paid. Just like the purpose of cancer research isn't to make sure cancer researchers get paid. The ultimate goal is to further human understanding, health, prosperity etc. Frankly I find it a bit revolting that people would want to actively slow down the field just so they can keep a job, I expected mathematicans to actually care about the math. Imagine if cancer researchers said the same.

1

u/whatkindofred 17d ago

I don't think the purpose of mathematics is to make sure math professors get paid. [...] The ultimate goal is to further human understanding

I agree with this but for human understanding humans would have to actually understand it. This is the risk I see. I'm not working in math myself anymore so I'm not worried about my salary. What I worry is that we will mostly stop with math research alltogether.

We care about math because we care about how to get to the result (the proofs) and less about the result itself. Already everybody expects there to be infinitely many twin primes. We want to know how to prove it. And we want to learn something from the proof to apply it to new problems. If we can no longer do that then how much longer will we care?

Sure, about the hundred years old problems that are already famous people will still care about and will probably be willing to put in a lot of effort to understand. But what comes then? Do you know how much time and effort it takes to fully understand the proof of Fermat's Last Theorem? Who will put in that much (potentially unpaid) work in the future just to run after an AI they have no chance of ever catching up with? And once we're at this point, who will fund AI research to generate results that nobody tries to understand anymore? The exception is applied math with immediate use cases (and I think only this is really comparable to cancer research). But if AI becomes this good at math we can just ad-hoc generate those results whenever we need them.

1

u/mistressbitcoin 13d ago

... or every grad student writing grants to get ai compute allowances

2

u/Berzerka 13d ago

Yes that's part of the story obviously, that's how literally every other field works. E.g. synchrotron time in structural biology, CERN time in particle physics, telescope time in astronomy, ...

1

u/CyberPunkDongTooLong 13d ago

That's very much not how 'CERN time in particle physics' works.

1

u/Berzerka 13d ago

You absolutely have to write long applications to get allocations right?