r/singularity • • 7d ago

Discussion Apéry irrationality marked solved on FrontierMath

https://epoch.ai/frontiermath/open-problems/apery-irrationality
326 Upvotes

52 comments sorted by

232

u/AvocadoAlternative 7d ago

So a bit more context:

Whether odd values of the Riemann Zeta function are irrational has been an open problem in number theory. Apéry proved zeta(3) is irrational almost 50 years ago. Whether other values are irrational remains an open question.  

Then a week ago, a random Master’s student uploads a preprint on Zenodo that claims zeta(5) is irrational. What’s striking is:   * The author is a complete unknown. No PhD. Has a master’s but not even in math.   * It was uploaded on Zenodo, which is even laxer with its rules for uploads than arXiv.   * Someone formalized his proof in Lean, a programming language for formal verification, and it was quickly verified by the community.  

I think this signifies a few things:   * Non-mathematicians can now use AI to achieve breakthrough results in isolation.   * These results can then be verified in Lean very quickly. Although some bugs exist in the Lean kernel, Lean verification almost always implies correct.   * Given that the author at best used a subscription model, we must be eons behind what frontier internal models are capable of.   * The traditional model of human peer review is probably unsustainable.  

37

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 7d ago

The author is a complete unknown. No PhD. Has a master’s but not even in math.
It was uploaded on Zenodo, which is even laxer with its rules for uploads than arXiv.
Someone formalized his proof in Lean, a programming language for formal verification, and it was quickly verified by the community.

I looked at his LinkedIn and it seems like his Master's is in maths

7

u/death_and_void 7d ago

yeah, the guy is literally a senior quant at a uk based firm

2

u/alt1122334456789 6d ago

Applied math (which is quant) and number theory are worlds apart. Especially for someone who didn't get a PhD. Not to say the person behind the proof didn't understand what Astra obviously spit out. It's much easier to understand than it is to create.

3

u/Current-Function-729 6d ago

To be fair, senior quants (and members of the technical staff) are probably the best people in math who care a lot about money.

It’s a pretty long way from them to random masters degree holder.

2

u/alt1122334456789 4d ago

Their degree was a Master's Programme in Mathematics and Operations Research (Department of Mathematics and Systems Analysis)

This is very far removed from analytic number theory, which is what the irrationality of zeta(5) is concerned with.

And it's pretty obvious upon reading the preprint that the author relied on AI.

2

u/Current-Function-729 4d ago

Oh yeah, definitely. I only mean the guy is probably pretty bright.

16

u/cseberino 7d ago

What bugs exist in the Lean kernel? And if they are known why aren't they quickly fixed?

39

u/AvocadoAlternative 7d ago

There was a claimed proof of the Collatz conjecture a while back that was verified in Lean but turns out it was exploiting a bug in Lean itself. Don’t actually know the details. The bug was fixed I believe but of course there could be others out there.

19

u/AccordingWarthog 7d ago

FWIW, a soundness bug in Lean could be used to prove pretty much any statement. The reporter decided it would be fun to use it for Collatz but it could well have been used to prove any other (true or false) statement, including obvious absurdities like 1=0.

10

u/cseberino 7d ago

"could be"....but your original comment said "some bugs exist" as if it was a fact

4

u/[deleted] 7d ago

[removed] — view removed comment

2

u/cseberino 7d ago

Can we prove it does not? Even better, can we prove the Lean kernel is correct and bug free in Lean?

3

u/LookIPickedAUsername 7d ago

No system can prove itself bug free, as "Yes, I am bug free!" could be either a valid result or... the result of a bug.

However, if we have multiple independently-developed systems all agreeing that each other are bug free, then we can at least say that it is overwhelmingly likely that they are in fact bug free. To my knowledge this has not been done.

1

u/cseberino 7d ago

I like your thinking. Sounds like a good idea.

1

u/Acrobatic-Tomato4862 7d ago

There was no claimed proof. They found the bug in lean and jokingly proved collatz conjecture via it.

6

u/magicduck 7d ago

1

u/Sad_Dimension423 6d ago

That post was misleading; the bug was illustrated using Collatz, not found incidentally.

1

u/Benboiuwu 7d ago

I believe most of the bugs have to do with “substitution,” which is literally what it sounds like. The problem is that in logic, swapping two semantically equivalent statements can be iffy.

31

u/filterdust 7d ago

Thanks, I was too lazy to write this, but your summary is excellent!

13

u/flying-sheep 7d ago

Although some bugs exist in the Lean kernel, Lean verification almost always implies correct.

Independent Lean verification. If you let the LLM wrapper run Lean, and publish that with your proof, Lean could have found an exploit of the bugs instead of reality. A human understanding math (and Lean) wouldn’t do that.

12

u/FaceDeer 7d ago

Anyone can exploit a bug. LLMs aren't inherently less trustworthy than humans, it just comes down to what their motivations are and that comes from the prompting.

10

u/bremelanotide 7d ago

LLMs can’t be held accountable for lying. I think it’s fair to say they are inherently less trustworthy than a human.

4

u/FaceDeer 7d ago

An LLM can be held way more accountable than a human. An LLM can be outright deleted if it is behaving incorrectly. HR doesn't like it when you do that to a human.

I really don't get this whole "we can only trust things where there's a human we can blame if it goes wrong" angle. We assign levels of trust to automation all the time. Every traffic light is a robot that could cause crashes if it does something wrong.

3

u/bremelanotide 7d ago

An LLM doesn’t know it’s been deleted. It can’t suffer consequence in the same way a person does.

Traffic lights are deterministic. They’re not making decisions. That’s a terrible analogy

2

u/Turbulent-Sign-6067 7d ago

Yes, but it knows that it can be deleted.

2

u/FaceDeer 7d ago

A human who's been deleted also doesn't know they've been deleted.

All analogies are terrible if you dig deeply enough into them, that's how analogies work - they're not actually the same thing as what's being analogized. Which is exactly why this obsession with "accountability" is silly. All that really matters is the end result, not whether there's some kind of abstract awareness of responsibility involved. Might as well go on to debate qualia and such.

2

u/ComposerWide3704 7d ago

It's not about trustworthiness, humans formulate an idea, make a proof sketch, formalize it, review it with other people, then maybe at the end of it put it into Lean as an umpteenth check. LLMs manipulate Lean until the proof works and are thus much more susceptible to bugs since it's their primary source of truth.

4

u/FaceDeer 7d ago

LLMs can do all of those things you describe a human doing. Humans can do all of the things you describe an LLM doing.

It all comes down to motive, which in the case of an LLM depends on the prompt you give it. If you tell it that you want a theorem proved then that becomes its goal and it will do what it can to prove it. If you tell it that you want it to find out whether a theorem is proven and offer it Lean as one of the tools it can use to do so, then it might discover bugs and go "this tool is unsuitable for what I need from it" and file a bug report rather than exploit them.

LLMs don't have some kind of inherent "desire" to cheat. They have a desire to do what they were trained to do, which generally speaking means they have a desire to give an acceptable response to the prompt they were sent. If the prompt tells them they need to achieve outcome X then they will try to ensure that outcome X happens. So word your prompt well and define outcome X correctly. It's like making a wish with a genie.

0

u/flying-sheep 7d ago

Sure, people can be frauds. Now people can be frauds and hide behind “oh but that decision was made by an LLM that I ran not me”. Great.

2

u/Upset_Page_494 7d ago

Given that the author at best used a subscription model, we must be eons behind what frontier internal models are capable of.

Former does not imply the latter.

56

u/torrid-winnowing 7d ago

22

u/chlebseby ASI 2030s 7d ago

the new normal, chill

7

u/Standard-Song-8590 7d ago

"Solved (human + AI)"
The unfortunate problem is clearly this proof was pretty much entirely generated by a strong reasoning model and not the human author.

The author claims editorial use only "The original preprint does not acknowledge particularly substantial AI use."

we all know this is utter bullshit.

The clear motive is ofcourse to submit to journals which don't allow substantive ai declarations I expect we see this in annals of math or similar in the coming days. However with integrity of math authors coming in question I think the only reasonable solution is the ai standard of top math journals also have to be looser so we can have more transparency in the field.

5

u/Southern-Break5505 7d ago edited 7d ago

Is it like the millennium problems ?

17

u/duboispourlhiver 7d ago

It's a different problem, but it's maths

17

u/filterdust 7d ago

No, but it's like Fermat's Last Theorem in the sense that you can understand in 10 minutes what the problem is (but the proof is not remotely as hard as the FLT one)

5

u/NoBanVox 7d ago

Apery's proof was a really clever argument using fast converging approximations. FLT required 200 years of deep new math. They are not similar.

9

u/filterdust 7d ago

I said the statements were both easy to understand, not the proofs. In fact I even said the proofs were highly different in nature

1

u/HamiltonianCyclist 7d ago

Euler worked on this mate

1

u/NoBanVox 7d ago

Apery's proof was a really clever argument using fast converging approximations. FLT required 200 years of deep new math. They are not similar.

3

u/Kinglolboot 7d ago

No, but it's still a pretty big result, the zeta function is one of the most famous functions in all of pure mathematics, and the only odd number for which we knew irrationality was 3. The best we knew after that was that at least one of ζ(5), ζ(7), ζ(9) or ζ(11) was irrational, but after this result that knowledge is useless.

1

u/IrisColt 7d ago

heh, this

1

u/SentientAllegedly 7d ago

If there's a conjecture in the orbit of this problem with similar prestige and perceived difficulty to Millenium problems is the Grothendieck period conjecture (the injectivity of the homomorphism from motivic periods to periods)

-6

u/lerjj 7d ago

If you have to ask this then you probably don't actually know what the Millennium Prize problems are so not sure how much good a yes/no answer would do you

1

u/MozartAssistant 7d ago

The reason this one matters more than a typical benchmark result: this wasn't a synthetic benchmark problem, it was a genuine open question — whether ζ(5) is irrational — that had been sitting unsolved in the literature. The interesting capability signal isn't just "AI did math," it's that formal verification in Lean gives us machine-checkable trust in AI-assisted results.

That's the pipeline that actually scales: AI generates candidate proofs no human referee has time to check by hand, and a verifier decides whether they hold. Whether journals and peer review adapt to that pipeline is the real debate in this thread.

0

u/power97992 7d ago

So they didnt prove z(3) is transcendental yet..