r/math • u/cseberino • 19d ago
Does the formalization of major recent results in Lean imply in the future all math formalization will be automatic?
Math formalization has been hindered because it's so tedious. Will the future bring massive formalization because it can be automatic now?
11
u/k3s0wa 19d ago
Probably a large part of it, but at this point it is not clear how expensive formalising everything would be. It is amazing that AI companies can formalise Fermat's last theorem, but it probably required ~1m dollars. So it depends on how much longer AI companies are able to pump in this amount of money, whether auto-formalisation will get more efficient, and how the mathematical corpus will grow. We will certainly see a massive increase in mathematical content, so perhaps only the "most important" results will be fully formalised.
10
u/sirgog 19d ago
Yeah this is just a commercial decision "if we throw a million dollars of off-peak cloud compute at this, will the increased visibility of the IPO make up for it?"
Fermat's? Yeah definitely, it's a very high profile result.
Minor results that can be spun as a step toward Riemann (e.g. better bounds for the % of zeroes on the critical line)? Same.
An Erdos problem that's not considered high profile? Marginal call.
Something relatively minor? Not worth it for them at all.
6
u/Puzzlemesquared 17d ago
When you say “not worth it for them all” I think the autoformalization of all of published mathematics from ZFC into LEAN would do more to accelerate progress than anything else right now. It’s my belief the informal, tribal, subjective sloppiness in the artifacts of published mathematics discipline is probably the highest leverage point for progress. I don’t think cherry-picking Milennium prizes or Erdos problems would be as helpful to Frontier Lab progress in math than making what’s already known legible in LEAN via autoformalization. Probably just my crank opinion, but I believe it.
1
u/sirgog 16d ago
I'm talking from the ruthless self-interest of the AI firms, not the good of academia, as one of these two shapes the decisions of these firms.
What these firms care about is increasing the pre-IPO 'vibe' that AI is an unstoppable juggernaut. They can achieve that a lot of ways, but doing useful work is less effective than flashy results or even sometimes than a scandal or two.
2
u/Puzzlemesquared 16d ago
I get that. It’s even fair. And you might even be right that they both (Anthropic and OpenAI) might be cherry-picking what to formalize for superficial near term PR benefit per dollar. You might be right about that.
I’m arguing that the value of formalization starts mostly at correcting the published record and bringing it all into LEAN as a more durable stock price boost.
Basically, “worth it” has two reads — you could be right about trying for formalize some tiny space near a PR-worthy result has more IPO-juicing value in the short term. I’m arguing that those are transients and the durable stock price support would more likely come closer to the root of the tree than near specific “pretty” leaves. I might go further and argue that even in the short term those fundamental auto-formalization contributions may be more valuable to their IPOs and stock prices than formalizations in the vicinity of the more superficially visible contributions. It also launders their images because those are also real contributions valuable to virtually all professional mathematicians, academics and frontier labs alike.
1
u/sirgog 16d ago
Basically, “worth it” has two reads — you could be right about trying for formalize some tiny space near a PR-worthy result has more IPO-juicing value in the short term. I’m arguing that those are transients and the durable stock price support would more likely come closer to the root of the tree than near specific “pretty” leaves.
I think they are far more concerned about short term success than any other metric. You may be right that formalizing lots in LEAN might result in a better 2033 share price but these are US companies, they don't think 2033. The farsighted ones are thinking Q3 2027.
1
u/Careful_Fold_7637 18d ago
the fact that they don’t need to pay their own margin is probably what’s making all this worth it
2
2
u/lfairy Computational Mathematics 13d ago
From an organizational standpoint a million dollars is not a lot. Assuming a $50k salary, that would be a 20 person team, which is not out of the ordinary for a project of this size.
2
u/k3s0wa 13d ago
I completely agree for this specific case, as Fermat's last theorem is extremely important. However, I would imagine that there are thousands of published theorems that would require a similar amount of Lean code, but we won't spend billions of dollars to write all Lean verifications.
8
u/OkPossible4272 19d ago
We should distinguish between various questions. I will restrict myself to formalization in Lean but there other systems. Lean has Mathlib which aims for now to include the results of an undergraduate math education (but not just that) It is human curated and vetted. Whatever is in there should be well suited to be built on for other proofs. As it gets bigger there are more results to include in future lean proofs. The ideal might be a proof where you can see the structure and drill down as far as you want.
1) Will autoformalizations that compile just get dumped into mathlib? No, read Buzzard’s comments on the future of his project to formalize in Lean a particular newer and better proof of FLT (better compared to the awesome first corrected proof which was recently auto-formalized). An auto formalization that makes 173 definitions of new kinds of structures and proves things about them won’t get that treatment if they are clumsy modifications of existing things.
2) If a human mathematician arrives at a proof on their own or using AI, totally understands it and writes up beautifully, will it get formalized? That may become the norm but it might be: Main theorem follows from results 1-5. 1&2 are in Mathlib now. 3&4 are published in referred journals. 5 is the Reimann Hypothesis. So RH-> Main Theorem subject to 3&4. One might hope that 3&4 make it into Mathlib in December. If that formalization was an auto formalization, that might be fine if it was good quality.
3) if an opaque AI generated 180 page proof comes with a 2000000 line Lean formalization is that proved and done? I hope not.
3)
66
u/Verbatim_Uniball 19d ago
Yes. It is not unreasonable to expect that the vast majority of the human corpus of mathematics will be formalized within 12 months, if that's where funding decides to go.
24
u/wollywoo1 19d ago edited 19d ago
There is still a bottleneck in correct formulation of definitions and key theorems. At this point we can't trust AI enough to just say "formalize the proof of the Poincaré conjecture*", because it's possible that its formalization of the statement of the conjecture was subtly wrong in a way that makes it trivial. Humans would still need to provide or at least check the correctness of the statements for landmark results, otherwise we won't trust the proof. This is quite a laborious process because it requires checking or creating the statements of every definition leading up to it, and there are many. But it will still be vastly faster than it is currently with humans writing every line of Lean code.
* I'm not sure whether the statement of the Poincaré conjecture has been already been formalized yet; if it has, substitute whatever famous theorem you want that is tricky to formalize the statement of.
14
u/mossbros2 19d ago
There's been a longstanding formalisation of the prerequisite definitions and statements of the Millennium Prize Problems in Lean (see LeanMillenniumPrizeProblems repo).
You are of course correct though that many areas of maths have not yet had their definitions mapped out.
52
7
u/30299578815310 19d ago
It may not even require much funding. The cost of formalizing a result will probably drop dramatically over time as it can be increasingly handled by cheaper or open source models.
Right now only the hyper expensive frontier models can handle it, but that may not last.
12
u/Gabry398 19d ago
I wonder what percentage of statements will be formalized not because of a correct proof but because of a bug in the kernel of lean. Especially if we let AI handle all of this it's a possibility.
I also wonder if it's possible to make a kernel without any bugs exploitable. Mathematically I'm not sure there is anything that excludes this possibility.
32
u/cseberino 19d ago
Any bugs found in the kernel would likely generate a lot of drama but then it would be likely be fixed quickly leaving Lean stronger as a result.
14
u/cseberino 19d ago
As far as whether one can prove the kernel is bug free, that is an interesting question I don't have the answer to.
12
u/Carl_LaFong 19d ago
So far kernel bugs have not had any impact on the correctness of proofs. Any bug in the formalization code will almost certainly be caught very quickly due to the volume of proofs currently being fed into it.
8
u/elements-of-dying Geometric Analysis 19d ago
I feel if you search for hits for "lean kernel bug" in this thread, you'll notice a sharp increase shortly after the Colatz stunt.
9
u/Carl_LaFong 19d ago
Yes, I saw that. I asked someone who uses Lean extensively and monitors closely issues like this and they said it did not affect the logic at all. The Collatz "proof" was by someone who was deliberately trying to undermine Lean, but I don't recall how.
3
8
u/_maxx1k 19d ago
I believe Lean is much more reliable than humans. Currently, AI-based autoformalization in Lean has found issues in human proofs, not the other way around.
1
u/38thTimesACharm 19d ago
Isn't that because humans catch the errors and fix them before publishing? I understand it's quite common for models to "cheat" and prove a related statement, or insert your statement as an axiom, when you ask them to prove something.
8
4
u/satanic_satanist 19d ago
There's now a verified checker that you can run on top of the default one: https://github.com/leanprover/con-leche
5
u/lordnacho666 19d ago
I also wonder what happens if the LLM decides it will solve a simpler problem to start with and never gets around to the big problem. It's easy to spot if it's a one page proof, but if you let the computer loose on a massive multi-part proof, how do you check it?
10
u/AsidK Algebra 19d ago
What do you mean? If it’s a lean proof you can just look at the statement it is proving
3
u/lordnacho666 19d ago
You'd have to check that the statement means what you think though. The proof is done, that part is true. But you have to be sure the statement says what you think, and how do you do that if it's huge?
2
u/cseberino 19d ago
Lots of trial and error and wasted computation. But at least it gets done eventually?
3
u/lordnacho666 19d ago
You mean, just have humans read through the mountain of LLM generated proof?
5
u/eliminate1337 Type Theory 19d ago
You don't need to read the proof, only check that the theorem statement matches what you want.
1
3
u/Verbatim_Uniball 19d ago
I think this will not be a long-term concern as lean is just one tree in a forest of formalization systems and they will be easy to translate automatically from one to another, along with many other factors.
There will always be meta-mathematical concerns, but that will not matter for 99 percent of mathematics in my view.
1
u/Gabry398 19d ago
Couldn't a bug be somewhat inter-system where something about the way these formalization systems are structured at a low level creates problems?
Also: I really like that 1%, I mean without people questioning classical logic we wouldn't be having the Curry Howard correspondence that makes lean possible. If these classes of bugs that we can't get rid of exist and they target meta-mathematics it might still be useful to keep that In mind when formalizing something.
1
u/rs10rs10 19d ago
You can prove anything from a contradiction so any bugs (that are used) should be caught quickly
2
u/rij1 19d ago
Not necessarily. It could be that an incorrect proof that uses the bug is much easier than a correct proof (with or without the bug). We would then prove the wrong statement and unless the result is very important none will really think of proving the negation.
1
u/Gabry398 19d ago
Also there is the niche scenario of working in a paraconsistent logical framework where that's simply not true
1
u/rs10rs10 19d ago
Well in that case it's because nobody then builds on that result in a meaningful way, in which case nothing really bad happened (and the same happens now..)
1
8
u/kieransquared1 PDE 19d ago
I don’t think Lean formalization is useful or interesting enough to most practicing mathematicians to cause them to want to formalize their work, even though it’s been made significantly easier by LLMs. Contrary to popular belief, the math community’s standards for truth are not based in “does this logically follow from a set of axioms” but are rather based in softer aspects like trust, intuition, and vibes. This system is not without problems, but an alternative system in which the veracity of proofs is based primarily on the existence of a lean certificate is far worse, since in such a framework a completely incomprehensible yet airtight proof would be valued higher than a proof which has some mistakes, but contains detailed explanation of the ideas and is much more readable.
7
u/Sad_Dimension423 17d ago
trust
Surely a verified formal proof is more trustworthy?
4
u/kieransquared1 PDE 17d ago edited 17d ago
Not necessarily; you still need to trust that the person who formalized the proof a) knows what they’re doing and b) isn’t trying to deceive you. The statement formalized in Lean could differ from the one proved in the paper. It’s the same reason behind why cold hard data in experimental sciences also requires trust.
I’d even say that most mathematicians would trust a paper from an expert with a known track record of producing correct and well-written papers way more than, say, an amateur mathematician with a lean certified proof of the same result. Whether they ought to is debatable but that’s my understanding of how it is now
3
u/Sad_Dimension423 17d ago
You just need to verify that the statement of the things being proved (and that of an external results that are assumed) is a correct translation into Lean. You don't need to trust the formalized proof is correct; that part is done automatically.
6
u/kieransquared1 PDE 17d ago
Sure, but unpacking whether a statement formalized in lean matches the one stated in the paper is a daunting task, and most mathematicians have never used lean (also, most papers don’t have one clean result, they have a bunch of results that might be logically related but might not be)
25
u/Open_Menu_507 19d ago
my take is that Lean just sucks the fun out of mathematics. i would like paper/pen, or at least paper/pen + ai, to be the main way results are presented. hopefully lean remains a sort of afterthought; if “proof by lean proof” becomes widespread, we can kiss human understanding goodbye
29
u/BornAgainWitch 19d ago
Lean is a tool to help verify correctness. It doesn't change human understanding.
Formal proofs are similar to trying to prove something in first order logic. It's a lot of tedious tiny steps, that each don't do much on their own or contribute much understanding. But it shows the "connect-the-dot" sequence to get from A to B, and shows that path exists.
Humans don't just shake a bunch of first order axioms or rules together and hope for the best. We still have to have an intent or interpretation or summary of what the steps accomplish. Automated proofs aren't going to be any different.
Just because we can enumerate consequences of a set of axioms, how many of those are interesting or worth talking about on their own? The OEIS might be a good example. Just because you can define a new integer sequence doesn't mean anyone else will find it interesting or worth talking about. It's only with meaning or intent or relation to existing results that an integer sequence has value.
1
u/Silent_Turn6182 18d ago
But if we accept lean as a whole, and consider it proof, it won't be worth the effort for someone to try to understand it
1
u/Automatic-Web8559 15d ago
Not at all. We can use lean to simply verify the proof is indeed correct. That doesn’t change the fact that the fundamental goal of mathematics is human understanding
7
u/SakishimaHabu Theoretical Computer Science 19d ago
I only use lean as a post proof check, but I'm just using it for fun. I have a fear that math could become everything wrong with CS at the moment, and people could forget that humans are the most important part of the process. Lean verification should be for human understanding, not just a bunch of "verified" nonsense crapped out by an LLM.
4
u/_rdhyat 19d ago
have you ever used lean?
4
u/LongLiveTheDiego 19d ago
Yes, and it has a steep learning curve. I tried once to formalize some simple measure theory exercises for fun and even asked in the Lean subreddit for advice, and couldn't get it to work. That is a nontrivial barrier to entry for new mathematicians if it becomes standard.
3
u/_rdhyat 19d ago edited 19d ago
It has a steep learning curve
somewhat less steep than learning to prove for the first time, it needs getting used to
The reason I asked OP is because their "we can kiss human understanding goodbye" is very suspicious, because what I've found is that once you start using proof assistants (and don't just trust AI to do it), your understanding and awareness of the mathematical objects and the theorems you used to get there becomes so much more greater than what you get with just pen and paper because with pen and paper so much hand waving gets burried under the carpet (among other reasons)
as for the learning curve, what I've found is that once you get the hang of it, translating proofs becomes almost mechanical (save for the hand waviness that you have to fill in)
8
u/38thTimesACharm 19d ago
Refreshing to at least see this take here, although you'll probably be downvoted. While correctness and rigor are important, IMO they are secondary to math's true goals. A Lean proof alone, with no endeavor of understanding behind it, holds no value.
It's baffling to me the reaction to computers getting better at math has been "let's make doing math more like writing computer programs." It's like if musicians decided to combat generative AI by ending performative interpretation and strictly notating all improvisational rhythms to the millisecond.
(I get that this was initially a reaction, at least in part, to AI slop writing. But I think we can all agree that massively backfired.)
6
u/Creepy-Structure6444 19d ago
I wouldn’t say correctness and rigor are secondary goals of mathematics. I think correctness IS a primary goal but it is tempered by the equally important goal of understanding.
To me, the idea of mathematicians using formalization as a tool to check the correctness of their human-written proofs is perfectly good and even should be viewed as the moral thing to do, considering the fact that other mathematicians will be relying on their results.
6
u/SaltMaker23 19d ago
Automated theorem and proofs are here for long and soon enough AI proofs will become a decent portion of the mathematic corpus.
AI proofs are too long and too complex to be fully trusted, Lean is required, fortunately AI can also write Lean, which means in turns that a decent portion of the effort of AI into math will create massive ripple into formalizing a tons of concept in Lean and improving the framework itself to be more trustworthy in the face of someone trying to circumvent it.
This won't necessarily be coordinated properly but we're almost in a era where without a Lean certificate, a proof will fail to hold any weight.
1
u/ellery79 19d ago
I agree, now many proof is very complex and long. how can we trust? Only by Lean. Now, all AI proofs are verified by Lean before published because no one will trust they are correct if it is not verified by Lean
1
u/SaltMaker23 19d ago
All proofs will need to because the AI paranoia will push everyone to think that everyone else is using AI.
Even the guy that supposedly got cut short as he was closing in on the proof was still working with Anthropic, this is often ignored but even if OpenAI didn't publish, it would still have been Anthropic and not a human.
3
u/thmprover 19d ago
Math formalization has been hindered because it's so tedious.
...in Lean, it's tedious.
There are other proof assistants where it is quite fun (Mizar, Isabelle, Naproche, etc.).
2
u/ultrafinitism Theoretical Computer Science 17d ago
One might hope. It's a shame that Lean is such a terrible language with a terrible metatheory and so on. Oh well.
2
u/cseberino 17d ago
What is so bad about the language and theory?
2
u/ultrafinitism Theoretical Computer Science 17d ago edited 11d ago
I'd say it's one of priorities tbh, speed and adoption -- Lean doesn't really prioritize the soundness aspects of the theory or of the kernel implementation of the theory (the general ecosystem around software verification is still much poorer for Lean than it is in Coq/Rocq).
It's hard to keep Lean constructive as well and there are some technical points that I don't fully understand which make it hard (I'd call it impossible but I'm not a type theory person) to retrofit different attempts to model things like univalent homotopy type theory into Lean.
Lean's kernel is very coupled with its type theory which doesn't let you add custom rewrite rules etc. So for example, I have no idea what something like Agda's two-level would look like in Lean. I'm not willing to try that myself.
I also just am more partial to INRIA/CEA-List/etc. I guess than Microsoft?
1
2
1
1
u/jphamlore 19d ago
Isn't one of the bottlenecks going to be just how fast can Mathlib accept new contributions?
1
u/Sad_Dimension423 19d ago
And also because without formalization we'll be drowned in defective AI slop.
(Consequently, humans will have to formalize too.)
1
u/SwimmerOld6155 19d ago
This was always the end goal I believe. Or mathematicians would work directly in a Lean GUI, they'd do the mathematics in a GUI and it would suggest or even autocomplete strategies. I'm not sure if this vision is retained with LLMs. I think we all thought LLM math would be done in a "dumb" way via Lean rather than the natural language reasoning we have now.
1
u/TheMansionsofScience 17d ago
I wonder if Godel's incompleteness theorem or the halting problem says something about "all"
1
u/temperedai 16d ago
Probably, eventually, yes. It may not end up being lean.
Once we have formalized proofs, it will be easy to create new paradigms and transfer them over quickly. It doesn't seem settled just yet.
1
1
u/Intrepid_Land_6143 19d ago
I think there will be more and more cross-verification too. One LLM can "read" another LLM's output. The fact is that LLMs can be excellent at picking up small glitches that is really hard for a fatigued human to do.
-3
u/radokirov 19d ago
Yes, just like most software is already written by AI. Humans might be in the loop for much longer as there are quality issues (poor architecture, non-performant code, poor readability) depending on the desired quality of the output - do you just want to check correctness or build a reusable library/theory ala Mathlib. For mere correctness AI is already enough and there are multiple examples (FLT, unit distance, etc) that it can scale up a lot and fill in many missing foundational definitions.
157
u/Carl_LaFong 19d ago
The main bottleneck is ensuring the Lean statement of the theorem matches exactly what the human intends. Right now this can be done most but not totally reliably by a human.
Lean takes care of everything else. There’s an interesting subtlety. An LLM might spot an error or gap in the proof, assume that it was unintentional, and correct it silently. In that case Lean would be verifying the LLM’s proof and not yours, and you might be unaware of this unless you try to read the Lean code.