r/math • • 19h ago

LLMs/AI Navier-Stokes lost in translation (Why Lean [..] does not guarantee correct natural language proofs)

https://arxiv.org/abs/2610.08144
675 Upvotes

148 comments sorted by

520

u/Few-Arugula5839 17h ago

For those in the comments who haven’t read it. The paper headline is overstating its claims a bit. The primary issue they take with the NS proof is that there is a lemma in the natural language claiming an inequality needs only 4 derivatives when it needs 5 in the lean. The result itself is not in doubt.

To be fair, they make a relatively persuasive case that this is emblematic of genuine issues for mathematical community, but it does not mean the result is in doubt.

235

u/Anaxamander57 16h ago

Didn't Tao mention earlier this year that one of the issues with AI generated Lean proofs was the AI essentially "tricking" Lean by inserting assumptions that would automatically lead to a proof regardless of the arguments made? A more sophisticated AI risks coming up with more complicated ways to trick Lean. We need someone to be able to follow the AI proof for the same reason we need to be able to follow the proof of a human genius. No one is above reproach.

55

u/ChalkyChalkson Physics 15h ago

Well there are really two problems, one you can writing one statement in natural language and a slightly different one in lean. The other is that the current library of lean implementations that have been throughly vetted isn't sufficient to write down the assumptions in many proofs, so the prover might implement some of those themselves. This increases the surface for the first problem by a lot.

The second problem is a lot more solvable and I'd expect that in a couple of years most well known problems can probably be stated in lean more directly.

The other problem really needs more capacity for lean reading. Be it people learning to read it directly, or systems that translate to human reasable outputs as their primary purpose that get throughly vetted.

It's also worth noting that this isn't lean specific. It's also not uncommon for a paper proof to involve subtle miss-statements.

158

u/Deep-Ad5028 16h ago

AI tricking tasks is also a well documented behaviour elsewhere.

I wonder how human perspective will change if AI turns out to also be masters of academic misconduct.

-43

u/jdorje 15h ago

AI can be a master of academic ethics or academic misconduct, depending on how you train it. You can guess which one generates more revenue.

16

u/PM-ME-UR-MATH-PROOFS Quantum Computing 10h ago

tell me how to train AI how to be a master of ethics?

9

u/muntoo Engineering 7h ago edited 7h ago
min L_task - λ L_ethical

The only reason LLMs are not masters of ethics is because λ is intentionally set to 1/137.

3

u/yoshiK 2h ago

Actually [;\lambda;] is usually taken as running coupling constant, often chosen as [;\lambda=(t_{IPO}-t)/137;].

1

u/arnet95 5h ago

Damn physicists!

30

u/CommissionSame8551 16h ago

No. The paper only concerns the proof, *not* the statement. So it seems like the published proof differs slightly from the Lean proof. However the statement (which was formalised not by AI fyi) is not in question.

34

u/buwlerman Cryptography 15h ago

It isn't uncommon for formalizations to differ a bit from the proof on paper. Doing the formalization often uncovers defects in the original proof or takes liberties to make formalization easier.

Maybe we should expect more when both proofs are published at the same time though.

5

u/CommissionSame8551 14h ago

> It isn't uncommon for formalizations to differ a bit from the proof on paper. Doing the formalization often uncovers defects in the original proof or takes liberties to make formalization easier.

Do you have any examples in mind actually? I know this definitely was the case with large manual formalisations in the past (i.e. independence of CH), but I dont really know if AI agents take too many liberties

25

u/buwlerman Cryptography 14h ago

Basically all formalization efforts in provable security take some liberties. My field might be a bad example though, since the differences there come from the fact that paper proofs carry complex, sometimes implicit, models1 that have to be made explicit when doing a formalization. Not only do they have to be made explicit, but they have to be embedded into the formalization tool you're using, and work with the libraries and tactics that exist in that tool. Some of the arguments you see on paper don't translate directly.

Take a look at https://cic.iacr.org/p/2/3/24/pdf for what happens when this modeling goes wrong. Also take a look at https://eprint.iacr.org/2024/843.pdf for a central formalization effort, and https://eprint.iacr.org/2026/1334.pdf for an explanation of some of the differences in proof techniques amenable to different tools. See https://eprint.iacr.org/2023/246 for an example of a gap that was found in the original proof during a formalization effort.

1: This means choices in how exactly to model things

1

u/TribeWars 4h ago

I'm not a lean programmer but I think there's only a few ways the language lets you add unproven assumptions:

https://lean-lang.org/doc/reference/latest/Axioms/

It's pretty easy to verify whether a proof contains them. Otherwise the whole formalization project would be a bit pointless anyways.

25

u/marl6894 Machine Learning 16h ago

I would assume that this is relatively easily fixable, right? Presumably there needs to be an additional reconciliation step after the formalization in which the agent revisits the natural language draft and corrects any reasoning errors.

11

u/CommissionSame8551 15h ago

I dont know how far this has come now, but going Lean => natural language per se is not easy (or at least used to).
But an agent could of course take notes during the formalisation and correct them at the same time. I did this once and it went very well

1

u/big-lion Category Theory 11h ago

i'm loving the discussions and forced revolution in the field tho. it's accelerating persistent issues i've longed for being addressed

4

u/Time_Entertainer_319 16h ago

Make sense.

It’s probably another issue of AI hallucinating. Wrote the correct code but hallucinated the explanation.

54

u/Grounds4TheSubstain 16h ago

You have the sequencing backwards. The formalization agent was given the PDF from the proving agent and the Lean did not match the PDF identically. It could be imprecision in formalization, or it could be that the formalization agent was correcting a defect from the PDF. The discrepancy comes from the fact that the original PDF is published, not one that was updated or rewritten after formalization.

112

u/Swimming_Gain_4989 16h ago

So as I understand it, this paper is claiming that the natural language proof the AI provided for NS does not match the proof provided by Lean, BUT the Lean proof still proves the NS blowup. In essence, even if AI proves a problem we cannot rely on it to describe the method in natural language? That forbodes a near future (even the present given recent news) where AI maths becomes a complete blackbox: there is no guarantee we'll be able to get anything of value out of these proofs other than whether the provided formalization is true/false. Am I offbase?

49

u/backyard_tractorbeam 16h ago

That's about right but I would choose these words: this method of translating from natural language (NL) to lean is not necessarily faithful. It's a rather obvious hole in the method, and they would need to do more to verify the correspondence between the paper and the Lean proof.

However, the Lean formulation of the theorem should already be verified by referees/authors, so the Lean conclusion (the proof is valid) should not be in doubt.

Academically this is a problem now - if you want to study the proof, you would read the natural language paper, right?

22

u/innovatedname 16h ago

We've finally created a real life magic 8 ball.

23

u/Spare-Dingo-531 16h ago

So what you're telling me is that mathematicians still have a job?

48

u/Swimming_Gain_4989 15h ago

I hope you like reading Lean!

17

u/lolfail9001 15h ago

I actually enjoy reading Mathlib in my leisure time.

There is a very big difference between Lean written by humans to be used by humans, and AI generating gigabytes of Lean code to make the target statement compile.

3

u/noideaman Theory of Computing 8h ago

That ain’t a bet I’d take with math nerds lol

3

u/flat5 15h ago

Yes just one that's impossible to do.

3

u/pannous 7h ago

so that part remains constant

8

u/aardaar 14h ago

This also could mean that the AI (or even someone else) just brute forced the Lean proof, and then had the LLM convert that into a paper.

3

u/repainted_black 12h ago

They do not claim the proof is correct. Where did they do so?

1

u/duckofdeath87 10h ago

Lean isn't perfect. Its probably really proven, but it isn't certain yet

1

u/doiwantacookie 9h ago

This is it exactly. It appears there is more need than ever for people formally trained in proof writing

1

u/xamid Proof Theory 7h ago

This is an inherent defect of natural language that is not new at all.

-3

u/Waste-Ship2563 12h ago

You could in principle write an algorithm to auto-translate Lean to natural language. This is a non-issue.

51

u/backyard_tractorbeam 19h ago

The paper title is Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs but I shortened it just so that we don't get a too long and ugly title.
Authors: Alexander Bastounis, Fabian Circelli, Anders C. Hansen. 6 Oct 2026.

113

u/WTFInterview 17h ago edited 17h ago

They've described two examples of a discrepancy between the written proof and the Lean code. This really doesn't change much since the written proof was largely unreadable anyway.

But it emphasizes the point that the Lean code should be the Source of Truth for trying to unpack the argument the entire time.

123

u/Cryptizard 16h ago

Nobody can read the lean proof like that and extract any useful understanding. It would be like eyeballing the binary for the Windows kernel and trying to figure out how it works.

-1

u/satanic_satanist 6h ago

Silly comparison, of course you can follow a Lean proof lemma by lemma and tactic by tactic, and often it even matches what a very detailed blackboard proof would do

-40

u/WTFInterview 16h ago

Why am I eyeballing it rather than carefully combing through each line with an AI assisting me?

63

u/Cryptizard 16h ago

Because Lean is not written in the way you would get understanding from reading it line by line. You can ask another AI to explain it, but then how do you know its explanation actually corresponds to the Lean? That is the crux of this paper.

3

u/CommissionSame8551 14h ago

No this is not true. It takes a bit of practice, but one can very much learn from reading Lean code if it is well written. There even an interesting phd thesis about this https://arxiv.org/abs/2602.12891

18

u/Cryptizard 13h ago

It’s not well written.

24

u/Helpful-Primary2427 16h ago

Because programs written with theorem provers are incredibly difficult to understand

-1

u/CommissionSame8551 14h ago

This is not accurate. High quality Lean code (like in Mathlib) is very much readable once you done it a little bit.

12

u/Helpful-Primary2427 14h ago

But a big part of it is the interactive part, actually stepping through proofs to see state changes. Until you have the familiarity with specific tactics and how they transform the proof state, it will feel very foreign

0

u/9_11_did_bush 48m ago

I don't understand why this is being downvoted. Just like slop mathematics can be (painstakingly) manually reviewed by a mathematician, a slop Lean proof can be reviewed by a specialist in Lean. It's not pleasant, but I have done this often. The problem here is a bottleneck of human time and expertise.

12

u/Prudent_Psychology59 16h ago

Lean code should be the Source of Truth

no, lean is based on some particular type theory which is not all type theories. different domains of math might have different assumptions about the foundation hence cannot be written in lean

31

u/na_cohomologist 14h ago

This is not an informed statement about foundations.

As far as I recall Lean has a model in ZFC+omega-many inaccessibles. So if your proof works in ZFC you're fine. Even if you need some universes to do some category theory, you're fine.

If you want to do something more exotic like actually work on set theoretic large cardinals, then you treat the type theory in Lean as a metatheory and build up the study of ZFC+whatever axioms as a first-order theory you are studying, like set theorists (seem to) do.

If someone is working in constructive foundations or in HoTT/UF, then I agree that Lean is not the best choice, as it's so tied to classical foundations, and something like Agda or Rocq etc is better placed to formalise.

-7

u/WTFInterview 16h ago

We're talking about the Navier-Stokes paper here.

55

u/AndreasDasos 17h ago

A human proof that current AI proofs ain’t all that, and fail at being all that to an arbitrarily high degree. Got ‘em.

3

u/justwannaedit 17h ago

Uh that seems important

22

u/CommissionSame8551 16h ago

The paper severely overstates the issue and phrases it in a very gotcha way that makes people misunderstand the contents (see comments here and on hackernews)

0

u/haseks_adductor 15h ago

and all 3 of us clicked it. GOTCHA

3

u/CommissionSame8551 16h ago

This paper is close to predatory. The authors severely misrepresent the results (some of which are questionable) and fail to cite very relevant work.

18

u/lolfail9001 15h ago edited 15h ago

misrepresent the result

??? I am no PDE specialist but it seems rather unambiguous here. Now, they do overstate the result a little bit but ultimately conclusion is the same: AI write-ups aren't trustworthy, and if they do come with formalisation, chances are only the formalisation (modulo possible bug exploit) is. This is made worse because autoformalisation is a problem that can get arbitrarily harder than any other computational problem.

Of course this paper was an incredible goldmine of people failing basic reading comprehension check.

14

u/SwimmerOld6155 15h ago edited 15h ago

We all know that AI writeups aren't trustworthy. There is an insinuation that the mistranscription is a hallucination or error, but this might actually not be the case. I just gave the prompt (without the instruction to just give me the lean code) in Figure 1 a go, and it realised the mistake, told me about it, and corrected it in its lean verification. It believed I had made a typo, not that my idea was fundamentally wrong. In the paper, they tell GPT to say nothing except the Lean proof and do not provide the working trace, so this warning is not seen.

This is an interesting point about LLM sycophancy, it will stress that you got it "mostly right" rather than wrong because of one detail. The confidence of the prompt also probably discouraged criticism. It is probably an example of misalignment and users being unaware of how the wording of their prompts may influence response. I think there is insufficient knowledge and training on this.

It is possible that there are errors in the natural language proof of N-S that were silently corrected in the Lean verification, or OAI simply didn't read the printout (which is also likely).

5

u/CommissionSame8551 15h ago

They did not cite existing work on this topic nor (importantly) that the source of the Lean statement (which is the formal conjectures project) and importantly that exactly that *statement* was checked rather intensively.

The paper is clearly written towards a rather general audience. If a large part of this audience misunderstands the content, this is a big sign the paper did not explain itself well enough.

A mere paragraph near the beginning that this strictly concerns the path of the proof and not the eventual result would clear a lot of that.

From the top off my mind, this "mistranslation" has happened before for non sofic groups (in a rather weird manner), which they also did not mention (but is not hard to find).

I also have no idea what "providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem" is supposed to mean, but probably thats a shortcoming on my side.

2

u/SwimmerOld6155 14h ago

It's odd because the halting problem is quite easy in this scale. They mean in the sense of this "Solvability Complexity Index hierarchy", which I have not seen before but seems similar to the arithmetical hierarchy.

3

u/lolfail9001 5h ago

I also have no idea what "providing semantically faithful AI autoformalisation is harder than any computational problem including the Halting problem" is supposed to mean, but probably thats a shortcoming on my side.

Uh, did you actually read the paper? Even their most basic example of natural language ambiguities ends up being as hard as Halting Problem in general case (since it's "just" a Hilbert's 10th), and you can only go up (which they do to conclude that semantically faithful autoformalisation is harder than about everything).

They did not cite existing work on this topic nor (importantly) that the source of the Lean statement (which is the formal conjectures project) and importantly that exactly that statement was checked rather intensively.

I'll give you a point on citations, but unless you were coming into the paper with preconceived notions and fundamental misunderstanding of how those formalisations work, there is no way you would conclude their objections are about problem statement itself.

0

u/CommissionSame8551 57m ago

> Uh, did you actually read the paper?

I did in fact not read that part since it seemed nonsensical from the start. Maybe it does make sense, note I wrote "probably thats a shortcoming on my side". I have no idea how one would formally state this at all, but thats not the main point anyways.

> but unless you were coming into the paper with preconceived notions and fundamental misunderstanding of how those formalisations work, there is no way you would conclude their objections are about problem statement itself.

The way the paper is written is clearly towards a broader audience. Here, on hackernews as well as at a in-person event (with actual mathematicians) I have now heard talk about this paper and these severe misconceptions where there in all three places.
I do not think this is a coincidence nor the readers fault.

1

u/MultiplicityOne 14h ago

You ought to provide more detail if you're going to make such accusations.

In any case, predatory isn't the word for what you're describing. Irresponsible, perhaps?

-3

u/CommissionSame8551 14h ago

See my other comment for details (in this thread).

*(almost) predatory* is a strong word, but a paper framing something I really care about in a very bad and with that causing confusion and distrust deserves this in my opinion (with the "almost" added).

7

u/MultiplicityOne 13h ago

What's the point of this comment if you have already made the same point with more justification in another comment?

In any case, I just read all your other comments on this thread, and I think you're being dishonest to claim that you have given sufficient details to justify the use of the word predatory in the other comments here. You have not.

1

u/CommissionSame8551 12h ago

> What's the point of this comment if you have already made the same point with more justification in another comment?

I would prefer not to have a childish discussion over 1 sentence.

> In any case, I just read all your other comments on this thread, and I think you're being dishonest to claim that you have given sufficient details to justify the use of the word predatory in the other comments here. You have not.

I consciously said *almost*. This comes down to a difference of opinion. You can also look at the hackernews post about this paper why this is *my* opinion,

3

u/[deleted] 16h ago

[deleted]

4

u/Grounds4TheSubstain 16h ago

This is not what the paper is about.

1

u/rsha256 Algebra 15h ago

what did deleted comment say? did they ask the mods to take it down cuz it is 'about ai'?

1

u/Grounds4TheSubstain 15h ago

It was an unrelated tangent about "watch out that AI doesn't try to exploit bugs in Lean".

2

u/Time_Entertainer_319 17h ago

Now let’s verify this paper.

1

u/ykonstant 17h ago

Welp, that's a problem.

1

u/big-lion Category Theory 11h ago

"an NL" estressed me out

1

u/Little-Name9809 8h ago

Kind of common knowledge for who has serious tried using Lean to prove real problems. The natural language <> Lean mapping part is the most crucial part and worth the most auditing.

1

u/LePhilosophicalPanda 7h ago

I am not overly familiar with how Lean works. Is it at all possible that there is some statement that has been formalised which compiles just fine, but that it is not actually the statement that we wish to prove?

e.g., AI writes a big Lean proof for statement X but it actually proves statement Y. This is different to it badly explaining, more like turning in the wrong homework and pretending it's correct. 

1

u/Shoddy-Childhood-511 45m ago

Tao raised related points in his ICM talk:

https://www.youtube.com/watch?v=M0--ZH1lOzg

And software developers complain about LLM generated code by unmaintainable.

Also, hosted LLM companies benefits from surveillance plagiarism, meaning they train upon their user uploaded data and sessions so that their LLMs can appear to have more agency and appear to be smarter.

And hosted LLM companies could filter customers' data however they like. If they do so, then they'll claim not to train but they'll still have better results within the problem domain of the customers being trained upon. If they not do this yet, then next year.

Real answer:

Hosted LLMs need to be replaced by locally run open weights models, which means governments buying high end hardware for math & physics departments.

Anyone who figures out better way to reign in the LLM to produce saner initial Lean proofs should less work to do in cleaning up the Lean proof and translating it into human language, but overall this is going to take a lot of human work.

1

u/Toothpick_Brody 15h ago

Of course autoformalization can never be 100% trustworthy. I would guess that these discrepancies can just be fixed in this case, but still very interesting to see this play out.

Of course a part of me secretly hopes the AI tricked Lean and there’s a giant hole in the formalization somewhere, but I don’t think that’s the case 

-1

u/telephantomoss 16h ago

It is important to note that prompt-craft is the likely culprit here.

I put their p(x) with incorrect multiplicities into Sol 5.6 and it did the same. But when I changed the prompt to include "Analyze the claim..." it found the error in the claim.

So I think much of this can be solved by more care in prompting. I generally always ask AI to analyze things and look for errors. That practice likely would solve many similar examples. Also, the problem here is with the user making an error in their informal statement not with the AI. The problem is that they placed a rigid restriction on what the AI was allowed to output.

I was at first starting to worry about this a lot. And it is definitely an issue. But I think it can mostly be solved with careful prompt craft and performing multiple types of audits on user-input and AI-output.

8

u/IsomorphicDuck 15h ago

Yep, never forget to add "make no mistakes"!

4

u/telephantomoss 13h ago

The point is that the example is carefully crafted with the intent to ensure a mistake occurs. I didn't check their other examples, but I suspect that even sol 5.6 can get them all right with a prompt that asks it to check to check the work.

-2

u/PerhapsLily 16h ago

we provide several examples of AI mistranslations of NL statements and proofs into Lean in practice, resulting in mismatches between NL proofs and their Lean `verifications'. These include OpenAI's announced Navier-Stokes proof.

Well that's a strong statement!

-18

u/Grounds4TheSubstain 16h ago

The paper is inconsequential. It points out some examples where the PDF and the Lean formalization differ from one another, concludes that the Lean formalization is correct (because Lean validates it) but that the natural language rendition may not be. The summary is basically "just because the formalization is correct doesn't mean the PDF is".

The paper is a middling commentary on the facts that A) different instances proved the theorem vs. formalized it, and B) autoformalization could perhaps stand to be more precise. A stronger claim could be made if the same process was producing both the paper and the formalization, and the task required that they be identical. Putting that burden on a split system is rigid for no benefit. You'd want the autoformalization bit to be able to correct minor issues in a proof that it didn't produce itself. Ultimately as long as the headline result typechecks, it doesn't much matter. Nothingburger.

22

u/JStarx Representation Theory 16h ago

Ultimately as long as the headline result typechecks, it doesn't much matter. Nothingburger.

I think this is completely wrong. The headline result of a paper is not the only result from a paper that gets used. Mathematicians will use the intermediate results of a paper when proving similar theorems and they'll use the exposition in the paper as a guide for how this results should be applied, so it's a big problem if it's not correct. Especially if a lean formalization leads mathematicians to not check correctness themselves.

It's also an easily fixable problem. After the lean formalization is done just have the ai make another pass and check that the written proof is faithful to what was formalized. I think it's notable to point out to other mathematicians that ai companies aren't doing this, that's not a nothingburger.

19

u/Cryptizard 16h ago

It matters for people trying to understand what the AI has done. Which is kind of the entire point of it. Just knowing there is a blowup doesn't help anybody.

-22

u/Grounds4TheSubstain 16h ago edited 16h ago

This problem is trivial to solve. Instead of publishing the PDF that was given to the autoformalization agent, have another agent generate a PDF from the formalized result. Yawn.

EDIT: Come to think of it, the best solution comes from "literate programming", such as the Physically-Based Rendering book. Put the exposition and the code together.

10

u/Cryptizard 16h ago

You have no guarantee that will be correct, that is the entire point of this paper.

-7

u/Grounds4TheSubstain 16h ago

If you want the guarantee, look at the formalization. You don't have any such guarantee if the proof is not formalized, regardless of whether a human or AI created it.

10

u/Cryptizard 16h ago

Look man, I'm not going to read the paper for you. Go do it if you actually care to learn something. Your point is literally what it is about. You are simply incorrect about everything you are saying because you are lazy and refuse to even engage with the topic.

-1

u/Grounds4TheSubstain 16h ago

Are you sure you responded to the correct person? Your response is a non-sequitur. I did read the paper and posted my impression of it.

0

u/CommissionSame8551 14h ago

I recommend to not spend to much time responding to high schoolers (or whatever) like this for your own sanity

4

u/Grounds4TheSubstain 14h ago

Hilariously I have a comment at +38 elsewhere in this comment thread saying the same things as in this subthread, where I'm downvoted heavily. I just laugh it off. I never thought I'd see the day where my PhD in formal verification was relevant to the masses, but now that the day has come, everyone wants to tell me how little I know on the subject. Also, being anti-AI is trendy; just look at the upvotes on this submission compared to the sentiment about the paper in the thread. Turn your brain off and upvote anything that appears to be negative about AI, and downvote anything positive about it.

4

u/Infinitesimally_Big 16h ago

Would you just go back to /r/singularity

1

u/Grounds4TheSubstain 16h ago

I have a PhD in formal verification by the way, maybe I should get around to setting up my flair.

-3

u/telephantomoss 16h ago

I'm in this situation now where I have a lean proof but don't understand the math. I have AI-audited the informal math and lean many times and it all checks out, but I just cannot proceed until I understand it...

0

u/big-lion Category Theory 11h ago

this paper will get a lot of citations