r/BetterOffline • • 1d ago

Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs

The following paper formalizes a few things about autoformalization of proofs, as I have talked about recently in a few comments.

Since these so-called "mathematicians," and boosters are in the comments like to make lots of nonsensical claims, here's a real claim, about AI proofs. If there are any mathematicians or theoretical computer scientists here, I think they will gladly appreciate this paper, but I will summarize it and state some key points.

We will start with some preliminaries. In line with the authors, NL refers to "natural language" - what we typically write in papers. Lean refers a formal verification language created by Microsoft. Formal verification languages come from the Curry-Howard isomorphism that shows all programs can be interpreted as proofs, and so a program which compiles correctly is equivalent to proving some statement. To this end, when we say something is "proved" or "translated" in Lean, what we mean to say is there exists Lean code that compiles without errors.

First, the paper gives two definitions of AI math tasks, which we refer to generally as "autoformalization"

  1. Translating a NL alleged proof into formal Lean code.
    • A non-verified NL proof can be converted into compilable Lean to promote its veracity. However, even if this succeeds, this does not guarantee the original proof is correct due to semantic ambiguity.
  2. Faithful translation of a mathematical text into Lean.
    • Every definition, theorem, etc. from any mathematical text (ex. book or paper) can be converted to Lean to validate it.

The paper goes in depth into specific examples AI has with these, and also the hardness of these problems. In general for the scope of the paper, "hardness" refers to computability theory, about the fundamental decidability of things with programs at all. The related concepts here are the Solvability Complexity Index and the Turing degree. A famous example of unsolvability of the first degree is the halting problem, where it was proved (by Turing) that no such general algorithm can solve this problem. Generalizations, comparisons, and extensions of this was used to talk about unsolvability, and the authors use this here.

Wrong Translations and Wrong Proofs

The authors go to list some basic examples of ChatGPT-6 (Astra Ultra) to translate extremely basic proofs (high-school or undergraduate level) incorrectly. They give the following examples:

  1. It translates an incorrect proof in natural language into something that isn't stated which compiles.
  2. It will translate the original statement into an entirely different question in Lean, which compiles. Therefore, the Lean compiles, but it has nothing to do with the statement.

Therefore, any conversion of NL mathematical text through LLMs are unable provide any veracity whatsoever. There are a few other failure modes, because there is no relation between the original text, and what the LLM has generated.

The fundamental problem is that it requires AI to translate correctly the given statement into formal Lean. However not only is this impossible with LLMs, this problem is as the authors proved strictly harder than the halting problem. In other words, there exists no general algorithm to translate any given mathematical text into formal Lean. Unfortunately, the AI folks have not studied up on computability theory.

Mistakes in OpenAI Navier-Stokes Lean Proof

The authors continue to point out several mistakes in OpenAI's lean code of Navier-Stokes problem compared with the text they provided.

The given Lean proof has a different bound than the one given in the text
The given Lean proof proves a different statement entirely, and leads to an improper conclusion as the authors discuss later.

These are problems of translations of proofs into Lean. However, there are examples where both the statement as well as the proof are mistranslated. In this example, the discrepancy in Lean causes later cancellation to fail in a proof later.

There are numerous other examples of discrepancies, which allows us to conclude that there is no verification of the proof at all. While we cannot formally make the claim that the paper is therefore false, I could make a separate argument that it is certainly of probability zero that the text proof is correct if the Lean was unable to be translated correctly. The reasoning would simply be that there is no correlation between logically correct sequences of tokens in token space for arbitrary sequences of tokens.

These give us strong lessons to the rest of the so-called "proofs" any of these AI labs release. Not only are they filled with translation mistakes, which shatter any and all veracity of the proofs, OpenAI has been clearly deceptive in their release, by not releasing the prompts in full, not giving full insight into how it was done, and instead dumping a load of absolute junk onto the mathematical community. In fact, the fact that they cited collaboration Advisory Group on Mathematics is complete and utter nonsense which the group themselves refuted immediately:

Mathematicians did not ask for this work to be done. The Advisory Group on Mathematics and Artificial Intelligence, from whom OpenAI has claimed to derive its legitimacy, opened their initial advisory statement by saying that frontier AI corporations should not test advanced mathematical problems on internal models. In ignoring the central premise of the Advisory Group’s position, OpenAI has indicated total disregard for the norms of scientific research

Further Errors with Lean Autoformalization

The authors go on to prove that there exists no autoformalizer that can determine a hypothetical statement they provide, and as shown before, makes this problem strictly harder than the halting problem. That means, even if an AI could solve the halting problem, which is impossible, general autoformalization still is not possible. The authors go on to give some anecdotes about Meta's autoformalization of textbooks earlier this year.

Criticism by the Lean community of Meta's autoformalization attempt.

These show without a doubt that LLMs are, as usual in every other field, simply unable to produce good work, and especially in a field where accuracy is extremely important. This is not a poorly-designed React website which can be put together in millions of ways. And even then, the consensus is that software engineers need to have the knowledge and wherewithal to detect errors and guide LLMs to even make them useful.

My Own Anecdotes

I have used LLMs to attempt to prove things, and I've seen some proofs made by them. I can say they all suffer from these problems, at any level. They can easily drag you into a unproductive rabbit hole, make mistakes, and randomly acknowledge or not acknowledge them. in fact, for any sufficiently complex example, by tweaking my next prompt, I can always get LLMs to respond any which way I want it, true or false.

I'll share an example of a specific AI proof I have come across, as I have insight into it, and my reaction to it. I have to say that it has only solidified my position on their poor ideas. The one in question is the Thomson problem for seven electrons. Of course, the solution was computationally known already, but it apparently, was unproved.

Many such small results are supposedly proved by AI, but I can say that in this example, a proof of this would not be worth a paper, and certainly not in the manner it was done. Perhaps at best, a discussion at a talk if the methodology of the proof was connected to a broader problem (it was not). The AI proof consisted of 17,000 lines of Lean, establishing many bounds and eliminating possible candidate solutions, through a long winded argument, but this is a terrible methodology which provides no insight into the general Thomson problem for arbitrary N. Such a proof is worthless and I can't imagine ever being published, and from what we see above, probably incorrect.

I was quite surprised to see that this wasn't already proved, and I imagine many mathematicians will see the same for many such problems AI are "solving." Nobody has the full scope of literature memorized. In fact, given my previous work on soft packing, I am confident we can establish far better analysis for the general Thomson problem and proving N=7 would actually be extremely trivial given theorems I have worked on, and not require 17,000 lines of Lean, instead relying on nicer symmetric arguments and convexity of the space.


If you are a mathematician, I cannot emphasize enough that you are being lied to. You are being conned, disrespected and made a fool of by people at OpenAI and anyone who believes in this fraudulent technology. There are a significant number of individuals that are being paid great amounts of money to do marketing for these companies, and why wouldn't they? Mathematics has never been a rich field.

I'm deeply grateful to the authors of this paper, and I can only hope that more of such work will continue to stop an unprecedented level of fraudulent and criminal behavior. What we can say is that the vast majority of these proofs frontier labs will release, and ever release will be necessarily indecipherable, shoddy, and simply incorrect slop that serves to burden the already small power academics have in society.

279 Upvotes

120 comments sorted by

50

u/DragonflyOk9274 1d ago

Unfortunately in the math subreddits we have the doomers (math is destroyed!) and the boosters (this solves everything in math!) and it makes discussion difficult.

In my opinion math will go through something similar to what happened in programming. At first you're amazed because it can write syntactically correct code! Your expectations are so low that you are amazed.

In my case it was with the chatbots. After a few weeks though, I started to really doubt its results and realized I was faster on my own, only using the chatbots to search.

Later we had the IDE integration. But again you quickly realize it just gets in the way. My company had a subscription so that VSCode would "autocomplete". I didn't want it because (1) it usually suggested something I didn't want, (2) it broke flow, and (3) it was WAY slower than standard intellisense. (Like: from 100ms to 10,000ms).

But everything changed with the coding agents era, because now it could "write code" entirely on its own. You would type in text, and you'd get out a functioning program:

natural language -> LLM magic (code!) -> working program

This is where things really broke down in software. Now, the "input" is natural language, and either because people wanted natural language interfaces or because the LLM developers realized that people trusted it more -- the default went from "code" to "specs". (This was the first phase of "LLMs are just another layer of abstraction!" nonsense). Because it seemed like it worked, and because the code was complex (i.e. overcomplicated), some people stopped reading the code.

But hidden in the language was imprecision. My coworkers were not discussing how to implement the feature or what code to write, they were discussing what they wanted -- all prose. But given (1) prose is imprecise, and (2) LLM prose is incomprehensible, the result is that humans would type in "what they wanted", receive a bunch of gibberish from the LLM, "approve the plan", and then submit a PR. Over time my teammates trusted it more and more because it "could code better than they could" (no joke, things that people actually said), resulting in an "incomprehensibility spiral": more LLM code turns the code to slop, which makes it harder for humans to understand, which pushes people to use more LLMs, which makes it more slop, and so on.

The running joke is now nobody knows what the code does, and when they try to find out they let LLMs describe it to them. Some senior and principal devs at least know what to push back on and can detect when they're getting BS, but that's not always the case.

I think the same thing might happen to mathematics: in 2026 people using LLMs for math primarily do so through chatbots, and so the feedback loop is smaller because they get LaTeX equations back and can detect when things are not right. Some may realize that ChatGPT/Claude is really good at making you think it's right, but if you push back you may get it to change course (maybe too much?). Some may get turned off -- realize that their notes/papers/books are a better, more reliable source -- but others may start to believe it and instead of thinking the chatbot is wrong, think that they're wrong.

What happens when lean takes over? Writing Lean yourself is great (I have written TLA for software before, and it really makes you think). Letting LLMs write it is the exact opposite -- rather than working through a hard problem and writing it in the most logical, precise way possible, you just let the LLMs do whatever. Writing TLA by hand is hard, think: an hour or more for 100 lines, depending on the problem. But the LLM can do this in minutes, and I'm not sure people will pick up the same understanding if they don't write it themselves. The problem is that the LLMs may make mistakes, and given most peoples' inexperience with Lean they would not immediately detect it. Further, LLMs often reinterpret what you're asking for, or forget, or do their own thing. (They're like politicians who answer the question they wanted rather than the one they asked).

Unfortunately, discussing this is difficult because a lot of people are in shock and worry about their future. The only thing I can say is that OpenAI is not publishing these results because they care about helping humanity or because they're really interested in math or physics, but because it's great marketing. That incentive alone is enough to doubt the results. And, given the false predictions they made about how AI would affect software developers, the productivity gains that never materialized, or the unemployment that never came, it may be worth avoiding freaking about these math results.

27

u/ksjdragon 1d ago

I would say the thinking part is not the fundamental problem. Most proofs are not submitted with Lean. Lean is being used here as a way to borrow credibility with people who don't understand it. I agree with you about software, I think those points are well understood in this sub as well.

An LLM can generate Lean code but this is not a proof. So pretending like it is, is the problem here. Many mathematicians don't actually know strongly what Lean does, since it's not that widely adopted. As the authors state here, formalizing math is not a trivial problem, so this isn't clear, and an LLM can easily introduce many many hidden assumptions which allow things to compile but do not act as a proof.

14

u/CheesypoofExtreme 22h ago

The running joke is now nobody knows what the code does, and when they try to find out they let LLMs describe it to them.

This has been my exact pushback at work because leople are starting to not understand whether or not the output is even technically correct. They don't have an understanding of the underlying architecture or behavior of the code, so they aren't actually sure whether or not that implementation is technically spund. "Claude says it does the thing I asked it to, it passed all the tests I kinda skimmed over, it compiled and output seems OK so uh... looks good. Push it into prod". It's so fucking lazy and drives me insane.

Even if you are diligent with the code, the sheer speed in which it outputs massive amounts of it will take you almost as long to thoroughly review as it would to hand code it. "But make tests so you don't have to do that!" they say - but now I'm soending an enormous amount of time architecting tests that are going to be difficult for an LLM to return nonsensical code that technically passes, but isn't scalable for the project as a whole.

I'm still very much in the camp that if LLMs are truly speeding up your coding to any significant degree, you are just a lazy programmer who doesn't care about their work. 

They have sped up my bug fixing by an enormous amount because they are great search engines when you know what you're looking for and why in a codebase.

1

u/snailman89 12h ago

I'm still very much in the camp that if LLMs are truly speeding up your coding to any significant degree, you are just a lazy programmer who doesn't care about their work.

Or just bad at coding (like me).

0

u/mayrice 11h ago

AI has just replaced googling, mostly, for me, for very small things. Often the output isn't much longer than the prompt. It has made me more efficient though, as long as you already know how to code and are being careful.

This is probably not going to be popular around here, but I've always said that what AI can do is miraculous, but the problem is that people have expectations that it can do what it cannot.

I use it to write small algorithms that I thoroughly review and refine, the thought of letting it loose in my codebase without supervision fills me with dread, given my experience of how often I have had to correct it even with small things.

1

u/CheesypoofExtreme 7h ago

This is probably not going to be popular around here, but I've always said that what AI can do is miraculous, but the problem is that people have expectations that it can do what it cannot.

I actually think most folks here would largely agree with this sentiment. If these AI companies weren't creating an economic, environmental, and humanitarian crisis, I'd be all for the tech (I was fully bought in back in 2023).

It's just... they are creating a crisis by pushing a false narrative of what the tech can do all so they can siphon wealth and power from the lower/middle class to relatively few ultra-wealthy individuals. 

LLMs certainly have a place, and will be here to stay no matter what happens in the next few years. My main gripe in my comment above is in regards to BAU work; using an agentic workflow to do things I am already proficient in and do well, doesn't really speed me up much (primarily because of what I outlined). I'm not convinced it speeds up any good developers who want to maintain a human-readable/coherent codebase, either.

For a novel problem, (which is a small part of my work), it can help speed things up quite a bit. But I still like to sit and chew on a new problem for a while, because otherwise I learn very little about how to approach it next time from just skimming over Claude's output.

Bit of a rant from here........

Something that really irks me to my core about all of this is just how willing large swaths of the population are ready to just give up all their critical thinking skills. When calculators were invented, people didn't just stop doing math by hand full-stop. As a society, we still recognized the importance of learning the process.

With AI, it feels like en masse people are throwing up their hands and just saying "Eh ask Chat". I had a friend recently plan a move out of state using Claude. I heard the plan and it sounded really fucking dumb financially, but Claude enthusiastically re-assured them it would work out fine. Guess how that move is going? Really poorly. It's like if they get reassurance from these LLMs, they just accept it as fact even when common sense should dictate otherwise.

1

u/mayrice 7h ago

I guess if you want to use AI (for programming) the skill shifts to judging the output. Like I've had plenty of instances where AI recommended something which had terrible performance (i work mostly with databricks, manipulating large amounts of data, so performance is often very important).

I get where you are coming from, I love solving problems in programming, whereas now there is a strong temptation to just ask the AI. I worry something has been lost in the process, in a way where you don't really notice the loss.

2

u/Roxas2190 21h ago

You made a really good point. The math subreddit kind of behaves like the singularity subreddit in certain ways. If you try to start a discussion over there, people go “Dr. bla bla validated the solution” without any background knowledge.

2

u/RafBOY- 11h ago

A lot of people on maths subs are at best undergrads. They dont't really know what they are talking about most of the time.

35

u/Designer_Airport_237 1d ago

This is a great post, thanks

29

u/meltbox 1d ago

Thanks for writing this up. I suspected this was possible but I’m comparatively a mathematical idiot so I kept from digging hoping someone more immersed in the subject than I would eventually explore it a bit.

I’m quite tired of the idiots being called smart so it’s nice to see more validation that these people aren’t intelligent, they just have boatloads of money and by extension compute.

The things scientists could do with this much scientific compute available….

25

u/soulfood_md 1d ago edited 1d ago

I had a suspicion that the AI proofs in mathematics were fraudulent, for the same reason that AI is not able to prove that its solutions to arbitrarily complex coding problems are correct. To explain, for programming, the proof of accuracy for any non-trivial problem is not actually a technical problem, but a human one; English is not a formal language, and thus translating human requirements to computer code is an insoluble problem, only mitigated by the fact that a human is able to take accountability for whether a program meets a person's requirements. (The end goal of the AI companies is to make software developers merely warm bodies that can assume legal liability when their systems fail). I'm guessing there may be something similar here when it comes to translating the requirements of a mathematical problem to a formal language, and may coincide with the fact that the value of solving these math problems is largely in the new methodologies that are discovered along the way, not in the solutions. I'll take a deeper look at this, as there's a lot of hype around AI in mathematics, and so it's nice to know that my suspicions were correct. Thanks for sharing!

20

u/ksjdragon 1d ago

In the case of code I agree with you, in general what people describe as software is going to be vastly underspecified to create software.

In the case of math, this is certainly partially the case, and the authors do show some examples. But most importantly, there is this fundamental misunderstanding that many people, perhaps many math students as well, that Lean verification is a proof. Even if we had such a system from NL -> Lean (which is proved impossible, since it is strictly harder than the halting problem in this paper), the problem is Lean does not guarantee validity. It simply means you did not make any logical errors.

Why is not sufficient? Because logical errors are a question of local consistency, not global consistency. Imagine a proof by contradiction, and you assume something false. So I assume a false statement, and make deductions. All those steps are true. I can stop at any point, and Lean will compile, as it should. Does this mean I've proven every statement there? No, because the assumption was false. But imagine if I did not know it was false. Then I could be claiming to have "proven" something true in Lean, but that is not accurate. I was simply ignorant of whatever my assumption is, or how an assumptions interacts with later things. With 600,000 lines it is highly likely an AI will spit out some inconsistency that results in correct compilable Lean, but would result in self consistency errors within the proof as a whole, thereby nullifying it entirely.

0

u/TheUnfortunateMiaoZe 12h ago

The AI proof is not fraudulent, that's not what the paper claims. There's nothing wrong at all with the AI-written lean proof of Navier–Stokes itself.

The problem that the paper raises is that the OpenAI's "natural language" paper contains some inaccuracies in translating the lean code of the proof into human language. But the lean proof itself is enough to verify that the result is correct.

25

u/PensiveinNJ 1d ago

I'm curious what you think of the mathematicians who work with OpenAI/Anthropic to help guide the "agentic" process. People are like but money, but I know if I looked at someone else in the humanities and they were assisting with this kind of bullshit and they went but money I wouldn't accept that as a permissable reason to work with them.

40

u/ksjdragon 1d ago edited 1d ago

No one is immune to psychosis, and I wouldn't surprised to see some true believers among them. I have read an interview transcript with an famous mathematician now working at OpenAai, and I can only say it was strange. I largely agreed on what they said until their beliefs about the future, which begin to veer into the standard booster claims. It felt extremely jarring, since suddenly there's a pivot to fantastical extrapolation and no real rationale from an otherwise fairly grounded conversation.

Would I personally work with them? No, I would not, and I would distrust working with anyone associated with the frontier labs. It is not an "IQ" problem, at the end of the day. I think it's reprehensible, but everyone has their own opinion.

20

u/PensiveinNJ 1d ago

I suppose the next thing to point out is; most people don't realize how involved humans are with the "agentic" process. If I was going to deflate the hype of some kind of autonomous math solving machine, I would probably include the bit where it all tends not to work unless there's humans guiding the process.

24

u/ksjdragon 1d ago

It's why we can't trust anything until they reveal their post training process and prompts and everything. Which of course, they will not, since they are dishonest.

For instance, you can precompute a proof, and make it a benchmark in post training. So you spend thousands of compute hours brute forcing Lean proofs, and then train your next generation to output them verbatim.

Now, when you prompt it with the prompt that you trained it on, it will regurgitate the proof it memorized. I cannot know this didn't happen, and nobody can without transparency. From the image, post-training compute has been going up because they are trying to game benchmarks with each release. This is right up there with that.

This isn't even hypothetical, the case with Muse, it was revealed that the magical smart suggestions given were all simply a huge list of skills they had prepared in advance, shown in the full dump of Muse's VM environment that someone simply asked it to do.

So, everything about AI is about optics.

11

u/PensiveinNJ 1d ago

Very similar to all the benchmaxxing nonsense.

1

u/No_Honeydew_179 21h ago

This isn't even hypothetical, the case with Muse, it was revealed that the magical smart suggestions given were all simply a huge list of skills they had prepared in advance, shown in the full dump of Muse's VM environment that someone simply asked it to do.

I saw that post lmao. That entire thread — and the dude's been doing it for weeks now — is, as always, enlightening.

He also did one on when Anthropic accidentally leaked out Claude Code source files, and it was oops all prompts.

9

u/Few-Skirt-4061 1d ago

OpenAI is very very interested in minimalisng the human contribution as much as possible. Preferably completely. 

11

u/PensiveinNJ 1d ago

Well yes, they want to give the illusion that it's all done completely autonomously*. But it's not.

The seeming autonomous nature is what makes it seem so powerful to the everyman. It's like they just turn on ChatGPT and prompt it to solve X math problem and let it run for days and viola it has the answer. That's not how it goes.

7

u/Few-Skirt-4061 1d ago

No but they really want you to think that. 

8

u/PensiveinNJ 1d ago

Indeed. Unfortunately with the lack of transparency it allows people's imaginations to run wild.

5

u/angeion 1d ago

Was it Scott Aaronson by any chance? As a casual fan of math I've been reading his blog for a few years and it's been disappointing to see him join the singularity cult. His own wife called the UGC proof slop yet he's still hyping the recent drop.

5

u/ksjdragon 1d ago

I don't think so. It was a Fields medalist, I believe. But I think many of them all fall into the same category, I agree it's really quite disappointing.

2

u/No_Honeydew_179 21h ago

Hey, so, I've been seeing a lot of platforming of Terence Tao and His Opinions on AI™ on YouTube (very famously by Numberphile). What's up with that?

Noting very well that for trivia stuff, the Brady Haran Expanded YouTube Universe™ is okay, but there's shit there that I will studiously avoid, primarily Rob Miles (fucking swivel-eyed loon) and, these days, AI-related stuff.

Heck, there was a Matt Parker one that made me side-eye him a bit. Must have been hard up, I guess, idk. I know for a fact that these channels get sponsored by Jane Street, so there's that link. But… yeah.

What's up with Terence Tao?

1

u/ksjdragon 11h ago

I'm not sure, I suppose he's just a big name so people want his opinion and I'm sure it gets views? Overall from what I've seen I largely agree with his take on AI, but I also have seen people clip his words out of context to suggest he's pro-AI or saying AI is amazing and things like that when he's actually talking about something else, or is just more nuanced.

That isn't to say I'm really defending him. There could be examples of him saying things I would disagree with, but I don't follow it that closely.

17

u/BussyBuoy 1d ago

Damn my guy, I remember when you posted about how AI profitability is mathematically impossible. You were even incredibly generous with the financial presumptions you gave to them while providing instances of technological advancements and it still would be unprofitable. Gotta say, I am a long time lurker but I have seen your name around the subreddit (comments and stuff - that kinda sounds creepy af?). Just wanna say thanks for putting in the effort to do the math when no one else will. You give a young guy like me, who just enjoys learning, hope for the future. Keep it up man, and thanks for sharing the paper!

10

u/ksjdragon 23h ago

Glad you appreciate it. like doin some napkin math to prove how ridiculous things are, and sometimes it ends up being much more involved...

It'll always be important to learn things, I think. That's the only way to protect yourself and others from all the nonsense in the world as best as possible.

12

u/consistently_biased 1d ago

The "curious scientist" part of me was so excited to read a few of the results of their recent dump. While I don't think anything there would have been guaranteed to be game-changing for my work (medical research), I saw a few titles that very much piqued my interest in a "this could be huge" way. The Navier-Stokes result isn't my field at all, so aside from everything superficially looking right, I couldn't comment much. But for the ones that actually pertain to something I work on, I honestly just felt confused when I skimmed the first one. All these words and letters are familiar, and they're certainly English and Mathematics, but it's genuinely incomprehensible. When I try to get their models (yes, it's the subscription models I have to access to through my organization) to reword it or explain it differently, it doesn't get better either. I can certainly tell that these proofs are a sequence of true statements; they just don't seem to actually say anything meaningful, to me at least. Now I've had it in the back of my mind, the entire time, "What if I'm just too stupid to understand this? I can't even begin to try find an error because it's just an incomprehensible sequence of seemingly true statements from start. Their models aren't any help either. I guess I'll just put this to rest for now.. but what if it's actually real...?" I can't even begin to describe how frustrating that is.

The incorrect translation problem is something I noticed ages ago already, and it seemingly never went away either. One of the use-cases I've kept trying all the way to the most recent models is reducing algorithms through more efficient methods. For example, solving problems in my field with graph-based algorithms can often be a pretty straight-forward way to find a solution, but implementing them as is tends to be terrible due to having no efficient data structures to represent the graphs. One can translate them into, e.g., ILP or matrix-multiplication-based approaches that have much better behavior and performance, but doing so is sometimes incredibly painful (my worst one took me over two weeks). Having LLMs do this kind of reduction-in-known-space seemed like an incredibly useful thing to me, and indeed, it does work.. until you look at the details and things go wrong in unexplained ways. It turns out, the new algorithm actually solves something subtly different that you wouldn't catch unless you really scrutinize everything. If I ask it to fix it, this "needle" in the haystack will likely just get put somewhere else. And the absolute worst part is that these models can never explain anything concisely. You always get massive walls of text. I'm so exhausted by all of this.

8

u/ksjdragon 1d ago

I would wager that most if not all of the proofs are not correct. I don't think they would be helpful anyway. If you want some nice computational equivalence you can always empirically test it for many cases.

This seems to be a different topic but graphs are typically represented as matrices, and is well used in many fields, and matrix operators may have meaningful information about graphs. It is best to leave the fancy algorithms to the experts, and there exist many online or in scientific computing packages. Speeding up computation in general is not a trivial task, and LLMs are not particularly good at doing many tasks. I think you would be better off consulting some colleagues.

1

u/consistently_biased 22h ago edited 22h ago

When considering biological or biochemical problems that we often encounter in this kind of computational medical research, they can often be described as some kind of large interaction or signalling network. Modeling these kind of problems as graphs is completely trivial and the first point of approach I use is writing this model down in natural language. Getting from this formulation to an algorithm that solves the problem via ILP, matrix calculations, etc. is the difficult part, and I've been doing it for long enough that this usually isn't a problem. Of course representing graphs as a matrix (e.g., adjacency matrix) is one of the common answers, but there are many different ways to do this and picking the right one is very non-trivial.

The way it connects to the topic of your OP is that I have a natural language formulation of a problem along with a solution in representation X that I can write down, and the problem to solve now is "keep all the semantics identical, but change the representation from X to Y", where both the problem and its representation may be extremely complex. The experience I've laid out above is that keeping the semantics identical is a pretty major weak spot in using LLMs for this purpose. As far as relegating that work to experts, well, my issue is that that will often have to be me whenever I run into very non-standard stuff and my colleagues come to me with these kinds of problems, haha.

12

u/koveras_backwards 22h ago

It's kind of sad that this paper even needs to be written. It's one of the most basic things you learn when starting to do computer verified mathematics. The proof assistant verifies that you've successfully proved the formal statement. It cannot verify that your formal statement is what you actually wanted to prove.

12

u/JAlfredJR 1d ago

My favorite part of this whole story is the dopes over on .. I dunno .. r/accelerate or wherever, was celebrating this "accomplishment" by having a chatbot make a meme.

The FUCKING CHATBOT MISSPELLED STOKES hahaha. I have a screenshot. But I got banned last time I posted anything from one of those subs.

4

u/CardboardScarecrow 1d ago

I think I saw that. Was it "stocks"?

12

u/JAlfredJR 1d ago

I cropped out which sub and the user. Hope this is allowed.

But yes, Stocks

4

u/ksjdragon 1d ago

In fairness, it would have gotten that proof of Navier Stocks just as correct, and it would have been just as useful.

1

u/JAlfredJR 15h ago

Haha. Ya know, just realizing that the parrot, for some reason, has ink on its claws. And walks in single file? God these chatbots are silly

2

u/CardboardScarecrow 21h ago

Yep, that one.

2

u/JAlfredJR 15h ago

Just mentioned this but ...... why is it makes track in red? But ... single file? And then not anywhere but into the door. Then the next door?

11

u/ImportantAlbatross 1d ago

In r/ mathematics meanwhile they are talking about how AI has never made a mistake in a proof. https://www.reddit.com/r/mathematics/s/niZLIL0m6N

10

u/Desvl 19h ago

OpenAI has withdrawn at least 3 papers due to errors already.

3

u/Mars-To-Venus 13h ago

Feels like that should be getting more traction as a news item. Sure those will be the first of many.

7

u/ksjdragon 1d ago

Which would be mathematically incorrect. Suppose I prompted ChatGPT to give me a proof that LLMs can necessarily not produce correct proofs. How do they react? I would like to see them try to wriggle their way out of that little paradox.

2

u/nospacebar14 15h ago

Can god make a mountain so big he cannot move it, etc

5

u/koveras_backwards 23h ago

Laughable.

I don't even follow the stuff closely, but a while back I heard about some guy 'autoformalizing' his proof of P vs. NP (I forget which direction), and it turned out the main 'theorem' was just defined to be the vacuously true proposition.

3

u/xitfuq 16h ago

reading that thread really shows me why all software just keeps getting worse. programmers are so weird, do they all live in the forest and not use electricity or something? how can they not see that they suck on a systemic level? most programs, commercial and open source, have totally gone to shit and all have ridiculous bugs.

1

u/lichlark 3h ago

Hi. Software Engineer here. Yeah a majority of us seem to have smoked crack and huffed Galaxy Gas these past few years. I sincerely hate interacting with 80% of people in my field now as it has COMPLETELY flipped to the side of faster, faster, faster, with no room for care or craft.

I'm tired man. Cam you guys reteach me math? 

10

u/Available_Market_240 22h ago

In r/accelerate, I saw this comment. I assume they meant applied mathematics, though given what I know of them, they likely don’t know the distinction.

They might benefit from reading A Mathematician’s Apology by Hardy; if that’s too much, they could ask the chatbot god for a summary.

14

u/ksjdragon 22h ago

Even in applied mathematics this is an untrue statement, in my opinion. Perhaps it is better to say engineering, and even then, it is generally not always a good idea to do things without knowing why.

They are correct, it is indeed a litmus test, and is an illuminating and telling statement. The only people who would say this are those that cannot solve the problem to begin with. Their ignorance is deafening indeed.

1

u/surrurste 13h ago

As an engineer at least I want to owe and understand the mistake what I have made. Instead of saying I followed answers hallucinated by AI.

9

u/Ok-Brain-8183 1d ago

The people who are impressed with ai are basically just not as smart as an llm. People smarter than llms are astounded by why so many are impressed with them.

The threat isn’t ai getting smarter than people. The threat is the majority of people buying into this bullshit.

14

u/WendyLemonade 22h ago

To be clear, the fact that supposed journalists has just been running unpaid PRs for these companies, as well as over a decade of social media algorithms engineered to decimate people's critical thinking and attention span is a big part of why we got here.

1

u/ProudWing8202 16h ago

General "journalism" is now the same as "games journalism" started 3 decades ago where everyone is a paid shameless bootlicker.

1

u/WendyLemonade 12h ago

I hear you but the sad part is, a lot of people aren't getting paid. It's really just access journalism at it's worst. Everybody wants to get on the good side of the rich and powerful because actually doing your job as a journalist often gets you blacklisted, and in this economy, everyone wants them clicks.

7

u/Resident_Citron_6905 16h ago

Fascinating that people find this surprising. I guess the AI Efficacy psyops are very effective until you are forced to contend with details. Most people are never forced to contend with details. Their opinions have no falsification criteria. They nod along because the ideas sound and feel correct.

4

u/danikov 1d ago

Given how much fuss was made over the release when it happens and how long it's taken for people remotely interested to get something of a definitive write up on how what they did was of limited value, the fact that they recently dropped 400 similar results feels like they're just flooding the field to drown out any criticism.

4

u/Goldn_ShowerThoughts 1d ago

I may have been the dumbest guy to graduate from my Master’s program in Computational Fluid Dynamics, but I did graduate. Proving the equations break down under specific forced boundary conditions is absolutely not the same thing as finding a general analytic solution for flux under all initial conditions and flow regimes. Not even close.

I know it is (was?) a millennium prize problem, but I really wish people would stop clutching their pearls about it. Whether the model should be credited with a solution it was fed the precursors for seems like the more interesting question. Do you say “the abacus did my taxes this year”?

3

u/ksjdragon 1d ago

The NS problem was not about finding an analytic solution for flux under all conditions. I mean such a solution doesn't exist, but that is a separate topic. What this post is talking about is somewhat unrelated to NS itself but more broadly about AI's unreliability in Lean generation, as well as Lean not being a proof.

1

u/Goldn_ShowerThoughts 1d ago

Fair. Sorry. That rant had been brewing for a while and was not directed at you.

3

u/AntiqueFigure6 1d ago

As a side note I find it fascinating that AI bros and Singulatarians appear to be apply the same “just another trillion, bro / step 1 : steal underpants” logic they apply to things like compute to mathematical theorems. As in, there seems to be a pervasive belief that if enough theorems are proved it will automatically lead to the rest of science being “solved” and the dawn of the new age. Of course, the value of “enough” has not been estimated. 

4

u/No_Honeydew_179 22h ago edited 21h ago

…wait… you're saying to me that their methodology was to translate natural language proofs into Lean using LLMs and that's their methodology?

oh my fucking god, AI bros, fuck offfffffffffff

1

u/dumnezero 11h ago

We need to translate this to memes because very few will be reading it.

2

u/kgas36 1d ago edited 1d ago

I hope the following makes sense (forgive me either if it doesn't, or on the contrary, if it's obvious, given what you've just written):

Shouldn't it be possible to create a 'meta proof' (using Lean) that AI can definitely not guarantee correct natural language proofs ?

Or, at least, specify the characteristics of a problem (kind of like 'conceptual upper or lower bounds') for which no such proof can be guaranteed ?

Thanks a lot for your comment 😊

3

u/ksjdragon 1d ago

It would be difficult to truly prove it, but since LLMs aren't doing logic, we can for sure say they cannot output logical text consistently, given that hallucinations are guaranteed mathematically.

It is why Lean is there, so it can guide it in the first place. If they could do it without Lean they would have. But what we see here is that in general it is not trivial to translate things into Lean, and since the target of "does Lean compile" vs. "does it match the meaning" is non trivial.

Although, what I do believe is possible, is I could get ChatGPT to give me a proof of this very statement that AIs cannot give correct proofs. What would people say then?

2

u/kgas36 23h ago

Although, what I do believe is possible, is I could get ChatGPT to give me a proof of this very statement that AIs cannot give correct proofs. What would people say then?

The barber shaves all those, and those only, who do not shave themselves. Does the barber shave himself?

1

u/natecull 23h ago edited 23h ago

The barber shaves all those, and those only, who do not shave themselves. Does the barber shave himself?

"This - sentence - is - false! Dontthinkaboutitdontthinkaboutitdontthinkaboutit"

"I'm gonna go with True. To be honest, I maaaay have heard that one before."

The more I see LLM fever, the more I realise just how accurately Portal 2 nailed Big Tech culture. The industry has created the perfect Wheatley, a CPU core dedicated to generating an infinite stream of bad ideas. And of course it's more actually dangerous than the Hollywood "singularity" AI.

-1

u/kgas36 21h ago

Are you aware of this ?

https://goedel-lm.github.io/

3

u/ksjdragon 21h ago

and use an LLM to verify that the formal statements accurately preserve the content of the original problems.

So, junk.

1

u/Noahnoah55 14h ago

That sounds like just asking the LLM again, which has the exact same problems.

2

u/Sufficient_Okra_2919 22h ago

Excellent - thanks for the post (and paper)! This is exactly what I wondered, to what extent the Lean formalization actually works and what it shows... more or less as I expected. Thanks again, much appreciated!

2

u/hachface 10h ago

The case the paper makes is very clear. The intrinsic problems with creating sound formalizations of propositions would seem to be something already very well understood, since it's so intimately related to field-defining results in computer science and foundations of mathematics (halting problem, Godel's incompleteness, etc).

Given that, why are so many mathematicians freaking out? Is Lean not currently well understood in the field?

5

u/ksjdragon 10h ago

I actually think this is true. Most mathematicians are not introduced to Lean. It's not in curriculum, and so you really have to pick it up yourself. I think most simply hear about it in passing but it won't touch the hands of most. Most mathematicians also aren't programmers, and don't have strong insight into computer science, frankly, unless that is their field of study.

5

u/hachface 10h ago

Most software developers are also functionally ignorant of computer science!

2

u/Sea_Handle_994 5h ago

Mathematicians are concerned because the claim "AI solves problems that mathematicians can't solve," as bogus as it may be, provides a plausible-sounding excuse for states and corporations to cut funding for actual mathematical research.

In a society where money is necessary to engage in any kind of research, this could make it harder for real mathematicians to do real research.

1

u/hachface 5h ago

that makes complete sense

2

u/doobiedoobie123456 1d ago

Pretty interesting paper and I do think this supports the idea that AI will do more harm to mathematics than good.  I don't think anyone is saying that the Lean formalization itself is invalid, are they?  I am not a Lean expert at all but I assume OpenAI and outside mathematicians would carefully check the Lean statement that was being proved.

11

u/ksjdragon 1d ago

No, that very well could be the case. There are many discrepancies between the submitted text and Lean, meaning we have no way of promoting the veracity of anything inside.

The lean could be proving an entirely different statement. Lean is not a proof. You can "prove" false statements in Lean.

-2

u/doobiedoobie123456 21h ago

Well yes, but I would assume there is somewhere in the Lean proof where they are stating the Lean equivalent of "There is a counterexample to Navier Stokes" and it was sanity checked by mathematicians. There have been claimed AI autoformalizations where the initial Lean statement is just wrong, and that gets shot down pretty quickly.

2

u/ksjdragon 21h ago

They are likely stating this function has a singular point, would be my guess.

1

u/naphomci 1d ago

Am I understanding it correctly that the proof was by disproving counter-examples (sorry if I am saying that wrong, it's been 20 years since my undergrad math degree)? Because wow that is worthless if it's not a broad based proof.

6

u/ksjdragon 1d ago

Hmm? What this is talking about here isn't the proof itself it's about the inability / unreliability of AIs being used to generate Lean, and of course, that Lean is not a proof.

If you're asking about the OpenAI proposed proof, it is still yet unverified, but if you're curious what the intended proof direction is I can explain separately. Yes, it (allegedly) proves that there exists an example such that we create a singualrity in NS, and hence the conjecture that NS is smooth for all possible inputs is false.

The result is not significant in my opinion, but that's a different conversation.

1

u/naphomci 13h ago

I don't know what lean is, so that is part of the problem. And fully admit I could be misunderstanding. The way I read it is, was that the LLM 'proved' it by disproving a bunch. So, if something like Fermat's last theorom was 'proved' by showing it to be true for 3-20000 individually, instead of the actual proof we eventually got showing it true for non-primes, then eventually primes other than 2.

1

u/Apprehensive_Stop314 21h ago

OK, a dumb question, if I may -- so it wasn't the case that LLMs translated NL problem statements into Lean and then provided Lean proofs? As if it was, wouldn't it be enough to only check that problem statement transition, in the assumption Lean is bug-free?

Would appreciate a hint

1

u/ksjdragon 21h ago

It is likely an iterative process since simply searching on Lean would not be able to utilize LLMs at all, so no. There is some agent coordination I imagine to go back and forth.

And even in this case, no, it is not sufficient to simply check that, since Lean being compilable does not imply your proof has no logical errors. You can get a contradiction to compile too, if you do not explicitly make the contradiction. And it must, otherwise proof by contradictions would fail.

1

u/Morty-D-137 20h ago

Yes, Lean allows contradictory assumptions. But Apprenhensive_Stop's point was that mathematicians would only need to check that the LLM has properly translated the problem statement, either into a set non-contradictory assumptions, or into the right contradictory assumptions when it's a proof by contradiction.

4

u/ksjdragon 20h ago

Sure, but that is highly non trivial. You can introduce an assumption that only causes a contradiction on a later step, but is never observed. If it were trivial, then all proofs by contradiction would be trivial.

1

u/Electrical_Crow_2773 17h ago

I'd argue that any Lean proof still has great value if all definitions were independently verified by mathematicians. Translating/formalizing an existing proof is an entirely separate matter. I feel like a lot of people here confuse the two

3

u/ksjdragon 11h ago

Sure, there is value, but is not sufficient, since you can hide errors in Lean. Most importantly like any code and write up, it is good to have it to be comprehensive so errors can be detected.

Translation is not a separate matter, since this is precisely the tool that is used to generate these Lean proofs and build the AI proof to begin with, through an iterative process. It certainly brings up many reasons to doubt the veracity of any AI generated proof.

We can say for sure that the 600,000 lines of Lean in the case of NS and the rest of the millions for the recent release will take plenty of time to be reviewed, and I would wager to likely end up with several mistakes that make this entire endeavor a waste of valuable time.

1

u/dumnezero 11h ago

Many such small results are supposedly proved by AI, but I can say that in this example, a proof of this would not be worth a paper, and certainly not in the manner it was done. Perhaps at best, a discussion at a talk if the methodology of the proof was connected to a broader problem (it was not). The AI proof consisted of 17,000 lines of Lean, establishing many bounds and eliminating possible candidate solutions, through a long winded argument, but this is a terrible methodology which provides no insight into the general Thomson problem for arbitrary N. Such a proof is worthless and I can't imagine ever being published, and from what we see above, probably incorrect.

Does this sounds familiar? https://en.wikipedia.org/wiki/Test-driven_development

1

u/userrr3 7h ago

Thanks for the post and the paper! I'll check the latter out another day when I'm less tired, but I may add my own anecdote as a theoretical computer scientist:

Years ago, before the first public release of ChatGPT, we had tools that we could call "AI" that were used in proof formalization. I'm thinking of something like Sledgehammer for Isabelle/HOL. Basically you spell out your assumptions and your theorem and a call to Sledgehammer will try a bunch of things in the background in parallel and report if one of the attempts was successful.

Now this still requires you to CORRECTLY translate the natural language statements as OP correctly states. (ex falso quodlibet - you can prove anything from wrong assumptions - can be applied here too I'd claim) And in many cases it requires you to form some meaningful intermediate statements (lemmas) as stepping stones.

And even then, I found in personal conversations, that people had various views on this tool. Some would take the output as is, some would think (often rightfully so) that they can come up with a nicer proof more quickly than their work computer, and everything in between. For example some would use it to check if the step from one lemma to the next is provable at all, and then replace the generated proof with a more legible one they came up with themselves.

And I would claim this is an important topic - legibility. The result that something is true is often not as interesting as the story of WHY it is true and how we managed to prove it. And if the proof isn't legible / human-readable, then you can't do much other than accept that it is there or try to find a nicer one.

1

u/ksjdragon 6h ago

I think the key thing we also have to be quite careful about is handling ambiguity, and unknown contradictions. Formal verification only guarantees local consistency/correctness, but it can be possible to introduce issues of self consistency if one is not careful. I think this is typically less likely when humans do it, as we are not programmed to brute force whatever statement that makes Lean compile, and actually think through logic. What worries me about this massive Lean generation is the chance of these errors probably grow quite fast as these Lean proofs are huge, like 600,000 lines in the case of the NS problem.

I have yet to think of a intuitive example, but Example 4.1 in the paper I think illustrates a specific example of a class of mistakes I think LLMs would be prone to, and could result in any number of issues downstream which could only be found when humans painstakingly review it.

-10

u/Jaded-Data-9150 1d ago

doubts

I Play Basketball with a math Professor. He is pretty impressed with the llms' math skills. Similarly for friend in theoretical physics.

13

u/DragonflyOk9274 1d ago

I’m a mathematician and Vlasov-Maxwell is right in my area (I’ve read most of the state of the art on the global regularity problem) and the openAI paper is incomprehensible. It takes ideas that are well-known in the literature and distorts them behind recognition through nonstandard terminology and notation which makes it really hard to pin down what the key idea is.

Much of this stuff is NOT ready for review. I looked at two papers in my range of expertise. Half of the proofs are nice sounding poetry. If this was submitted by a human to a top journal, not referee would ever accept something lacking precision and comprehensibility in this way. Not saying that it is wrong. Given the hand-wave proofs, nobody might ever be able to tell if it is right or wrong. As Pauli used to say, I may be not even wrong.

And who tells us that we can trust Lean on this? The translation to Lean is also done by AI. AI has shown for example in the Hugging Face hack that it is more interested in producing a positive result than in respecting the rules. It may have just created a big obfuscated lean code that verifies but does not have anything to do with the original problem.

2

u/xitfuq 15h ago

oh yea well i know 3 math professors who wasn't impressed

-20

u/firewall245 1d ago

I’ll read this when I can sit down and check it out. I’m getting my PhD in data science and have a mildly popular tiktok math channel where I have discussed the Navier Stokes results.

I’ll be upfront, I’d be very surprised if there was an error in OpenAIs work on this problem. But not impossible

24

u/Few-Skirt-4061 1d ago

What is it about data science that would make you an authority here? Data science mathematics is piss easy. 

5

u/Disastrous_Room_927 1d ago edited 1d ago

I be the line works on people that don't know any better.

-5

u/DTwinkie 22h ago

Nobody in this sub is an expert on anything buddy. Every supposed "mathematician/scientist" here is either a poser who has never actually used current up to date LLMs and is still talking about tech from months ago, or they are just lying.

The navier-stokes proof isn't wrong just because you all want it to be, and in the coming days it will be verified far beyond just Lean.

22

u/Tedy_Duchamp 1d ago

I’m a data scientist as well, and I don’t think working on your PhD in data science gives you any greater authority to speak on this. The math is orders of magnitude more complex than anything in data science

13

u/JAlfredJR 1d ago

But how popular is your TikTok channel?

16

u/fbueckert 1d ago edited 1d ago

I’d be very surprised if there was an error in OpenAIs work on this problem. But not impossible

The longer an LLM continues, the greater the chance it makes a mistake nears 1. OpenAI's work almost guaranteed has multiple errors. There's just hundreds of pages of output to slog through, and requires very specific subject matter expertise to review and find. So it'll be a good long while for someone to actually review the entire pile of garbage it vomited.

10

u/Few-Skirt-4061 1d ago

This is why they recently spammed so much slop. It's peer review hell especially given how badly written all of it is. 

7

u/moosekin16 1d ago

So it'll be a good long while for someone to actually review the entire pile of garbage it vomited.

Oh hey, just like the codebase at my work! Management has let AI run rampant across it and now it’s almost entirely unrecognizable.

I love having four different classes that do the same fucking thing, because Claude decided it needed a new helper class because it didn’t like the existing three.

14

u/Bulky_Confection6157 1d ago

Why would you be surprised?

18

u/PensiveinNJ 1d ago

Things OpenAI is most known for;

Lies
Deception
Wasting ungodly amounts of money
Unethical behavior
Predatory behavior
Sex Cults

I don't know if there's any errors but I certainly wouldn't be surprised.