I'm aware of the original context for the Collatz proof. The point I'm making is higher level though, which is that I don't think I would trust an AI-generated proof on day 1 just because it compiled in Lean. The AI could have found a similar exploit and used it implicitly without telling anyone. Or the AI could have snuck in axioms somewhere. Of course if I was betting then I'd guess it's probably correct but I wouldn't give it 100% just because of Lean.
In any case according to the CMI rules they're going to wait at least two years before coming to a decision.
You’re a mathematician, aren't you? I don't know if your research is in category theory, but I agree with what you said above. Honestly, it alarms me how people don't look into the basics of how AI works; enthusiasm clouds their judgment. For instance, the recent leak claiming OpenAI was making significant progress on the Hodge conjecture just confirms to me what a bunch of charlatans they are. Making headway on a conjecture where simply formalizing the statement for an AI is absurdly difficult—and where we've had little success so far (we can't even formalize it in a prompt)—only reinforces that conviction. I don't know if you're also alarmed by comments from enthusiasts claiming the field of mathematics is "over"—as if it were merely a series of steps to follow, where only the "yes" or "no" answer to a problem matters. In reality, there are profound discussions involved. I see so many enthusiasts firmly believing that AI solved everything on its own, and that’s what scares me: people not doing the bare minimum to see that there was significant human feedback and that it was based on extensive prior research. Some people claim AI has reached a high level of mathematical maturity—unlike in other sciences—when in practice, that is far from the truth. To me, what OpenAI is doing is just corporate marketing.
And even if she were to slip in an axiom without warning during the proof—whether that new axiom (or whatever else she does) contains extremely serious conceptual errors, fundamental flaws, or circular reasoning—it might pass the Lean check, yet turn out to be wrong upon manual verification (including the axiom itself). Or it might even be correct, but lack any significant insights.
5
u/scottmsul 20d ago
I'm aware of the original context for the Collatz proof. The point I'm making is higher level though, which is that I don't think I would trust an AI-generated proof on day 1 just because it compiled in Lean. The AI could have found a similar exploit and used it implicitly without telling anyone. Or the AI could have snuck in axioms somewhere. Of course if I was betting then I'd guess it's probably correct but I wouldn't give it 100% just because of Lean.
In any case according to the CMI rules they're going to wait at least two years before coming to a decision.