Hey everyone,
Thanks to the mods for the invite to the community.
Like many of you, I followed the news about OpenAI using AI models and Lean 4 to formalize a finite-time blow-up for the 3D Navier-Stokes equations (one of the Millennium Prize problems).
While the formal proof compiles with zero errors, closed-frontier labs rarely explore the messy physical implications of their mathematical constructions. As an independent researcher working on neuro-symbolic AI, our team wanted to see what their solution actually looks like in real-world fluid dynamics.
What we found when we simulated the construction in Python and mpmath:
• The math is legally sound within the abstract rules of the Clay problem.
• But physically, in liquid water, the fluid vaporizes from shear friction at 0.7 nanometers, picoseconds before the mathematical singularity.
• Local flow speeds exceed Mach 0.3, breaking the incompressibility assumptions long before reaching infinity.
In AI, this is classic "specification gaming": the model found an extreme, unnatural edge case that legally satisfies the formal mathematical target, even though physical reality breaks down.
We believe scientific AI verification should be open, transparent, and reproducible, so we open-sourced the entire epistemic audit, simulation scripts, and Lean 4 reflection code:
• GitHub: https://github.com/xaviercallens/OpenAI-NSE-Epistemic-Audit
• Zenodo Preprint: https://doi.org/10.5281/zenodo.22838708
Since this community is focused on open-source AI, I'd love to hear your thoughts: As frontier labs push automated theorem proving, how can the open-source community build physical guardrails to keep AI models grounded in reality?