Lean verification of AI autoformalisation doesn’t guarantee correct Navier‑Stokes proofs
openai
| Source: Mastodon | Original article
A recent AI autoformalisation of the Navier‑Stokes blow‑up theorem in Lean reveals a mismatch between the natural‑language proof and the formal proof, highlighting limits of verification.
OpenAI’s much‑talked‑about claim of a blow‑up proof for the Navier‑Stokes equations has hit a snag. A new pre‑print demonstrates that the Lean formalisation the company released does not correspond to the natural‑language (NL) argument it was meant to capture. While the Lean proof checks out mechanically, the NL proof – the version presented to mathematicians – appears to diverge, meaning the human‑readable reasoning could be incorrect.
The discrepancy matters because auto‑formalisation – the process of translating NL mathematics into a proof assistant like Lean – is increasingly touted as a way to let AI verify complex results. If the translation is not semantically faithful, the formal proof may verify a different statement altogether, undermining confidence in AI‑generated mathematics. The authors of the pre‑print place perfectly faithful auto‑formalisation at the top of the Solvability Complexity Index hierarchy (SCI = ∞), above even the Halting problem (SCI = 1), underscoring how hard it is to guarantee a one‑to‑one mapping between human prose and formal code.
As we reported on 2026‑10‑07, the mismatch between the two proofs was already noted in community comments, but the new analysis formalises the gap and calls it a “semantic mismatch.” The episode will likely prompt tighter scrutiny of AI‑driven proof pipelines, revisions to OpenAI’s submission, and renewed interest in methods that can certify the fidelity of auto‑formalisation. Watch for responses from OpenAI, possible updated Lean scripts, and broader discussions in the mathematical AI community about standards for linking human‑readable arguments to machine‑checked proofs.
Sources
Back to AIPULSEN