Lean verification of AI autoformalisation doesn't ensure correct Navier‑Stokes proofs
openai
| Source: Mastodon | Original article
Researchers demonstrate that a Lean‑verified AI‑generated formal proof of Navier‑Stokes blow‑up fails to match the corresponding natural‑language proof, highlighting limits of auto‑formalisation.
A new arXiv paper — *Navier‑Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs* — raises doubts about the reliability of formal verification pipelines that translate AI‑generated mathematics into the Lean proof assistant. The authors, Bastounis, Circelli and Hansen, demonstrate that the Lean proof produced for OpenAI’s announced claim of blow‑up solutions to the Navier‑Stokes equations does not correspond to the natural‑language argument that was publicised. In other words, the formal proof that passes Lean’s kernel checks proves a different statement than the one claimed in the original text.
The finding matters because auto‑formalisation has become a cornerstone of how researchers and companies, including OpenAI, seek to certify AI‑generated results. If the translation step can silently introduce mismatches, the mere existence of a verified Lean proof no longer guarantees that the underlying mathematical claim is sound. This undermines confidence in a growing body of AI‑produced preprints, such as the 700 mathematical papers OpenAI released earlier this month [2026‑10‑07].
The paper supplies concrete examples of mistranslations and highlights a systemic failure mode: a verified proof that “verifies” the formalisation, not the original argument. The authors call for tighter coupling between natural‑language reasoning and its formal counterpart, and for tools that can detect semantic drift during translation.
Going forward, the community will watch for OpenAI’s response—whether it will revise the Navier‑Stokes claim, improve its auto‑formalisation pipeline, or provide additional evidence of correctness. Researchers are also likely to develop new verification standards that require traceability between natural‑language statements and their formal proofs, and to scrutinise other AI‑generated results that have been “Lean‑verified” without such checks. The debate signals a broader reassessment of how much trust can be placed in AI‑assisted mathematics.
Sources
Back to AIPULSEN