OpenAI says it cannot rule out that de‑identified data from Buckmaster and Alpöge's product use improved its models
openai
| Source: Techmeme | Original article
OpenAI acknowledges that, though unlikely, anonymized data from Buckmaster's and Alpöge's use of its products could have contributed to improving its models.
OpenAI has issued a qualified disclaimer about the data that may have fed its recent claim of a Navier‑Stokes breakthrough. In a statement released alongside the paper that formalises a proof of finite‑time singularities in the Navier‑Stokes equations, the company said that, while it considers it “unlikely,” it “cannot rule out that de‑identified data derived from” the usage of its products by mathematicians Thomas Buckmaster and Sebastian Alpöge helped improve its models.
The comment follows OpenAI’s earlier announcement that an internal system had produced a proof of the Millennium‑Prize Navier‑Stokes problem – a claim we covered on 9 September 2026. The new wording addresses lingering doubts about whether the system’s training data included, even in anonymised form, the prompts or intermediate calculations supplied by Buckmaster and Alpöge as they worked on the problem. OpenAI stresses that its researchers did not directly see the users’ prompts and that any influence would be indirect, arising only from aggregated, de‑identified usage data.
The clarification matters because it touches on two hot‑button issues in AI: the provenance of training data and the attribution of scientific breakthroughs. If proprietary or unpublished research can inadvertently become part of a model’s training set, questions arise about intellectual‑property rights, academic credit, and the fairness of claiming a discovery as “AI‑generated.” The admission also fuels scrutiny from the broader research community, which has already debated the validity of OpenAI’s Navier‑Stokes proof.
Going forward, observers will watch for any formal response from Buckmaster and Alpöge, potential regulatory inquiries into data‑use practices, and whether OpenAI will adjust its training pipelines or disclosure policies. The episode could shape how future AI‑driven scientific claims are vetted and credited.
Sources
Back to AIPULSEN