Researcher claims OpenAI trained on conversations, heralds breakthrough
anthropic openai
| Source: HN | Original article
A researcher alleges OpenAI trained its models on conversational data before announcing a breakthrough.
OpenAI’s claim of a “big math breakthrough” has been challenged again, this time by a researcher who says the company may have trained its models on the very conversations that later underpinned the announcement. The researcher, who has remained unnamed in public reports, wrote to OpenAI scientists Sébastien Bubeck and Mark Sellke asking whether his chats with ChatGPT were part of the training data that helped solve the problem. OpenAI replied that no specific user data was accessed, but conceded that “de‑identified data derived from their usage of our products could have helped improve our models.” The response, critics say, does not address whether the researcher’s interactions entered the massive, opaque pools of data that fuel model updates.
The dispute revives a controversy first highlighted on 10 September, when we interviewed mathematician Tristan Buckmaster about his collaboration with Anthropic researcher Levent Alpöge and OpenAI’s subsequent claim of progress on a Millennium Prize problem. Buckmaster has repeatedly asked OpenAI to clarify whether his own conversations were used, and has accused the firm of pressuring him to co‑ordinate the public announcement. Anthropic’s Sholto Douglas defended OpenAI, insisting that user transcripts would not be harvested for training.
Why it matters is twofold. First, the episode raises fresh questions about the transparency of AI training pipelines and the extent to which proprietary models may incorporate user‑generated scientific content without explicit consent. Second, it touches on academic credit: if a breakthrough is partially derived from private exchanges, the rightful attribution of discovery becomes murky, potentially eroding trust between researchers and AI providers.
The story is still unfolding. Observers will watch for any formal clarification from OpenAI, possible responses from the research community, and whether regulators or institutional review boards will demand more rigorous data‑use disclosures. Further statements from the unnamed researcher and any follow‑up investigations could shape how AI companies handle confidential scientific dialogue in the future.
Sources
Back to AIPULSEN