AI models need more biological data; OpenAI funds its creation
funding openai
| Source: Mastodon | Original article
OpenAI is funding research to create new biological data, aiming to boost AI models’ capability in biomedicine and public‑health applications.
OpenAI’s nonprofit arm has announced a new funding programme aimed at bolstering the biological data that underpins medical‑focused artificial‑intelligence systems. The OpenAI Foundation will back an initiative dubbed “Data for Public Health” (also referred to as “Public Data for Health”), which will pay to create high‑quality scientific datasets that are currently scarce in the AI research ecosystem.
The move builds on a proposal floated last year by Ruxandra Teslo, a policy analyst specialising in clinical trials. Teslo suggested that AI developers could accelerate progress in predictive biology by acquiring data from biotech firms that have gone bankrupt. By bidding in those bankruptcy proceedings, she argued, it would be possible to obtain detailed regulatory filings, manufacturing strategies and safety data that are normally locked away as trade secrets. OpenAI has now taken up that idea, committing resources to purchase and curate the information so it can be safely used to train and fine‑tune medical models.
Why the funding matters is twofold. First, current large language models and other AI systems lack the depth of domain‑specific knowledge needed to generate reliable insights in drug discovery, diagnostics or public‑health forecasting. Access to richer, vetted biological datasets could close that gap and enable breakthroughs that are difficult to achieve with generic web‑scale data alone. Second, the approach raises questions about data provenance, intellectual‑property rights and the ethics of repurposing confidential corporate information, even when the original owners have entered liquidation.
What to watch next includes OpenAI’s criteria for selecting bankruptcy cases, the mechanisms it will use to ensure that the acquired data respects privacy and regulatory constraints, and how the broader AI‑health community responds. Industry observers will be looking for partnerships with legal firms that handle biotech insolvencies, as well as any guidance from regulators on the permissible reuse of such data. The success of the programme could set a precedent for how AI developers source specialised scientific information, potentially reshaping the relationship between the biotech sector and the rapidly expanding field of health‑focused artificial intelligence.
Sources
Back to AIPULSEN