Third Circuit says AI training on copyrighted material isn’t fair use
copyright training
| Source: Mastodon | Original article
The U.S. Third Circuit ruled that using copyrighted works to train AI models does not qualify as fair use, underscoring legal limits on AI training data.
The U.S. Court of Appeals for the Third Circuit has ruled that an artificial‑intelligence company’s use of a competitor’s copyrighted material for training does not qualify as fair use. The case involved ROSS Intelligence, a now‑defunct legal‑tech startup that copied 2,243 Westlaw headnotes to train its own legal‑research AI. The court concluded that the copying was a commercial act that infringed Westlaw’s copyright, rejecting ROSS’s argument that the training process was transformative.
The decision is significant because it marks one of the first appellate rulings directly addressing the legality of using copyrighted works as training data for AI systems that compete with the original source. By affirming that such copying can constitute infringement, the Third Circuit sets a potential benchmark for future disputes across the burgeoning AI market. Legal‑tech firms and other AI developers may now face heightened pressure to secure licenses for the texts, images, or code they use to train models, reshaping business models that have relied on freely scraping large corpora.
The opinion also drew a line between ROSS’s product and generative AI systems that create new expressive output, leaving the broader question of whether training on books, articles or other expressive works is fair use unresolved. That gap suggests further litigation could surface in other circuits, especially as large language models continue to dominate the AI landscape.
Watch for an appeal by ROSS’s successors, possible Supreme Court review, and legislative initiatives aimed at clarifying copyright rules for AI training data. Industry groups are already signaling a push toward clearer licensing frameworks, and the next few months are likely to see intensified debate over how intellectual‑property law will adapt to AI’s data‑hungry reality.
Sources
Back to AIPULSEN