info@josefelgueroso.com
2025-07-14
Artificial intelligence systems are rapidly transforming how information is created, accessed, and used—but their development raises complex questions about the fair use of copyrighted works in training data. As generative AI models become more sophisticated, courts and policymakers are grappling with whether and when using protected content to train AI constitutes copyright infringement or falls within the boundaries of fair use. This article examines the evolving legal standards that govern these questions, focusing on recent U.S. federal court decisions and one policy report.
The doctrine of fair use is a foundational principle in U.S. copyright law, codified at 17 U.S.C. § 107. It permits certain uses of copyrighted works without the rights holder’s consent, balancing the interests of creators with the public’s interest in access, innovation, and free expression. Fair use is an affirmative defense to copyright infringement and is assessed on a case-by-case basis, guided by four statutory factors.
Fair use exists to foster creativity, scholarship, and innovation by allowing limited, socially valuable uses of copyrighted material. It is intentionally flexible, enabling courts to adapt its application to new technologies and evolving societal needs. The doctrine is crucial in contexts such as criticism, commentary, news reporting, teaching, research, and—more recently—technological development, including artificial intelligence.
Courts analyze fair use by weighing four non-exclusive factors, considering the totality of circumstances:
No single factor is determinative: courts weigh all factors together in light of the purposes of copyright law. The analysis is highly contextual and may evolve as new technologies and uses emerge. The Supreme Court and lower courts have repeatedly emphasized that fair use is not a mechanical checklist but a flexible, equitable doctrine that adapts to changing circumstances.
This decision from the Federal District Court of Delaware addressed whether using copyrighted legal headnotes to train a legal AI tool constituted copyright infringement or fair use. The court found Ross Intelligence liable for direct copyright infringement, rejecting its fair use defense. The ruling clarified that editorial content, even if derived from public domain sources, can be protected by copyright, and that using such content for commercial AI training—especially when the AI tool competes with the original—does not qualify as fair use.
This comprehensive policy report analyzes the application of copyright law and fair use to the use of copyrighted works in generative AI training. It surveys technical practices, legal theories, and policy considerations, providing a framework for evaluating whether AI training on copyrighted works is infringing or fair use. The report does not take a categorical position but highlights the complexity and fact-specific nature of the current legal landscape.
This decision from the Federal District Court for Northern California evaluated whether Anthropic’s use of copyrighted books—acquired both by purchase and via shadow libraries—for training generative AI models was fair use. The court distinguished between different uses: it held that using lawfully acquired books for AI training was fair use due to the highly transformative nature of the use, but that maintaining a permanent library of pirated copies was not excused by fair use.
This decision addressed whether Meta’s use of copyrighted books—downloaded from shadow libraries—to train its Llama large language models constituted fair use. The court found Meta’s use highly transformative, but emphasized that the most important fair use factor is market harm. The court granted summary judgment for Meta, finding no evidence that Llama could output substantial portions of plaintiffs’ works or that AI training had caused market harm to the specific plaintiffs. However, the decision noted that future cases with better evidence of market dilution could reach a different result.
| Issue | Agreement | Disagreement |
|---|---|---|
| Copyrightability of Training Data | ▪ All sources recognize that original, creative editorial content (such as headnotes or authored books) is protected by copyright, even if derived from public domain materials. ▪ There is consensus that the originality threshold is low, but present for curated or synthesized content. | ▪ The Thomson Reuters v. Ross decision strongly affirms copyright in editorial content used for AI training, while the Copyright Office report and the California district court opinions acknowledge the issue but focus more on the nature of downstream uses and factual context. |
| Transformative Use and Fair Use | ▪ All sources agree that the transformative nature of AI training—whether it adds new purpose or meaning—is a central inquiry under the first fair use factor. ▪ There is consensus that copying for AI training may be transformative if it enables new uses or functionalities not intended by the original work. | ▪ Bartz v. Anthropic and Kadrey v. Meta both find AI training on books to be “quintessentially transformative” when outputs do not reproduce or substitute for the originals. ▪ Thomson Reuters v. Ross, by contrast, finds no transformative use where the AI system directly competes with the original work and serves the same market function. ▪ The Copyright Office report notes that transformativeness is highly fact-dependent and may not always be present, especially if the use is not necessary or outputs are substitutive. |
| Market Harm and Licensing | ▪ All sources emphasize the fourth fair use factor—market harm—as critical. ▪ There is general agreement that actual or potential market substitution, including lost licensing opportunities, weighs against fair use. | ▪ Thomson Reuters v. Ross holds that direct competition in the same market and lost licensing opportunities for AI training data are decisive against fair use. ▪ Kadrey v. Meta finds no actionable market harm absent evidence that AI outputs substitute for the originals or that a licensing market for training data is one copyright holders are entitled to control. ▪ Bartz v. Anthropic similarly finds no market harm where there is no evidence of output substitution, but distinguishes between lawful and pirated source copies. ▪ The Copyright Office report highlights uncertainty and evolving licensing markets, noting that voluntary and statutory licensing models are under consideration. |
| Use of Pirated or Lawfully Acquired Works | ▪ All sources agree that use of lawfully acquired works for AI training is more defensible under fair use than use of pirated or unauthorized copies. | ▪ Bartz v. Anthropic draws a sharp line: training on lawfully acquired books is fair use, but building a permanent library of pirated books is not excused by fair use. ▪ Kadrey v. Meta does not directly address the piracy issue but focuses on the absence of output substitution and market harm. ▪ The Copyright Office report discusses the legal and practical challenges of sourcing training data, including the prevalence of unauthorized sources. |
| Scope and Flexibility of Fair Use | ▪ All sources acknowledge that fair use is a flexible, fact-specific doctrine that must adapt to technological change. ▪ Courts and policymakers agree that no single factor is determinative; the analysis is holistic. | ▪ The sources differ in how much weight they assign to each factor and the circumstances under which AI training will or will not be deemed fair use. ▪ The Copyright Office report is more cautious, emphasizing the evolving and unsettled nature of the law, while the district court opinions reach more categorical conclusions based on the evidentiary record in each case. |
This section distills the key factual patterns from recent case law and policy analysis regarding the use of copyrighted works in AI training. The following points summarize which circumstances typically favor a finding of fair use, and which weigh against it.
| Factual Scenario | Tends Toward Fair Use | Tends Toward Unfair Use |
|---|---|---|
| Transformative AI training, no output substitution | ✔️ | |
| Direct competition with original market | ✔️ | |
| Lawful acquisition of works | ✔️ | |
| Use of pirated/unauthorized copies | ✔️ | |
| No evidence of market harm | ✔️ | |
| Substantial market/licensing harm | ✔️ | |
| Outputs regurgitate original content | ✔️ | |
| Factual or published works | ✔️ | |
| Highly creative/unpublished works | ✔️ |