Jose Felgueroso
Abogado | Attorney

  • Home
  • Blog


Fair Use of Copyrighted Works in AI Training: Legal Standards and Emerging Trends

info@josefelgueroso.com

Versión en español

2025-07-14

Artificial intelligence systems are rapidly transforming how information is created, accessed, and used—but their development raises complex questions about the fair use of copyrighted works in training data. As generative AI models become more sophisticated, courts and policymakers are grappling with whether and when using protected content to train AI constitutes copyright infringement or falls within the boundaries of fair use. This article examines the evolving legal standards that govern these questions, focusing on recent U.S. federal court decisions and one policy report.




The Doctrine of Fair Use and Its Factors

The doctrine of fair use is a foundational principle in U.S. copyright law, codified at 17 U.S.C. § 107. It permits certain uses of copyrighted works without the rights holder’s consent, balancing the interests of creators with the public’s interest in access, innovation, and free expression. Fair use is an affirmative defense to copyright infringement and is assessed on a case-by-case basis, guided by four statutory factors.

Statutory Basis and Purpose

Fair use exists to foster creativity, scholarship, and innovation by allowing limited, socially valuable uses of copyrighted material. It is intentionally flexible, enabling courts to adapt its application to new technologies and evolving societal needs. The doctrine is crucial in contexts such as criticism, commentary, news reporting, teaching, research, and—more recently—technological development, including artificial intelligence.

The Four Fair Use Factors

Courts analyze fair use by weighing four non-exclusive factors, considering the totality of circumstances:

  1. Purpose and Character of the Use
    • Considers whether the use is commercial or nonprofit, and whether it is transformative—i.e., does it add new expression, meaning, or message, or merely supersede the original?
    • Transformative uses are favored, but commercial uses are not automatically disqualified. The more transformative the use, the less significant other factors may become.
  2. Nature of the Copyrighted Work
    • Evaluates whether the original work is more factual or creative. Uses of factual or published works are more likely to be fair than uses of highly creative, unpublished works.
    • This factor typically plays a lesser role but can be decisive in close cases.
  3. Amount and Substantiality of the Portion Used
    • Assesses both the quantitative and qualitative extent of the use. Copying the “heart” of a work, even if a small portion, may weigh against fair use.
    • However, using the entire work can be fair if it is reasonable in light of the transformative purpose (e.g., for technological or archival uses).
  4. Effect of the Use Upon the Potential Market
    • Considers whether the use harms the actual or potential market for the original or its derivatives. This is often the most important factor.
    • Market harm can include lost sales, diminished licensing opportunities, or market dilution from widespread, substitutive uses.

Holistic and Contextual Analysis

No single factor is determinative: courts weigh all factors together in light of the purposes of copyright law. The analysis is highly contextual and may evolve as new technologies and uses emerge. The Supreme Court and lower courts have repeatedly emphasized that fair use is not a mechanical checklist but a flexible, equitable doctrine that adapts to changing circumstances.

Key Takeaways

  • Case-by-Case Inquiry: Fair use is assessed individually for each use, considering all relevant facts and circumstances.
  • Transformative Use Favored: Uses that add new meaning or purpose, rather than merely substituting for the original, are more likely to be fair.
  • Market Impact Critical: Uses that harm the market for the original or its derivatives are less likely to be fair, especially if they compete directly or undermine licensing opportunities.
  • Flexibility for Innovation: The doctrine’s adaptability is essential for addressing novel issues, such as AI training, that were not foreseen by earlier copyright regimes.

Summary and Key Points from Each Source

1. Thomson Reuters v. Ross Intelligence (D. Del., 2025)

This decision from the Federal District Court of Delaware addressed whether using copyrighted legal headnotes to train a legal AI tool constituted copyright infringement or fair use. The court found Ross Intelligence liable for direct copyright infringement, rejecting its fair use defense. The ruling clarified that editorial content, even if derived from public domain sources, can be protected by copyright, and that using such content for commercial AI training—especially when the AI tool competes with the original—does not qualify as fair use.

  • Copyrightability: Editorial content like Westlaw’s headnotes and Key Number System are original and copyrightable, even when summarizing public domain material.
  • Direct Infringement: Ross’s use of Westlaw headnotes in AI training constituted actual copying and substantial similarity.
  • Fair Use Factors: The court held that Ross’s use was commercial, non-transformative, and harmed both the actual and potential market for licensing AI training data, tipping the balance against fair use.
  • Market Impact: The competitive use of copyrighted material in AI tools that serve the same market as the original weighs heavily against fair use.
  • Defenses Rejected: Innocent infringement, copyright misuse, merger, and commonplace-features defenses all failed.

2. "Copyright and Artificial Intelligence, Part 3: Generative AI Training" (U.S. Copyright Office, May 2025)

This comprehensive policy report analyzes the application of copyright law and fair use to the use of copyrighted works in generative AI training. It surveys technical practices, legal theories, and policy considerations, providing a framework for evaluating whether AI training on copyrighted works is infringing or fair use. The report does not take a categorical position but highlights the complexity and fact-specific nature of the current legal landscape.

  • Fair Use Analysis: The report details how each fair use factor applies to AI training, emphasizing that outcomes are highly fact-dependent and evolving.
  • Transformativeness: AI training may be transformative if it enables new uses, but wholesale copying of works for training generally weighs against fair use unless justified by necessity and minimal public exposure of protected content.
  • Market Harm: The potential for lost sales, market dilution, and lost licensing opportunities is a central concern, especially if AI outputs can substitute for original works.
  • Licensing and Opt-Outs: The report discusses voluntary and statutory licensing models, collective licensing, and technical/legal opt-out mechanisms for rights holders.
  • International Context: U.S. law is contrasted with international approaches, some of which provide explicit TDM (text and data mining) exceptions.

3. Bartz v. Anthropic (N.D.Cal., 2025)

This decision from the Federal District Court for Northern California evaluated whether Anthropic’s use of copyrighted books—acquired both by purchase and via shadow libraries—for training generative AI models was fair use. The court distinguished between different uses: it held that using lawfully acquired books for AI training was fair use due to the highly transformative nature of the use, but that maintaining a permanent library of pirated copies was not excused by fair use.

  • Transformative Use: Training generative AI models on copyrighted books was deemed “quintessentially transformative” when outputs did not reproduce or substitute for the originals.
  • Pirated Copies: Building a permanent, general-purpose library of pirated books was not a fair use, as it was not transformative and displaced legitimate sales.
  • Format Shifting: Scanning purchased print books for internal research and training was allowed as fair use, provided originals were destroyed and no distribution occurred.
  • No Output Infringement: The court’s ruling was limited to cases where the AI system did not output infringing copies of the training works.
  • Market Effect: The court found no actionable market harm from training alone, absent evidence of output substitution.

4. Kadrey v. Meta (N.D.Cal., 2025)

This decision addressed whether Meta’s use of copyrighted books—downloaded from shadow libraries—to train its Llama large language models constituted fair use. The court found Meta’s use highly transformative, but emphasized that the most important fair use factor is market harm. The court granted summary judgment for Meta, finding no evidence that Llama could output substantial portions of plaintiffs’ works or that AI training had caused market harm to the specific plaintiffs. However, the decision noted that future cases with better evidence of market dilution could reach a different result.

  • Transformative Purpose: Training LLMs to generate diverse text was considered a highly transformative use, distinct from the original purpose of the books.
  • Market Harm Central: The court stressed that fair use will often fail if there is substantial market dilution or substitution, but found no such evidence in this case.
  • No Output Substitution: Llama was not shown to output substantial or infringing portions of plaintiffs’ books; thus, direct substitution was not established.
  • Licensing Market: The court held that copyright holders are not entitled to monopolize the market for licensing works as AI training data if the use is otherwise fair.
  • Fact-Specific Inquiry: The ruling was limited to the record before the court, leaving open the possibility that other plaintiffs could prevail with stronger evidence of market harm.

Points of Agreement and Disagreement Between the Sources

Issue Agreement Disagreement
Copyrightability of Training Data ▪ All sources recognize that original, creative editorial content (such as headnotes or authored books) is protected by copyright, even if derived from public domain materials.
▪ There is consensus that the originality threshold is low, but present for curated or synthesized content.
▪ The Thomson Reuters v. Ross decision strongly affirms copyright in editorial content used for AI training, while the Copyright Office report and the California district court opinions acknowledge the issue but focus more on the nature of downstream uses and factual context.
Transformative Use and Fair Use ▪ All sources agree that the transformative nature of AI training—whether it adds new purpose or meaning—is a central inquiry under the first fair use factor.
▪ There is consensus that copying for AI training may be transformative if it enables new uses or functionalities not intended by the original work.
▪ Bartz v. Anthropic and Kadrey v. Meta both find AI training on books to be “quintessentially transformative” when outputs do not reproduce or substitute for the originals.
▪ Thomson Reuters v. Ross, by contrast, finds no transformative use where the AI system directly competes with the original work and serves the same market function.
▪ The Copyright Office report notes that transformativeness is highly fact-dependent and may not always be present, especially if the use is not necessary or outputs are substitutive.
Market Harm and Licensing ▪ All sources emphasize the fourth fair use factor—market harm—as critical.
▪ There is general agreement that actual or potential market substitution, including lost licensing opportunities, weighs against fair use.
▪ Thomson Reuters v. Ross holds that direct competition in the same market and lost licensing opportunities for AI training data are decisive against fair use.
▪ Kadrey v. Meta finds no actionable market harm absent evidence that AI outputs substitute for the originals or that a licensing market for training data is one copyright holders are entitled to control.
▪ Bartz v. Anthropic similarly finds no market harm where there is no evidence of output substitution, but distinguishes between lawful and pirated source copies.
▪ The Copyright Office report highlights uncertainty and evolving licensing markets, noting that voluntary and statutory licensing models are under consideration.
Use of Pirated or Lawfully Acquired Works ▪ All sources agree that use of lawfully acquired works for AI training is more defensible under fair use than use of pirated or unauthorized copies.
▪ Bartz v. Anthropic draws a sharp line: training on lawfully acquired books is fair use, but building a permanent library of pirated books is not excused by fair use.
▪ Kadrey v. Meta does not directly address the piracy issue but focuses on the absence of output substitution and market harm.
▪ The Copyright Office report discusses the legal and practical challenges of sourcing training data, including the prevalence of unauthorized sources.
Scope and Flexibility of Fair Use ▪ All sources acknowledge that fair use is a flexible, fact-specific doctrine that must adapt to technological change.
▪ Courts and policymakers agree that no single factor is determinative; the analysis is holistic.
▪ The sources differ in how much weight they assign to each factor and the circumstances under which AI training will or will not be deemed fair use.
▪ The Copyright Office report is more cautious, emphasizing the evolving and unsettled nature of the law, while the district court opinions reach more categorical conclusions based on the evidentiary record in each case.

Key Takeaways

  • Consensus exists on the importance of originality, the centrality of the market harm factor, and the need for a holistic, context-sensitive fair use analysis.
  • Disagreements center on the degree of transformativeness in AI training, the significance of lost licensing markets, and the treatment of pirated versus lawfully acquired works.
  • Fact-specific outcomes: The cases diverge sharply based on whether the AI’s outputs substitute for the originals and whether the training data was lawfully obtained.
  • Policy uncertainty: The Copyright Office report underscores ongoing legal and policy debates, while the courts are beginning to draw sharper lines based on market impact and the nature of the AI system.

Facts That Tend to Make a Use Fair or Unfair

This section distills the key factual patterns from recent case law and policy analysis regarding the use of copyrighted works in AI training. The following points summarize which circumstances typically favor a finding of fair use, and which weigh against it.

Facts That Tend to Support a Finding of Fair Use

  • Highly Transformative Purpose: Uses that repurpose copyrighted works for a fundamentally new function—such as training generative AI models to produce diverse, non-infringing outputs—are more likely to be deemed fair. Courts have found that training large language models on books, when outputs do not reproduce or substitute for the originals, is “quintessentially transformative.”
  • No Output Substitution: If the AI system does not generate outputs that are substantially similar to or compete with the original works, this weighs in favor of fair use. The absence of evidence that the model regurgitates or enables access to the original content is critical.
  • Lawful Acquisition of Training Data: Using lawfully acquired or licensed works for AI training is more defensible under fair use than using pirated or unauthorized copies. Courts have distinguished between format-shifting lawfully purchased books for internal research (often fair) and building permanent libraries from pirated sources (not fair).
  • No Actionable Market Harm: Where there is no evidence that the AI’s training or outputs have caused actual or potential market harm—such as lost sales, diminished licensing opportunities, or market dilution—courts are more likely to find fair use. The inability of plaintiffs to show concrete market impact has been decisive.
  • Incidental or Necessary Copying: Copying that is incidental to a transformative technological process, and not for the purpose of supplanting the original market, may be permissible—especially if the amount taken is reasonable in relation to the new use.
  • Factual or Published Works: Uses involving factual, informational, or already published works are more likely to be fair than uses involving highly creative or unpublished works, though this factor is rarely dispositive.

Facts That Tend to Weigh Against Fair Use

  • Direct Market Competition: When the AI system or its outputs serve the same market function as the original work—such as a legal research tool trained on a competitor’s editorial content—this strongly weighs against fair use.
  • Non-Transformative or Redundant Use: Uses that do not add new meaning, purpose, or function, or that merely substitute for the original, are unlikely to be fair. Intermediate copying is not transformative if it is not necessary to achieve a new purpose.
  • Use of Pirated or Unauthorized Copies: Building a permanent library of pirated works, even if only some are used for training, is not excused by fair use. Courts have rejected the argument that a transformative end justifies unlawful acquisition of source material.
  • Substantial Market Harm or Lost Licensing Opportunities: Evidence of substantial market harm weighs heavily against fair use, as courts deem this factor the "single most important element" of fair use analysis. Courts are, however, divided on AI training licensing markets: some recognize harm to traditional markets (lost sales, market dilution from AI-generated content) while others reject licensing market harm as "not cognizable" to avoid circular reasoning. This remains an evolving area with conflicting judicial approaches.
  • Output Substitution or Regurgitation: If the AI model is shown to output substantial portions of the original works, or if its outputs serve as substitutes for the originals, this will likely defeat a fair use defense.
  • Bad Faith or Evasive Conduct: While not always dispositive, evidence of bad faith—such as intentionally circumventing licensing or using unauthorized sources—can tip the balance against fair use, but its weight in the fair use balancing act is a point of divergence among courts.

Key Factual Themes

Factual Scenario Tends Toward Fair Use Tends Toward Unfair Use
Transformative AI training, no output substitution ✔️
Direct competition with original market ✔️
Lawful acquisition of works ✔️
Use of pirated/unauthorized copies ✔️
No evidence of market harm ✔️
Substantial market/licensing harm ✔️
Outputs regurgitate original content ✔️
Factual or published works ✔️
Highly creative/unpublished works ✔️

Practical Implications

  • AI developers should prioritize using lawfully acquired or licensed data, avoid training on competitors’ proprietary content, and implement safeguards to prevent output substitution.
  • Rights holders should document actual or potential market harm, including lost licensing opportunities, to strengthen claims against unauthorized AI training.
  • Fair use remains a flexible, fact-intensive doctrine, and outcomes will continue to depend on the specific facts and evidence presented in each case.

Sources

  • 17 U.S.C. § 107 - Limitations on exclusive rights: Fair use
  • Thomson Reuters v. Ross Intelligence (D. Del., 2025)
  • "Copyright and Artificial Intelligence, Part 3: Generative AI Training" (U.S. Copyright Office, May 2025)
  • Bartz v. Anthropic (N.D.Cal., 2025)
  • Kadrey v. Meta (N.D.Cal., 2025)