Fair Use
Motivation
Generative AI’s improvement across text, audio, and video is in large part due to the availability of huge amounts of training data. For instance, GPT-3’s was trained on about 240 billion unique tokens (equivalent to a few million books), while more recent models like Llama 3 were trained on 15 trillion tokens.¹ This data-heavy approach enabled breakthrough capabilities but relied on the critical assumption that publicly-accessible web content could be freely scraped and used for training.
However, much of this data may be protected as copyrighted material, and training on it without the permission of the data’s creators assumes that AI training is considered “fair use.” In other words, creators of the data (e.g., writers and artists) have exclusive rights over it, subject to some exceptions, so does AI training count as one of those legal exceptions (namely, fair use)?
This question remains unresolved. There are several ongoing lawsuits that will determine whether AI companies will be permitted to claim training is fair use, including New York Times v. OpenAI, Authors Guild v. OpenAI, Andersen v. Stability AI, New York Times v. Perplexity, and UMG v. Anthropic. The Data Provenance Initiative's 2024 audit found that 70% of popular AI datasets lack proper licensing documentation and 50% contain licensing errors. This "documentation crisis" suggests that some companies may have trained their AI on data without attention to attribution, creating retroactive liability risks. There are some indications that industry itself recognizes that unlimited scraping may not be legally sustainable, as companies are increasingly turning to licensing deals—OpenAI alone has partnered with several publishers including Axios, The Atlantic, Vox Media, Conde Nast, the Financial Times, Le Monde, and News Corp. While the terms of these deals are not public, the agreement with News Corp reportedly entails $250 million in payments from OpenAI to News Corp over a 5 year period.
Understanding fair use within the broader picture of data ownership has significant implications on, for instance, worker displacement, licensing norms, the cost of data, and competition in the AI industry. However, the story is muddled with complexities, as we discuss in this piece.
Notable Definitions
Copyright Protection
Copyright is a legal right that protects original works of authorship. A work is protected as copyright under US Copyright Law if it possesses two characteristics:
- Original: The work must be “independently created by the author… [with at least] some minimal degree of creativity” (Feist Publications v. Rural Telephone Service Company 1991, p. 345).
- Fixed: It must be in a form that can be “perceived, reproduced, or otherwise communicated for a period of more than transitory duration” (Copyright Law of the United States 1990, 17 U.S.C. § 101).
The owner of copyright for a particular work is typically granted exclusive legal rights to the distribution and reproduction of the work for a fixed duration of time.² It protects against unauthorized copying, adaptation, public performance, and display of the work. Such “derivative works” include translations, dramatizations, abridgments, or “any other form in which a work may be recast, transformed, or adapted.” ³
Copyright applies broadly across creative fields, covering written works, music, film, software, and visual art, among others.⁴ Protection is automatic upon creation of the work and does not require formal registration (though registration is a necessary prerequisite to pursuing legal remedies in the event of infringement).⁵⁶⁷ It is one of several complementary intellectual property frameworks, working alongside patents, trademarks, and trade secrets. Once the copyright term on a work expires, the work enters the public domain and may be freely used by anyone.⁸
The Fair Use Exception
There are several exceptions under which copyrighted works can be used without infringement, one of which is known as “fair use.” There are four statutory factors that judges weigh when determining fair use:⁹
- “The purpose and character of the [secondary] use, including whether such use is of a commercial nature or is for nonprofit educational purposes.”
- “The nature of the copyrighted work.”
- “The amount and substantiality of the portion used in relation to the copyrighted work as a whole.”
- “The effect of the [secondary] use upon the potential market for or value of the copyrighted work.” The court weighs all four factors together when considering whether “fair use” applies. There is no fixed rubric for what qualifies as fair use, but past cases indicate how these four factors are weighed by the courts.
The first factor is commonly referred to as the “transformative use” factor. This asks how and why the secondary user employed the original work, including whether the use is commercial. Secondary works that add sufficiently new expression, meaning, or message to the copyrighted work are considered transformative. As the Supreme Court decided in Campbell v. Acuff-Rose Music, the more transformative a use is, “the less will be the significance of other factors, like commercialism, that may weigh against a finding of fair use.”¹⁰ This factor is highly debated, largely due to inherent subjectivity.
The second factor (“nature” of the copyrighted/original work) determines where a work falls on the “creativity” spectrum. As copyright is intended to cover creative expression, factual works receive comparatively less protection.¹¹
The third factor (the amount copied) considers both the quantity and quality of what was taken, and whether what was taken is the heart of the work. Smaller amounts are more likely to qualify as fair use. While quantity might be measurable, quality is often difficult to characterize.
The fourth factor, the “market effect” factor, examines whether the secondary use harms existing or potential markets for (i.e. value of) the copyrighted work. This factor is highly debated, as it is generally difficult to measure. For instance, it might require projections into the future, but such projections are inherently uncertain and therefore “attackable.” The later section Weighing the Four Factors for Generative AI Training discusses the application of speculative market effect projections within current fair use arguments in AI cases.
A detailed analysis of how each of the factors apply to generative AI is offered in section 4. Additionally, the history of pivotal cases as they relate to the four factors is discussed in section 5.
Analysis and Synthesis
Modern AI relies on massive amounts of training data. Language models, for instance, are trained on the equivalent of nearly a billion books of text. AI companies have two ways of legally obtaining this training data: (1) arranging to license the data from the data owners or (2) accessing the data through a legal exception to copyright such as fair use.
The training data underlying the most well known models was not licensed, as the data was scraped online. When copyright owners assert infringement, they must first show that their specific works were included in the training data without being licensed.¹² Only after this step does the burden of proof fall on the AI companies to claim that it is fair use.
Importantly, fair use and licensing can intersect in complicated ways: even if AI training is ruled as fair use, a license that explicitly prohibits AI training may still render the use illegal. This highlights the difference between copyright law (which provides the fair use legal exception) and contract law (which governs agreements such as licenses). For example, a creator may require anyone on their site to click “I Agree” before accessing any scrapable content, thereby entering them into a binding contract.
Copyright owners are leveraging contract law to strengthen licensing language, and AI companies are increasingly licensing content to avoid legal scrutiny. However, many of the current legal battles are about scraping that occurred prior to 2022, as creators were not aware of large-scale AI training and did not include explicit language around it. Even if such language is added now, there is a sense that the “damage has been done” since many AI companies have already trained their models on copyrighted data and used them to amass large user bases, revenues and investments.
The Two Outcomes (in Broad Strokes)
Broadly, rulings on fair use will result in one of two outcomes.¹³
- If AI training is ruled fair use, then scrapable data can be trained on as long as a valid license does not prohibit it. Creators may respond by paywalling their content, testing for bots, or removing their content altogether, which has implications for the openness of the internet. The remaining unblocked content is likely to be low-quality, so companies may turn to licensing deals with large, centralized rights‑holding platforms (like Sony or the New York Times) who can provide premium content bundles. Individual creators may be left behind because they have little negotiating power. The cost of licensing may drive smaller AI labs out of the AI industry, the remaining large AI companies may become further entrenched, and the cost of licensing may be passed down to the users. Creative workers will become increasingly displaced due to lack of attribution and compensation. Although this outcome may temporarily prevent companies from moving data operations out of the US, the lack of high-quality data may force this outcome in the long term.
- If AI training is ruled as not fair use, then licensing becomes mandatory. The price of high-quality data will likely still increase, but there will be a broader spectrum of low- to high-quality data that AI companies can purchase, which may provide consumers with different tiers of AI that are priced accordingly. Some companies may outsource training to other countries that are more permissive. AI companies will likely license mostly from groups with negotiating power (like media platforms and unions) rather than individuals. However, data markets for smaller creators may emerge, creating questions around data valuation. The online information ecosystem will likely remain relatively open, as there will be clearer separation between accessing for AI training and for other purposes. Thus, diverse and high-quality will likely remain online, though protected by licenses. Challenges will emerge around enforcement and auditing of compliance. Creative workers will still face the risk of displacement; much of this is determined by factors beyond fair use, such as unionization and negotiated worker protections.
Note that, regardless of rulings, several long-term implications depend on rewriting the rulebooks of licensing to cope with a new collective reality. In that sense, the presence of AI training has changed the dynamics of the creative industry in permanent ways.
Detailed Implications and Open Questions
We provide more detailed discussions of the implications and open questions below.
Creative Economy and Labor
- Worker Displacement and Future of Work
There is broad concern that generative AI will displace creative workers. Without exclusive rights, creative workers may not be credited or attributed for their contributions and instead find themselves replaced by AI. If creative labor becomes “up for grabs”, then there are few incentives for creative workers to produce. - Data Valuation and Markets
Creators may be able to claim compensation, e.g., through licensing. Already, different compensation schemes for training data are being tested and debated. Some seek to compensate creators with a flat, one-time payment. Others seek to pay creators based on “how much” the creator’s data improves the AI system (an outcome-based approach). Each scheme comes with implementation hurdles, including how to measure the value of individual pieces of data. The data market that eventually emerges will depend on these fair use rulings, which will in turn determine creator incentives and bargaining power. - Unions and Collective Bargaining
Labor unions such as WGA and SAG-AFTRA have begun using collective bargaining to compel employers to agree to AI guardrails, such as obtaining mandatory consent for training on worker outputs. The idea of leveraging data trusts is also gaining prominence: individual creators will not have much leverage to negotiate fair compensation. By pooling IP, data trusts could negotiate bulk deals and distribute royalties.
Transforming the Information Ecosystem
- Openness and Paywalls Online
If creative works that can be easily scraped online may be trained on without attribution or compensation, creative workers may become discouraged from publicly sharing their works (e.g., art, photography, writing portfolios) online. Such a change would have far-reaching consequences, including:- Information could become increasingly hidden behind paywalls or tests that certify a user is human, undoing efforts to create an open information ecosystem.
- Work that requires significant investment (e.g., investigative journalism) would be disincentivized due to (i) a lack of audience due to paywalls or (ii) a lack of pay if AI is permitted to provide the same information to users without compensating the original copyright owner.¹⁴
- Low-Quality Content: Slop
A potential consequence of more paywalls or restrictive licensing is that the only information that remains freely accessible is cheaply produced—for instance, AI “slop” created by generative AI itself. It may become more difficult to obtain high-quality content, which would reduce the diversity of available training data. - Monocultures
If models can only train on licensed data, they may over-index on content owned by large corporations with the capacity and legal teams to strike deals. Independent creators, smaller communities, and developing nations without centralized licensing bodies may be excluded, creating unidimensional AI systems that tend to regurgitate mainstream cultural norms and dominant attitudes. Further, as we move closer to a world where users are being fed daily information and advertisements by AI, the data that gets selected to reach the end user directly influences engagement and advertising, which in turn can shape mass opinions and beliefs.
Competition, Corporate Power, and Consumers
- Market Entrenchment, the Data Moat, and Network Effects
Stricter licensing is inevitable. However, the cost of data will likely be higher if training is ruled fair use. Large AI providers may become further entrenched while smaller companies face barriers to entry because they cannot afford licensed data. This asymmetry would reward large companies for behaving illegally (as they already own powerful models trained on copyrighted data), thereby encouraging monopoly. Another notable handicap for small companies would be the upward battle against network effects. It would further handicap small companies that already face an upward battle against network effects.¹⁵ - Media and Platform Power
Legacy media companies like Sony and Disney will become the gatekeepers of high-quality cultural content, particularly if training is not ruled as fair use. This gatekeeping ironically does not prevent generative AI from replacing creators, as studios and other media companies are actively deploying their own internal AI systems to accelerate their own media production. - Jurisdictional Shopping
Copyright is territorial. If US courts rule against fair use, AI companies may move their data mining operations to countries that have fewer restrictions (e.g., Japan). - Passing Costs to the Consumer
Data licensing costs may be passed to the consumer. High-quality AI could become gated behind expensive subscriptions, creating uneven access to AI.
Legal Implications
- Enforcement Challenges and Data Tracing
There will be challenges around auditing and enforcing licenses, especially if AI training is not fair use. One obstacle is data tracing: proving beyond doubt that an AI system was trained on a specific piece of data, which remains largely unsolved. - Licensing
There are also ongoing efforts to create new licenses with explicit provisions on AI use. New popular options include Responsible AI Licenses (RAIL) and the Montreal Data License, which allow owners to define exactly how their data can be used for machine learning. These frameworks were designed to bypass the present legal gray area of fair use, replacing unpredictable, case-by-case legal defenses with explicit rules for AI developers.
Weighing the Four Factors for Generative AI Training
As detailed above, fair use analysis under Section 107 of the Copyright Act employs a four-factor test.¹⁶ Determinations are highly fact-specific and require a case-by-case analysis.
In discussions, factors one and four have received the most attention. Assessing factor one (purpose and character of usage) requires an examination of whether generative AI is “transformative.” Early rulings indicate that courts are receptive to the argument that training AI to learn statistical patterns differs from the expressive purpose of original works, suggesting transformation. Factors weighing in favor of a fair use ruling include whether the use is for public-interest or non-commercial purposes and the degree to which the secondary work is transformed. Still, factor four (market effect) has, in recent AI cases, has proved to be the most promising avenue for copyrightholders hoping to protect their works. Courts may consider the following arguments: lost sales through direct substitution; market dilution, where AI-generated content floods markets and thereby reduces the value of originals; and lost licensing opportunities where viable licensing markets exist. We discuss cases that have employed some of these arguments and their degree of success in Section 5.
Factors two and three receive comparatively less attention. As long as the copyrighted work is sufficiently “creative”, factor two (nature of the original work) is taken as a given. Factor three (the amount used) examines how much is copied. Generally, the fact that AI training occurs on an entirely copied work counts against fair use.
2025 U.S. Copyright Office Report
In May 2025, the United States Copyright Office released a report addressing AI and copyright law, focusing specifically on whether the unauthorized use of copyrighted materials to train generative AI systems is defensible as fair use.¹⁷ While the report is “non-binding,” it offers a window into how the Copyright Office is thinking about the application of fair use to generative AI. The report concluded that multiple acts in AI training constitute prima facie copyright infringement: data collection and curation involve reproduction, the training process itself requires copying works to high-performance storage, and model weights that memorize training data potentially infringe reproduction and derivative work rights. Notably, the Office rejected the "human learning" analogy that AI companies frequently invoke, observing that humans retain imperfect impressions while AI systems create perfect digital copies. It also emphasized that using works for training is not automatically non-expressive and rejected fair use arguments that training is "inherently transformative." In short, there is potential but not certainty for infringement.
Initial Rulings: Anthropic & Meta Cases
Two landmark district court decisions in June 2025 provided initial judicial guidance but reached somewhat divergent conclusions. In both Bartz v. Anthropic and Kadrey v. Meta, the courts largely ruled in favor of the AI companies, but differed on their openness to the particulars of factor 1 and 4. Neither court accepted the plaintiffs’ argument that the existence of a potential market for licensing works as AI training data constituted the kind of market harm relevant to factor 4.
- Transformative Use: In Bartz v. Anthropic, the Northern District of California ruled that training on lawfully acquired books constitutes "quintessentially transformative" fair use, finding that the purpose of learning statistical relationships differs fundamentally from the expressive, human-readable purpose of the original texts. However, the same court held that Anthropic faces liability for downloading and retaining a digital library of pirated works, establishing that the source of data acquisition matters significantly for fair use analysis.¹⁸ In Kadrey v. Meta, Meta won summary judgment on fair use for reproduction/training for the 13 named plaintiffs (all published authors), but the judge stressed that the ruling was limited in scope to only those 13 authors.¹⁹
- Market Dilution: The Bartz court rejected market dilution arguments, finding that AI outputs don't substitute for books even when addressing similar topics. The Kadrey court showed more receptivity to these theories, though ultimately ruled for Meta based on insufficient evidence of harm, but limited the ruling to the Meta case. Still, the Kadrey court appeared open to a market dilution theory: if plaintiffs could provide more direct evidence that AI outputs substitute for their works and erode demand, such a claim could be viable. Additionally, the ruling noted that some types of works are more likely to support a strong market dilution argument than others–notably, news articles and print journalism, which may be replaced by LLM-generated summaries or reporting. This suggests that nonfiction genres like biographies, instructional writing, or how-to guides, where the content can be more easily mimicked may be especially vulnerable. By contrast, memoir and literary fiction, where author identity, voice, and brand matter more, may be less likely to face the same kind of substitutional pressure.
- Potential Licensing Markets: Past disruptive technologies have often been challenged on the basis of alleged market harm through lost revenue or the lost ability to license. However, both courts were unpersuaded by this argument. Rights holders argue they're losing licensing revenue because AI companies aren't paying to use their data for training. But as the Kadrey court pointed out, this reasoning is circular: the only reason a 'training data licensing market' exists at all is because LLMs created demand for it. You can't claim lost revenue in a market that wouldn't exist without the technology you're suing over.
These decisions establish important principles but remain limited: they are district court rulings subject to appeal and not binding on other jurisdictions. The holdings turn on the distinction between lawfully purchased books and pirated sources, with Anthropic later agreeing to a $1.5 billion dollar settlement on the basis of the ruling. Meta, by contrast, obtained a favorable fair-use ruling in that case, though the decision was explicitly narrow and fact-specific. More broadly, the two opinions converge on the view that LLM training is transformative, but disagree on how that finding interacts with commerciality, asserted market substitution, and ‘lost licensing’ theories under factor four.
Related Context and Cases
Many scholars argue that the fair use doctrine formally emerges with Folsom v. Marsh (1841), in which Justice Story articulated market harm and purpose as central considerations in determining infringement. The irony, one intellectual property scholar notes, is that what is now celebrated as a safety valve originally expanded the scope of copyright:
“Formerly, infringement was limited to near-verbatim reproduction and all other subsequent uses were considered legitimate. In the new fair use environment, all subsequent uses became presumptively infringing unless found to be fair use” .
In the next century, courts grappled with situations where copying produced no clearly demonstrable economic harm of copying, especially when such copying served a clear public benefit. Two cases illustrate the courts’ initial reluctance to treat potential licensing markets as evidence of harm when those markets were speculative. In Williams & Wilkins Co. v. United States (1973), the medical publisher W&W sued the federal government. The federal agencies National Institute of Health (NIH) and National Library of Medicine (NLM) had photocopied and distributed W&W medical articles to researchers. The majority emphasized the absence of demonstrated harm, writing that it was “very important that [the] plaintiff has failed to prove its assumption of economic detriment, in the past or potentially for the future” (W&W Co. v. U.S.). And, if the purpose of copyright is to spur innovation and serve the public good, surely increasing access to research is moving this initial purpose forward. Twenty years later, Sony’s Supreme Court case reaffirmed the speculative harm principle, concluding that the movie studios’ claims of future loss were unsubstantiated and therefore insufficient to defeat fair use (Sony Corp. of America v. Universal City Studios, Inc). The majority opinion noted that the expansion of broadcast availability was a clear public benefit.
Later cases would clarify the importance of analyzing the secondary work relative to the original work, particularly in understanding whether the secondary work can substitute in part or, in whole, the original in the market. Courts thus developed a more positive position on potential licensing markets, particularly as they emphasized the risk of “market substitution.” In Harper & Row v. Nation (1985) a magazine published an unauthorized excerpt of Ford’s unpublished memoir. The Court emphasized that the secondary use was unfair because people will buy the new product over the original, usurping the author’s right of first publication and the market value of the initial work. Harper & Row thus recentered “market substitution” as a decisive concern. Less than a decade later, however, Campbell v. Acuff-Rose Music (1994) reframed factor one of the fair use analysis around “transformative use,” asking whether the secondary work merely supersedes the original or instead adds something new with a further purpose and/or different character.
Yet interpretations of fair use continue to evolve. Ten years after the W&W case, a class action was brought against Texaco for systematically photocopying journal articles, prompting courts to clarify again whether harm must be shown and whether potential licensing markets count as cognizable economic loss. American Geophysical Union v. Texaco Inc became the hinge case that recognized “traditional, reasonable, or likely to be developed markets,” further evolving the fair-use-harm doctrine. Subsequent cases would go on to test the licensing market argument. Mass digitization cases like Authors Guild v. HathiTrust and Authors Guild v. Google showed courts upholding large-scale copying to enable search and accessibility.
The 2021 Google v. Oracle further indicated where fair use doctrine is heading: Oracle sued Google for copying its Java APIs, and arguments revolved around the copyrightability and licensability of software interfaces. A mere few years later, a slew of lawsuits would emerge as generative AI models offer a more direct analogy of large bodies of copyrighted data feeding an algorithm that produces new outputs derived from that same corpus.
Footnotes
A token is defined as a small chunk of text, roughly equivalent to a word or part of a word, that serves as the basic unit a model processes.
U.S. Copyright law typically protects works for duration of the life of the author plus another 70 years (Copyright Law of the United States 1990, 17 U.S.C. § 302)
Copyright Law of the United States 1990, 17 U.S.C. § 101
Copyright Law of the United States 1990, 17 U.S.C. § 411.
Copyright Law of the United States 1990, 17 U.S.C. § 410.
Copyright Law of the United States 1990, 17 U.S.C. § 412.
U.S. Copyright Office. “Frequently Asked Questions.” Accessed March 4, 2026. https://www.[copyright](/terms/copyright?internal=true).gov/help/faq/.
Copyright Law of the United States 1990 17 U.S.C. § 107.
Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569 (1994).
Courts also prefer to give the original creator the right to first publication, so whether a piece is published is also considered as part of the second factor. If the work is not yet published, the secondary use is unlikely to be fair use.
This step creates a significant hurdle. Although training on scraped data is an open secret, showing that a specific work is included in training is not only technically challenging, but sometimes impossible. The New York Times, in their lawsuit against OpenAI, provided circumstantial proof by prompting GPT‑4 and recording instances of it regurgitating near-verbatim snippets of their articles.
Although neither outcome is likely to retroactively protect creators whose data has already been used.
In this context, network effects refer to the phenomenon where AI companies with loyal user bases enjoy a steady stream of training data supplied by users. For example, each time a user chats with ChatGPT or Gemini, (i) they inherently provide OpenAI and Google, respectively, with information on user queries, and (ii) users often give some indication of whether the chatbot response is satisfactory, which serves as training data.
Copyright Law of the United States 1990 17 U.S.C. § 107.
U.S. Copyright Office. Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-publication version). Report of the Register of Copyrights. May 2025. https://www.[copyright](/terms/copyright?internal=true).gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf.
Authors Guild. “Mixed Decision in Anthropic AI Case: Authors Guild Responds to Summary Judgment in Bartz v. Anthropic.” June 25, 2025. https://authorsguild.org/news/mixed-decision-in-anthropic-ai-case/.
Kadrey v. Meta Platforms, Inc. “Order Denying the Plaintiffs’ Motion for Partial Summary Judgment and Granting Meta’s Cross-Motion for Partial Summary Judgment.” Case No. 23-cv-03417-VC, Doc. 598. U.S. District Court for the Northern District of California, June 25, 2025