Introduction to Meta Defeats US Copyright Claim
Some of you may have been following the occasional article on AI and copyright, so I noted with interest a case with significant implications for generative artificial intelligence (AI) and digital copyright law. Meta Platforms has successfully defeated a group of US authors’ claims brought under the Digital Millennium Copyright Act (DMCA). US District Judge Vince Chhabria granted Meta’s motion for partial summary judgement in the Northern District of California, finding that Meta’s use of the authors’ works was legal fair use.
https://law.justia.com/cases/federal/district-courts/california/candce/3:2023cv03417/415175/280/
The ruling not only absolves Meta from liability under the DMCA but also signals how courts may approach similar disputes concerning AI training data in the future. Below, we explore the background to the case, the key legal arguments, Judge Chhabria’s analysis, and the wider ramifications of this decision.

Background to the Dispute
At the heart of the case lies Meta’s use of copyrighted books and literary works to train its large language models. Like other major tech firms working on generative AI, Meta compiled vast datasets, which included works by numerous contemporary authors. These datasets were used to “teach” AI systems how to generate human-like text, answer questions, and perform other sophisticated language tasks.
A group of authors sued Meta, alleging that the company infringed their copyrights by reproducing and using their works without permission. In addition to direct infringement claims, the authors alleged violations of the DMCA’s provision against the removal or alteration of copyright management information (CMI).
According to the DMCA claim, Meta violated Section 1202(b) of the DMCA, which forbids the wilful removal of CMI with knowledge that it will facilitate infringement, by deleting or neglecting to include authorship or copyright metadata when incorporating the works into its training data.
The Authors’ Arguments
The authors contended that their works had been copied wholesale into Meta’s training datasets without permission and stripped of any metadata that might have identified them as copyrighted.
They argued that this removal of CMI was a crucial step that allowed Meta’s AI systems to process and reproduce their works without attribution or licensing, thereby directly facilitating infringement.
In their view, this conduct was emblematic of a wider problem in the AI industry: that large-scale language models rely on massive, indiscriminate ingestion of copyrighted material to function effectively, often at the expense of creators’ rights.
Meta’s Defense
Meta vigorously contested the allegations, asserting that its use of the authors’ works fell squarely within the doctrine of fair use.
According to 17 U.S.C. § 107, fair use permits some unauthorised uses of copyrighted works for teaching, research, criticism, commentary, news reporting, and scholarship. Four non-exclusive factors are examined by courts when evaluating fair use:
- The purpose and character of the use, including whether it is commercial and whether it is transformative.
- The nature of the copyrighted work.
- The amount and substantiality of the portion used.
- The effect of the use on the potential market for or value of the copyrighted work.
Because the works were used to teach AI systems to recognise linguistic patterns and generate unique outputs rather than to replicate or disseminate them to the general public, Meta claimed that its use of the works was extremely transformative.
Meta also maintained that the use had no negative effect on the market for the authors’ books; rather, it did not serve as a substitute for the original works and did not diminish their value.
The Court’s Analysis
In his opinion, Judge Chhabria sided with Meta, concluding that the use of the authors’ works was protected under fair use.
Transformative Purpose
The court’s reasoning was heavily reliant on the transformative nature of Meta’s use. The judge emphasised that Meta didn’t profit from the works by selling copies or republishing them. Rather than serving the authors’ original literary and expressive objectives, the books were analysed by algorithms to extract general principles of language and narrative structure.
This kind of change is consistent with earlier fair use rulings, like the historic ruling in Authors Guild v. Google, Inc. (2015), where it was decided that Google’s scanning of books to produce a searchable index qualified as fair use. In that instance, the Second Circuit ruled in favour of fair use because it believed that Google’s database had a different function than the books’ original expressive use.
Amount Used
Although Meta copied entire works, the court noted that copying entire works can still be fair use if the copying is necessary to achieve the transformative purpose. For example, in the Google Books case, full works were scanned, but that copying was deemed permissible because partial copying would not suffice to enable comprehensive indexing and search functions.
Analysing the entire texts was also necessary in Meta’s case in order to properly teach the AI about comprehensive language usage and structure.
Market Effect
The court rejected the argument that Meta’s use harmed the market for the authors’ works. There was no evidence that AI outputs from Meta’s systems served as a substitute for the original books or reduced their sales.
Instead of directly reproducing or disseminating copies of the original works, Meta’s models create new text by identifying patterns in a large dataset.
The court came to the conclusion that Meta would not be disadvantaged by the fair use analysis because there was no discernible market harm.
Implications for the DMCA Claim
Having found that Meta’s use of the works was not infringing, Judge Chhabria addressed the DMCA claim.
Under Section 1202(b), liability requires that the removal of CMI was done knowingly and that it facilitated or concealed an act of infringement. Here, because the underlying act — copying for AI training — was deemed non-infringing under fair use, the second requirement of Section 1202(b) was not satisfied.
Judge Chhabria succinctly concluded:
“Because Meta’s copying was not an infringement, its removal of [content management information] could not have furthered an act of infringement.”
The DM was successfully disposed of by this discovery.
Impact on AI and Copyright Law
This ruling represents a significant victory for technology companies developing large language models and other generative AI systems.
By reaffirming that large-scale text ingestion for training purposes can qualify as fair use, the decision provides a crucial legal foundation for AI research and development in the United States.
The decision is not without controversy, though. Allowing tech companies to copy entire books without permission, according to authors and copyright advocates, devalues creative works and discourages original authorship. They worry that these rulings will give tech companies broad immunity at the expense of the livelihoods of creators.
On the other hand, AI developers contend that fair use is essential to innovation and free expression. They argue that preventing access to large corpora of text would cripple AI development and ultimately harm the public, which benefits from the advances in natural language processing, accessibility tools, and educational applications.
Broader Legal Context
The Meta decision joins a growing body of case law addressing the tension between AI training and copyright protection. Courts have consistently leaned towards allowing data ingestion for transformative, non-expressive purposes.
In addition to Authors Guild v. Google, other relevant precedents include:
- Perfect 10, Inc. v. Amazon.com, Inc. (2007), where the Ninth Circuit held that using images as thumbnails in a search engine was a transformative fair use.
- Kelly v. Arriba Soft Corp. (2003), where copying images for indexing and search was deemed fair use.
However, each fair use determination is fact-specific, and future cases may turn on different factual records or technological contexts.
Moreover, legislative developments could impact the legal landscape. There have been increasing calls for clearer statutory frameworks governing AI and copyright, including proposals for opt-out mechanisms for authors and more transparent data provenance requirements.
What Happens Next?
The plaintiffs may still file an appeal with the Ninth Circuit, even though this decision represents Meta’s clear victory at the summary judgement stage. An appeal would be expected given the case’s high profile and its potential to influence future AI copyright law.
The ruling will probably encourage other tech firms to use copyrighted works to train AI models in the interim. Businesses that are facing comparable lawsuits, like OpenAI, Google, and Anthropic, will surely use Judge Chhabria’s logic to support their positions.
Key Takeaways for Authors and the Tech Industry
Fair Use Expanded for AI: Courts still interpret fair use broadly to support non-commercial, transformative analytical uses, particularly when doing so makes new technological functions possible, such as AI training.
DMCA Claims Face Higher Hurdles: Without underlying infringement, claims under the DMCA’s CMI provisions are unlikely to succeed.
Need for Legislative Clarity: The decision highlights the pressing need for lawmakers to revisit copyright laws in light of emerging AI technologies to balance creators’ rights and technological innovation.
Uncertainty for Authors: Authors and publishers may need to explore new avenues to protect their works, such as licensing schemes, collective rights management, or advocating for opt-out registries for AI training datasets.
Judicial Scepticism of Market Harm: Courts appear to require concrete evidence of actual market substitution or harm to weigh against fair use, which can be challenging for plaintiffs to demonstrate in the context of AI
UK Law? No “Fair Use” — Only “Fair Dealing”
In the US, Meta succeeded primarily by relying on the broad doctrine of fair use under § 107 of the US Copyright Act. Because of its flexibility and open-endedness, this doctrine enables courts to weigh four factors and take into account any purpose they believe to be “transformative”, including AI training.
The general fair use doctrine is absent from UK copyright law, in contrast. Instead, it offers a series of precise and limited exceptions known as “fair dealing”, which are thoroughly outlined in laws (most notably the Copyright, Designs and Patents Act 1988, or CDPA).
UK fair dealing exceptions cover purposes such as:
Research and private study.
Criticism or review.
Reporting current events.
Quotation.
Parody, caricature, and pastiche.
There is no general catch-all “transformative use” defence in the UK.
Could the Use of Meta Fit Under UK Fair Dealing?
The most likely exception that Meta could attempt to use is the Section 29A CDPA, which permits text and data mining for non-commercial research.
However, there are stringent requirements for that exception:
Only pertains to research that is not for profit.
Users must be able to access the works legally (for example, by obtaining a licence or subscription).
The use must be limited to information gathering through computational analysis.
Meta AI training is usually a business endeavour, with the goal of creating goods and services that can be used for profit. This would typically not be considered “non-commercial research” under UK law. So I don’t feel this has any legs whatsoever. Furthermore, copying entire works for large-scale model training that results in commercial outputs, like generative AI chatbots, is prohibited by section 29A.
In summary
In the developing relationship between AI and copyright law, Judge Chhabria’s ruling represents a turning point in US law. The UK courts have not yet been tasked with any such questions, but it is only a matter of time. The authors’ concerns regarding the unlicensed use of their works are still valid and urgent, and I wish the UK government would take notice.


