The Memo: The cost of using pirated material? Anthropic faces a record-breaking $1.5 billion copyright settlement

the memo.png

The cost of using pirated material? Anthropic faces a record-breaking $1.5 billion copyright settlement

Written by Erin Bradbury - 13 August 2026

Bloomsbury, one of the world's largest publishers, is set to receive a share of a record-breaking $1.5 billion settlement from Anthropic, a California-based AI developer, over its retention of pirated works. Ever since AI-assisted technologies, such as chatbots like Claude, became widely available questions have been raised about how these systems are developed. Large language models (LLMs) are trained on vast quantities of text and copyright has remained a central issue in the debate surrounding AI.

In 2024 a class-action lawsuit was filed in California under Bartz v Anthropic by three authors, Andrea Bartz, Charles Graeber and Kirk Wallace Johnson. The plaintiffs alleged that Anthropic used copyrighted materials without the authors’ permission. Their claims also centred on the approximately seven million books which were downloaded and retained from pirate libraries as part of Anthropic’s internal data base building efforts.  

Anthropic argued that its use of the plaintiffs’ books constituted ‘fair use’ and was therefore non-infringing. They also stated that these materials were not used in any commercial models. The judge agreed that using books to train AI models didn’t violate US copyright law, stating: “Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works, not to race ahead and replicate or supplant them, but to turn a hard corner and create something different […] If this training process reasonably required making copies within the LLM or otherwise, those copies were engaged in a transformative use.”

Under US copyright law, courts typically assess fair use using four statutory factors:

  1. the purpose and character of the use, including whether such use is of commercial nature or is for nonprofit educational purposes;
  2. the nature of the copyrighted work;
  3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
  4. the effect of the use upon the potential market for or value of the copyrighted work.

This case highlights that the court will not extend this protection to materials acquired from pirated libraries. Though Anthropic was partially successful in in their argument, the court still ordered them to pay approximately $3,000 per covered book as the case revealed that Anthropic later shifted towards purchasing physical books and digitising them at scale to develop its training library.

The settlement covers more than 482,000 books, around 91% of which have been claimed by authors and publishers. 350 class members opted out, preserving their right to bring separate claims against Anthropic, several of whom have already done so. In addition, the company must destroy any infringing data.

Anthropic’s partial success hinged on its argument for fair use for training its model, and it was treated differently due to how its material was obtained from pirated libraries. In a similar case, Meta was ruled in favour of as the judge focused on whether or not the company had harmed the market for the author’s work, indicating if there was enough evidence a different ruling might have been made.

Currently there are over 130 active cases relating to copyright and AI in the US. Numerous cases involving text, images, visual art, video and music remain ongoing, with plaintiffs including Getty Images and The New York Times bringing claims against technology companies such as Google, OpenAI and Microsoft.