
By Chuck Gallagher — Business Ethics Keynote Speaker and Trainer
TL;DR: Chuck Gallagher, AI ethics speaker and author, breaks down the $1.5 billion Anthropic copyright settlement granted final approval on July 20, 2026, and argues that where training data came from is not a paperwork problem — it is a product risk courts have now put a price on.
There is a moment in every company that nobody writes down. Somebody asks a question, and somebody else waves it off.
Where did these books come from?
We’ll sort that out later.
Nobody schedules that meeting. Nobody takes minutes. But some version of that exchange is where most of the ethics cases I study begin. Not with a villain. With a shrug.
What Did the Court Actually Approve?
On July 20, 2026, Judge Araceli Martinez-Olguin of the U.S. District Court for the Northern District of California granted final approval of a $1.5 billion class action settlement in Bartz v. Anthropic and entered final judgment. The case was filed in August 2024 by three writers — novelists Andrea Bartz and Charles Graeber, and nonfiction author Kirk Wallace Johnson — on behalf of a class covering roughly half a million works.
The court found the deal delivered substantial benefits given the novelty of the claims. Per-work payments land somewhere around $3,000. The order points out that figure is four times the $750 statutory minimum for ordinary infringement, which also happens to be the most common award in copyright cases.
Read that again. Three thousand dollars a book. Times half a million books.
Why Did “We’ll Sort Permissions Later” Cost $1.5 Billion?
Here is the part people keep getting wrong.
In June 2025, Judge William Alsup ruled that training an AI model on lawfully purchased books was fair use. He called it exceedingly transformative. That was a genuine win for the industry, and it is still the law today.
Then he drew a line. Downloading millions of titles from shadow libraries — Library Genesis and Pirate Library Mirror — and holding them in a permanent central library was not fair use. Court filings put that haul at more than seven million books.
So the technology was not the problem. The training was not the problem. The acquisition was the problem.
As an AI ethics speaker and author, I watch organizations pour money into model safety while treating data sourcing as a procurement footnote. That is backwards. Your model can be perfectly defensible and the pipeline that fed it can still take you to the mat. Statutory damages run as high as $150,000 per work. Do that arithmetic against half a million works and you understand why $1.5 billion started to look like the cheap door.
What Does the Fee Cut Say About Who the Court Was Watching Out For?
Something happened in this order that deserves more attention than it is getting.
Class counsel originally asked for 20 percent of the fund. They trimmed themselves back to 12.5 percent, or $187.5 million. The court cut them again — down to roughly 6.8 percent of the fund, about $101.56 million — and then held back 10 percent of even that until counsel files a post-distribution accounting, with room to cut further if the actual hours come in under projection. Service awards for the three named plaintiffs dropped from $50,000 each to $15,000.
Every dollar not paid in fees stays in the fund. The fund is non-reversionary. Anthropic does not get it back. It goes to authors.
That is a judge watching the money.
What Did This Settlement Not Buy?
This is where the discipline matters, and where I have had to correct more than one confident executive summary.
Class members released only claims tied to past acquisition and copying — the inputs side — through August 25, 2025. Claims about model outputs are not released. Claims about future conduct are not released. Works that never made the Works List release nothing at all. The court said so in plain language: a narrow release benefits the class.
Beyond the money, the settlement requires destruction of the original pirated files and any copies traced to them, subject to legal preservation duties. Anthropic has represented that neither dataset, nor any portion of either one, was in the training corpus of any of its commercially released large language models.
Let me be clear about what that means. A company agreed to pay $1.5 billion over books it says never made it into a shipped product. The getting and the keeping were the exposure. Not the output.
Where Does This Leave Everybody Else?
Every company deploying AI right now is running some smaller version of the same play. Scraped data. Vendor datasets nobody audited. A model fine-tuned on a folder somebody pulled from somewhere and never sourced. When organizations bring me in as an AI ethics speaker and author, this is the conversation nobody scheduled and everybody needed.
The rationalization always sounds reasonable inside the room. Everybody needs data. Training is transformative. We will clean it up before launch.
I know that voice. Mine sounded reasonable too. It’s just a loan. I’ll pay it back. That is the whole trick — the story you tell yourself that makes a wrong thing feel temporary.
Frankly, the fix is not complicated. It is just unglamorous. A written data sourcing policy. An auditable chain of custody for every dataset that touches a model. Named no-go acquisition channels, put in writing, so nobody has to interpret. And a model risk sign-off where the legality of training data is a launch-blocking control, not a legal memo that shows up after the press release. An ISO/IEC 42001-style AI management system gives you the scaffolding for all of it. It will not give you the will.
What Is the Lesson Here?
Inputs are not paperwork. Inputs are a product’s origin story, and origin stories eventually get told out loud, under oath, with a docket number attached.
You know exactly what I mean. Somewhere in your organization there is a dataset nobody can trace. And somebody knows. Somebody always knows.
Every choice has a consequence. Choosing not to ask where the data came from is still a choice. The consequence just takes a few years to show up.
Frequently Asked Questions
What did the Anthropic copyright settlement actually cover?
The $1.5 billion settlement in Bartz v. Anthropic, granted final approval on July 20, 2026, resolves claims about Anthropic’s past acquisition and copying of books obtained from pirate sites, through August 25, 2025. It does not release claims based on AI outputs, and it does not release any claims about future conduct. Works that were not on the court-approved Works List are unaffected by the settlement entirely.
Did the court rule that training AI on books is illegal?
No. In June 2025, Judge William Alsup ruled that training an AI model on lawfully purchased books is fair use, describing the use as exceedingly transformative. What he did not excuse was downloading millions of books from shadow libraries such as Library Genesis and Pirate Library Mirror and keeping them in a permanent internal library. The settlement resolves that second category, not the first.
How much money do authors get from the Anthropic settlement?
Payments work out to roughly $3,000 per work, distributed pro rata among valid claimants and split between authors and publishers according to elected default splits or the terms of their publishing contracts. The court noted that amount is about four times the $750 statutory minimum for ordinary infringement. A court-appointed Special Master will resolve disputes between claimants over any given work.
What should companies do about AI training data provenance?
Treat data sourcing as a launch-blocking control rather than a legal afterthought. As an AI ethics speaker and author, I tell leadership teams to put four things in place: a documented sourcing policy, an auditable chain of custody for every dataset, named no-go acquisition channels, and a model risk sign-off that no one can waive quietly. An ISO/IEC 42001-style AI management system provides a recognized structure for that discipline.
Does this settlement protect Anthropic from future copyright claims?
It does not. The release is narrow and backward-looking, covering only past acquisition and copying through August 25, 2025. Claims about model outputs remain live, as do claims about anything the company does going forward. The court specifically contrasted this with the rejected Google Books settlement, which had tried to release future claims.
Bring This Conversation to Your Organization
Here is the uncomfortable truth for any leader reading this: the exposure in your organization is almost never the decision somebody agonized over. It is the one nobody bothered to write down. I spend my time in front of boards, compliance teams, and technology leaders helping them find those unwritten moments before a court does it for them — and Chuck Gallagher brings that work to conferences, leadership retreats, and internal ethics programs across industries. If your organization is deploying AI faster than it is documenting where the data came from, let’s have that conversation now rather than after the docket number. Learn more or start the conversation at ChuckGallagher.com.
Five Questions for Reflection
1. In your organization, who is actually accountable for knowing where a training dataset came from — and could that person produce the answer today, in writing?
2. The rationalization in this case was “we’ll sort permissions later.” What is the equivalent phrase people use in your company to defer a question they do not want to answer?
3. The court released only past acquisition claims, not output claims or future conduct. What ongoing risks does your team assume are resolved simply because something was settled once?
4. Anthropic represented the pirated files never entered a shipped model’s training corpus, and still paid $1.5 billion. What does that say about the difference between legal harm and reputational harm?
5. If a judge examined the provenance of every dataset your organization touched in the past twenty-four months, what would you want to fix before that review started — and what is stopping you from fixing it now?
