Anthropic and the billion-dollar receipt
https://sangtd.net/anthropic-and-the-billion-dollar-receipt/July 20, 2026, a federal judge in San Francisco signed the final order approving a $1.5 billion settlement between Anthropic and a class of authors whose books were used to train the company’s Claude models without permission. The approval was the last procedural step in a lawsuit that began in August 2024, when novelist Andrea Bartz and two other authors accused Anthropic of downloading nearly two hundred thousand pirated ebooks from the Books3 dataset — a collection distributed through BitTorrent — converting them to training data, and feeding them into a model that now generates billions of dollars in revenue. The final order triggers the payout: three thousand dollars per book, shared between authors and publishers, across an estimated half a million works. Ninety-one percent of the eligible books have been claimed. The National Writers Union, which represents many of the affected authors, had already called the settlement “not the settlement we want” when it was first proposed last year. The checks are now being cut, and the litigation is over. The settlement was a receipt. And the receipt cost $1.5 billion.
What was being purchased was not forgiveness. It was provenance. Before the settlement, the books in Claude’s training data had a clear origin: they were stolen. They came from pirate sites, downloaded in bulk, converted without permission, fed into a training pipeline without attribution. Every output Claude produced was downstream of that act. After the settlement, the origin story changed. The books were no longer stolen — they were acquired as part of a legally resolved dispute. The settlement did not undo the act of downloading from pirate sites. It did not restore the authors’ control over their work. It did not remove the pirated books from the training data. What it did was issue a receipt. Pay $1.5 billion, and the same training data that was illegal in August 2024 became a legitimate business expense by July 2026. The receipt transforms the narrative: the company did not steal — it settled. And settlement implies a transaction, not a crime.
This is not how penalties are supposed to work. A fine for theft is supposed to deter the thief and compensate the victim. It is not supposed to convey ownership of the stolen goods to the thief. If someone steals a car, pays a fine, and then receives a title deed to the car, the system has stopped being a justice system and started being a marketplace. But that is precisely what the $1.5 billion settlement achieves. Anthropic did not delete the pirated books from its training data. It did not retrain Claude from scratch on licensed content. It wrote a check, and the check purchased something more valuable than any individual book: the right to say that the question of where the data came from has been resolved. The settlement closed the question. The receipt was issued. The goods are now, for all practical and legal purposes, theirs.
Six months before the preliminary settlement, Judge Alsup had already issued a mixed ruling that shaped everything that followed. Training AI models on copyrighted books, he found, was fair use — a landmark decision, the first time a federal court had given credence to the industry’s claim that ingesting the entire written output of humanity qualified as transformative under a doctrine last updated in 1976. But Alsup drew a line at the method of acquisition. Downloading millions of books from pirate sites like Library Genesis and Pirate Library Mirror, he ruled, was illegal on its own terms, separate from the question of training. That question — the piracy question — was set for trial. A jury would hear evidence about how Anthropic obtained its training data. A verdict could establish binding precedent: not just that training on copyrighted works is fair use, but that acquiring those works through mass piracy is not. The settlement preempted the trial. No jury would hear the evidence. No appeals court would review the ruling. No precedent would be set. Anthropic’s deputy general counsel, Aparna Sridhar, issued a statement highlighting the fair use finding as a landmark — spinning a settlement that cost the company $1.5 billion as a vindication of its legal position. The other question, the one about the pirate sites, was buried with the check. The receipt covered everything.
The legal concept at work here is not complex. It is the same principle that governs every settlement in every industry: paying to make a problem go away. But the scale and the nature of the problem change the meaning of the act. When a pharmaceutical company settles a lawsuit over a defective drug, the settlement compensates victims but does not retroactively make the drug safe. When an oil company settles an environmental lawsuit, the settlement funds cleanup but does not make the spill unhappen. When an AI company settles a copyright lawsuit over training data, the settlement does something unique: it launders the provenance of the data itself. The data remains in the model. The model continues to generate revenue. The settlement is simply the price of wiping the data’s criminal record clean. The stolen books are now, in the eyes of the market and the law, legally acquired assets. And legally acquired assets can be defended.
This is where the story pivots from settlement to accusation. In February 2026, Anthropic joined OpenAI in publicly flagging what they called “distillation campaigns” by Chinese AI firms. Anthropic posted on its official account: “We have proof of distillation at scale by MiniMax, DeepSeek, and Moonshot.” OpenAI submitted a memo to the U.S. House Select Committee on China accusing DeepSeek of “free-riding on the capabilities developed by OpenAI and other U.S. frontier labs.” The U.S. State Department followed in April with a diplomatic cable instructing embassies worldwide to warn foreign governments about Chinese firms allegedly distilling American AI models, naming the same three companies. The language of theft, of stolen property, of illicit extraction, suddenly filled the policy discourse. The companies that had downloaded the entire written record of humanity without permission from pirate sites were now demanding that the world respect their intellectual property.
Distillation is a technical process. A large, powerful model — Claude, GPT-5, Gemini — generates outputs. A smaller model is trained on those outputs to replicate the larger model’s capabilities at a fraction of the cost. The outputs are not copied verbatim; the smaller model learns the patterns, the reasoning chains, the behavioral characteristics of the larger one. It is, in essence, the same process that produced the large model in the first place, shifted one layer up the stack. The large model was trained on human-written text — books, articles, forum posts, code. The small model is trained on large-model-generated text. The principle is identical: learn from existing outputs to produce new ones. The only difference is whose outputs are being learned from.
The AI industry’s position on distillation is that it is theft when done without permission. Using Claude’s API to generate millions of responses, then training a competitor model on those responses, violates the terms of service and constitutes unauthorized use of proprietary technology. This argument has a surface logic. If a company invests billions of dollars in developing a model, and a competitor captures that investment by training on its outputs, the competitor has extracted value without bearing the cost. But the surface logic collapses the moment the question is asked: where did the original model’s capabilities come from? They came from human-written text. Books, articles, code, essays, forum discussions — the accumulated written output of millions of people who were never asked for permission, never compensated, never even informed. The original extraction was orders of magnitude larger than any distillation campaign. The only difference is that the original extraction has now been receipted.
The receipt is the key that unlocks the double standard. Before the settlement, Anthropic’s training data had a provenance problem — it was acquired from pirate sites, and the authors were suing. If Anthropic had lost the lawsuit, the training data would have been legally tainted, and the company’s entire position on intellectual property would have been undermined. After the settlement, the provenance problem is resolved. The data is legally clean. And from legally clean data, legally clean models are built. And from legally clean models, a claim of ownership arises — a claim that can be enforced against anyone who would do to Anthropic what Anthropic did to the authors. The receipt does not just close a lawsuit. It opens the door to new ones.
The Chinese firms Anthropic accuses of distillation — DeepSeek, MiniMax, Moonshot — are not monolithic. Some of them release their models as open-weight, freely downloadable, self-hostable. DeepSeek’s V3 and R1 models are available for anyone to run on their own hardware. Qwen, from Alibaba, publishes its weights under permissive licenses. These models are smaller than their American counterparts, cheaper to run, and increasingly competitive on benchmarks. A person with a reasonably powerful desktop can run a distilled model locally, without sending data to any company’s servers. This is not a side note. It is the difference that makes the entire argument about distillation collapse into contradiction.
If distillation is theft, then theft comes in two varieties. The first variety — the American variety — takes human knowledge, trains a proprietary model on it, keeps the weights secret, charges for access by the token, and uses the law to prevent anyone from replicating the model or even understanding how it works. The second variety — the Chinese open-source variety — also takes outputs from larger models, but releases the resulting model freely, with weights available for anyone to inspect, modify, and run. Both acts involve training on outputs the original creator did not authorize. But one of them returns the result to the commons. The other encloses it.
The distinction matters because it reframes the entire debate. The AI industry wants the conversation to be about theft: they created something valuable, and Chinese firms stole it. But if the outputs of a language model are stolen when used for training, then every human writer whose work was ingested to train the original model was also stolen from. The difference is not in the act — training on unlicensed outputs — but in the direction of the flow. When knowledge flows from the commons into a proprietary model, it is called innovation. When knowledge flows from a proprietary model into an open model, it is called theft. The direction, not the act, determines the label.
This asymmetry is not an accident of law or a quirk of international relations. It is the structural purpose of the settlement. The $1.5 billion was not just the price of making a lawsuit go away. It was the price of reversing the flow. Before the settlement, knowledge flowed from the commons into Anthropic’s servers, and the legality of that flow was contested. After the settlement, the flow in that direction is settled — paid for, receipted, closed. The only remaining question is the flow in the other direction: from Anthropic’s servers back into the commons. And that flow, the company argues, is illegal. The receipt purchased the right to build a one-way valve.
There is a version of this story that does not end in contradiction. In that version, Anthropic trains Claude on the world’s books, releases the model weights openly, and says: here is what humanity’s collective knowledge can produce when synthesized through computation. Authors might still object to the training — and they would have every right to — but the objection would be political, not economic. The model would be a public resource, not a private product. And in that world, distillation would not be a crime. It would be optimization. A smaller, faster, cheaper model that anyone can run is an improvement on a larger model that requires a datacenter. Distillation would be celebrated as efficiency, not condemned as theft. The only reason it is condemned is that the original model was never meant to be a public resource. It was always meant to be a product. And products have owners. And owners call the police.
The tragedy is not that the ideal world is technically impossible. It is that it is economically impossible under the current incentives. An AI company that open-sources its most capable model gives away its primary revenue stream. A company that keeps its model closed can charge for access, build a moat, and use the accumulated capital to train ever-larger models that deepen the moat. The incentives point relentlessly toward enclosure. The settlement fits perfectly into this logic. It is not a deviation from the business model. It is a line item in it. $1.5 billion for perpetual rights to the training data, amortized over decades of API revenue. At scale, it is cheap.
And so the cycle begins again. The next generation of models will be trained not just on human-written text but on synthetic data — outputs of earlier models, including those that were themselves trained on human-written text. The provenance chains will grow longer and more opaque. Each layer will be receipted through new settlements, new terms of service, new laws. And at each layer, the owners of the receipt will claim the right to control the next layer. The commons will be mined, processed into proprietary output, receipted, and then fenced off. Extraction followed by enclosure followed by enforcement. The billion-dollar receipt is not the end of the story. It is the template for every chapter that follows.