If you’ve been waiting for a clean yes-or-no answer, you’ll be waiting a while longer. In 2025 and 2026, U.S. courts finally started ruling on this question, and the answer that emerged is more nuanced than either side wanted: it depends on how the books were obtained, not just what they were used for.
The Big Split: Training vs. Piracy
The clearest signal came from Bartz v. Anthropic. Judge William Alsup ruled that training a large language model on legitimately acquired books counts as fair use, describing the process as transformative because the model learns patterns rather than reproducing the original text. But he drew a hard line at how the books got into the training set: building a permanent library from pirated copies was not protected, regardless of what happened afterward. That distinction — acquisition versus use — proved costly. Anthropic settled the underlying piracy claims for $1.5 billion, one of the largest copyright settlements on record, while the fair-use finding on training itself stood.
A similar pattern showed up in Kadrey v. Meta, where the court also found training to be transformative, since a book’s original purpose is to be read, while its use as training data serves a different function entirely. But the judge flagged a more troubling argument: an AI model trained on an author’s work could flood the market with similar content and undercut sales of the original — the kind of market substitution fair use is meant to guard against. The authors in that case didn’t build enough evidence to prove it, so Meta prevailed, but the reasoning signals the theory is very much alive for future plaintiffs.
Not Every Case Goes the AI Company’s Way
Fair use hasn’t been a blanket shield. In Thomson Reuters v. Ross Intelligence, a court found that training a legal research tool on Westlaw’s copyrighted headnotes was not fair use, in part because the output competed directly with Westlaw’s own product. The case is now on appeal, but it’s a reminder that outcomes shift with the type of content, the AI’s purpose, and how closely its output resembles the market for the original.
The Common Thread
Across these rulings, a few factors keep deciding outcomes: whether the training data was lawfully acquired, whether the use is genuinely transformative, and whether the resulting AI competes with or substitutes for the original works. The U.S. Copyright Office’s own analysis, released in 2025, reached a similarly conditional conclusion — some AI training will qualify as fair use, and some won’t, depending largely on whether the output competes with the source material.
Why “Complicated” Is the Honest Answer
Several major cases — including disputes involving The New York Times, Getty Images, and various music publishers — are still working through the courts, and each involves different content types, business models, and claims about market harm. Appeals are pending. International law adds another layer, since the EU, UK, and other jurisdictions apply their own frameworks, some more permissive of text-and-data mining than U.S. fair use doctrine.
For now, the practical takeaway is this: courts appear increasingly comfortable with the idea of training on copyrighted books, but increasingly unforgiving about the sourcing of that data. Piracy is a liability regardless of how the model is later used. Licensing deals, once optional, are becoming a genuine risk-management strategy for AI companies. And for authors and publishers, the fight is shifting from “can they train on my book at all” to “how was my book obtained, and does the output compete with me.”
The law isn’t settled. It’s converging — slowly, and unevenly — toward a rule that separates lawful use from lawful acquisition.




