The ruling does change the fundamental nature of what the books were intended for: they were intended to be read by human beings, not used as training for machines to replace human thought and ingenuity.
The revelations came out in a lawsuit brought by authors who said the destruction violated the Copyright Act. Federal judge for the Northern District of California William Alsup ruled that the bulk buy, scanning, and destruction is fair use under Section 107 of the Copyright Act. As far back as 2024, Anthropic, which has come under fire by the Trump administration over national security concerns, said "We don’t want it to be known that we are working on this." It's not a good look to be destroying books.
The suit reads that authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson were a few of the authors whose books were bought by Anthropic. For use in Claude, Anthropic "assembled these copies into a central library of its own, copied further various sets and subsets of those library copies to include in various 'data mixes,' and used these mixes to train various LLMs. Anthropic kept the library copies in place as a permanent, general-purpose resource even after deciding it would not use certain copies to train LLMs or would never use them again to do so." The authors had not agreed to this and these copies of their books were taken permanently out of circulation, never to be read by human eyes again.
Anthropic is one of these companies that has been buying physical books in bulk to rip apart, scan, then destroy, in service to its AI model Claude. Any books will do, even if that book is the very last copy of that book in the world. Legally, this is allowed under the first-sale doctrine, which permits a book-buyer to do whatever they want with that book object, without a need for the copyright owner's permission. It's an object, like a pair of shoes or a blender.
The ruling reads: "every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company."
In the suit, a judge found that because Anthropic was using the books in a way that fundamentally transformed them into something else, in this case turning them from books into raw data that will have no reference to their original self, and that this constituted fair use. Previous to the bulk buying of books, Anthropic pirated books or bought pirated books. They looked to former head of Google Books' scan project Tom Turvey to come in and get more data to feed the LLM. That data was not just information but also style and tone, proprietary authorial elements.
There's lots of companies that are willing to do the selling. ISBNdb claims to have the "world's largest book database," and says "the world's best AI training data is sitting on a shelf." They also promise not to reveal who is doing the bulk buying, knowing that no AI company wants to be the subject of a headline about how they're destroying millions of books at a go.
ISBNdb says that AI training on AI material results in data degradation. "Not all data degradation is intentional," says ISBNdb. "When AI systems train on text that was itself AI-generated, a documented phenomenon called model collapse occurs: subtle linguistic nuances vanish, systematic errors compound, and outputs converge on repetitive patterns. Each generation trained on synthetic data is slightly worse than the last. Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage."
404 Media spoke to a bookseller who said that bulk orders from his store have resulted in the destruction of out-of-print books for which there are only single copies remaining. "I personally have mixed feelings about all of this," said the bookseller. "It benefits me financially as well as by clearing out old inventory that is otherwise unlikely to sell. I’ve been well-suited for these sales with inventory from overseas and foreign language books. On the other hand, I don’t like the end-use, and I don’t like that uncommon books are being pulped."
The ruling reads: "This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies." The ruling does change the fundamental nature, however, of what the books were intended for: they were intended to be read by human beings, not used as training for machines to replace human thought and ingenuity.
As it stands, Anthropic and other AI companies that are developing and furthering their large language models for use both as information and writing tools are permitted to take books out of circulation, feed them to the LLM, and destroy them, without concern for copyright law or, as it turns out, the preservation of the scope and breadth of human history, knowledge, experience, and creativity.
Powered by The Post Millennial CMS™ Comments
Join and support independent free thinkers!
We’re independent and can’t be cancelled. The establishment media is increasingly dedicated to divisive cancel culture, corporate wokeism, and political correctness, all while covering up corruption from the corridors of power. The need for fact-based journalism and thoughtful analysis has never been greater. When you support The Post Millennial, you support freedom of the press at a time when it's under direct attack. Join the ranks of independent, free thinkers by supporting us today for as little as $1.
Remind me next month
To find out what personal data we collect and how we use it, please visit our Privacy Policy


Comments