The Great Book Shredding
There is something uniquely unsettling about destroying a book in order to teach a machine how to write.
Not copying it. Not borrowing it. Not photographing it and returning it to a shelf.
Destroying it.
That is the increasingly visible end of the artificial-intelligence industry's hunt for high-quality training data: millions of physical books purchased in bulk, their bindings removed, their spines cut away, their pages fed through industrial scanners, and the originals discarded.
The story recently circulated widely through a post by Hedgie Markets, which described AI companies buying old and rare books for destructive scanning and argued that an uncomfortable new business model had emerged around quietly turning the world's printed heritage into machine-learning data. Read the original Hedgie Markets post .
The post is deliberately provocative, and some of its claims need qualification. But the underlying practice is not speculation.
It is documented in federal court records.
And the economics behind it may be only beginning to take shape.
The machine wants books
Large language models need enormous quantities of text.
The internet provided the first great reservoir. Web pages, forums, Wikipedia, news articles, code repositories, academic papers and digitized books could all be collected at extraordinary scale.
But there is a problem with the internet of the mid-2020s: increasingly, it contains material produced by AI systems themselves.
For companies building the next generation of models, that makes older human-written material unusually attractive. A book printed in 1987 cannot secretly have been written by ChatGPT. A technical manual from 1974 contains no synthetic prose generated by Claude. A forgotten academic monograph from 1996 represents something increasingly valuable: uncontaminated human-produced text.
That was essentially the pitch uncovered by 404 Media. The publication reported that book-data company ISBNdb was advertising bulk sourcing of printed books for AI training, arguing that pre-2022 books offered professionally edited human knowledge without contamination from the flood of AI-generated material now appearing online. Read 404 Media's reporting .
Its marketing language was remarkably candid. According to the reporting, ISBNdb described printed books as some of the world's best remaining AI training data and promoted acquisition at enormous scale. The company also acknowledged the public-relations problem created when those books are destroyed during scanning.
There is an important qualification to that story.
After 404 Media published its report, ISBNdb removed the relevant pages and said the proposed service had never actually launched. The company stated that it had never purchased, scanned or sold a book for AI training and described the webpages as a test of market interest.
So it would be inaccurate to say that ISBNdb is currently operating a proven pipeline that has already delivered millions of books to anonymous AI labs.
But what the marketing page revealed is arguably more interesting than whether ISBNdb itself completed an order.
Someone believed there was a market for this.
And we already know that the underlying process exists.
Anthropic already did it
The clearest evidence comes from Bartz v. Anthropic, the copyright case involving the company behind Claude.
The court record describes Anthropic's acquisition strategy in unusually direct terms.
After initially obtaining millions of books from pirate libraries, Anthropic became more cautious about the legal implications of those sources. According to Judge William Alsup's ruling, the company hired Tom Turvey, formerly responsible for partnerships around Google's book-scanning project.
His mission was extraordinary in its simplicity: obtain "all the books in the world."
Anthropic then contacted major book distributors and retailers and spent millions of dollars purchasing millions of physical books, many of them used.
What happened next is not an allegation. It is described in the federal court's decision.
Service providers removed the books from their bindings, cut the pages down to scanner-compatible dimensions, digitized them, converted the scans into machine-readable files, and discarded the paper originals.
Millions of books went in.
A digital research library came out.
The physical books did not.
Read Judge William Alsup's ruling in Bartz v. Anthropic .
The legal twist
The strangest part of the story is that destroying the original helped Anthropic's legal argument.
Judge Alsup concluded that Anthropic's conversion of legally purchased print books into internal digital copies constituted fair use under the particular facts of the case.
The reasoning was narrow but consequential.
Anthropic had bought a physical copy. It converted that copy into a digital version. The physical original was destroyed. The digital replacement remained internal and was not sold or distributed.
"The print original was destroyed. One replaced the other."
Because the conversion did not increase the number of library copies—one physical copy effectively became one digital copy—the court found the format conversion permissible under the fair-use analysis.
The ruling did not say that anything an AI company does with books is automatically fair use.
In fact, the same decision rejected Anthropic's attempt to justify its separately acquired library of millions of pirated books. The distinction mattered: buying a book and converting it was treated differently from downloading an unauthorized copy and keeping it indefinitely.
But the practical incentive created by the ruling is difficult to ignore.
If you want a legally safer digital corpus, buying physical books and destroying them as you scan them may be more defensible than obtaining unauthorized ebooks.
That is an extraordinary outcome.
Copyright law may have inadvertently made the shredder part of the compliance strategy.
And now the book market is behaving strangely
This might remain merely an odd story about one AI company were booksellers not now seeing unusual purchasing patterns around the world.
The Guardian has reported that secondhand booksellers in Britain and Ireland received large orders with no obvious thematic relationship: obscure technical books, old biographies, specific editions, foreign-language titles and decades-old publications bought together in apparently machine-selected batches.
Similar patterns have been reported elsewhere, with booksellers describing buyers paying full prices for large, eclectic collections and sometimes operating through opaque identities or freight addresses.
Read The Guardian's reporting on unusual bulk book purchases .
The connection between every one of those orders and AI training has not been established.
That distinction matters.
But the behavior resembles exactly what one would expect from a system attempting to acquire text rather than books.
A collector cares about condition.
A reader cares about subject.
A dealer cares about resale value.
A training-data pipeline cares whether the text is new to the dataset.
That produces buying behavior that looks absurd to a traditional bookseller: a manual on agricultural machinery beside an obscure biography beside an academic title beside a forgotten novel, all purchased with equal enthusiasm.
The books may have almost nothing in common culturally.
As data, they have one important property in common:
They contain words the model may not have seen before.
The rare-book question
This is where the story becomes more complicated.
The claim that AI companies are systematically destroying the world's rarest books is stronger than the currently available evidence supports.
Anthropic has said that its acquisition programs do not buy and destroy rare or antiquarian books. The original court ruling documenting Anthropic's destructive-scanning program does not establish that the millions of books it purchased were rare.
That qualification should not be buried.
At the same time, the preservation concern is no longer hypothetical.
Reporting connected a large and unusual order of secondhand books to an Amazon facility in Las Vegas associated with book scanning. An investigation originating with 404 Media used a tracking device hidden in a book to follow the shipment to the facility, while subsequent reporting described destructive book-processing activity there.
Read the reporting on the tracked book shipment .
None of this establishes the most dramatic version of the story—that AI companies are knowingly hunting down and shredding the final surviving copies of unique historical works.
But it creates a preservation problem that should be obvious.
At industrial scale, rarity is not always visible.
A book does not need to be a Gutenberg Bible to be culturally scarce.
It can be a provincial history printed in 1913.
A privately published memoir.
A small-run scientific proceedings volume.
A technical manual from a company that disappeared fifty years ago.
A regional cookbook.
A local-language translation.
A university monograph that was printed once and never reissued.
There may be twenty copies in the world.
Or five.
Or one.
And a purchasing algorithm working from ISBNs, metadata and availability does not necessarily know the difference between a worthless duplicate and an artifact that has quietly become scarce.
Digitization is not the same thing as preservation
This distinction deserves more attention.
Scanning a book can preserve its information.
It does not necessarily preserve the book.
Libraries and archives understand this distinction very well.
A physical book contains more than the sequence of characters printed on its pages. Editions differ. Marginalia matters. Paper matters. Bindings matter. Printing errors matter. Inscriptions and ownership marks matter. Illustrations, plates, foldouts, typesetting and physical construction can all become important to historians.
The artifact itself carries information.
And even when only the text matters, there is another problem.
A private corporate dataset is not a public archive.
If a company purchases a scarce book, destroys it, scans it and stores the resulting text inside a proprietary training corpus, humanity has not necessarily gained a preserved digital copy.
The company has.
Those are not the same thing.
A preserved work normally remains available to researchers, libraries, historians and future readers.
A training dataset may never be published at all.
The physical artifact disappears while its digital ghost becomes corporate infrastructure.
Calling that "digital preservation" stretches the term almost beyond recognition.
There is a counterargument
Books are destroyed every day.
Publishers pulp unsold inventory. Libraries withdraw duplicate or damaged copies. Used-book dealers recycle stock that has sat unsold for years. Warehouses dispose of enormous quantities of printed material because storing books costs money.
Most copies of most books are not museum objects.
And destructive scanning is efficient. Once a binding is removed, loose pages can pass rapidly through automated sheet-fed scanners. For a project involving millions of volumes, the difference in cost and throughput between destructive and non-destructive scanning can be enormous.
There is also something undeniably valuable about digitizing neglected books.
An obscure book that has sat unread for seventy years may become searchable for the first time.
Its ideas may become discoverable again.
In principle, mass digitization could be one of the greatest preservation projects ever undertaken.
Google Books demonstrated a different version of that possibility: books could be digitized at enormous scale while the physical volumes, many borrowed from libraries, continued to exist.
The problem is not digitization.
The problem is an incentive structure in which destruction is cheaper, secrecy is commercially useful, the resulting archive is private, and the cultural cost is externalized to everyone else.
The irony is almost too perfect
Artificial-intelligence companies are increasingly concerned about the quality of the internet because artificial-intelligence companies have filled the internet with artificial-intelligence output.
So they are looking backward.
Toward books.
Toward writing created before the current generative-AI boom.
Toward texts selected, edited and published through human institutions.
In other words, the machines require precisely the cultural substrate that preceded them.
The older and more human the material, the more useful it may become.
That creates a grotesque circularity.
We built systems to synthesize human writing.
Those systems produced so much synthetic writing that human writing became a premium resource.
And now companies building the systems are purchasing physical repositories of that human writing so machines can ingest them before the corpus becomes further contaminated by machines.
There is a technical term for contaminated training data.
There should probably be one for the cultural absurdity of the solution.
Ownership is not stewardship
If you legally purchase a book, you are generally free to tear it apart.
You can write in it.
You can use it as a doorstop.
You can burn it.
You can feed it into a scanner and throw the pieces into a recycling bin.
That is part of what ownership means.
But legality and stewardship are different concepts.
Industrial AI changes the scale of the question.
One person destroying one old book is usually irrelevant.
A billion-dollar company designing an acquisition system capable of purchasing and processing millions of books is something else entirely.
At that scale, individual property rights collide with collective preservation.
Markets are good at determining how much somebody is willing to pay for a book.
They are much less reliable at determining how important it will be that a copy of that book still exists fifty years from now.
That is why libraries, archives and museums exist in the first place.
They preserve things whose future value cannot be expressed by today's resale price.
The real danger
The Great Book Shredding is not frightening because every destroyed book is priceless.
Most aren't.
It is frightening because once books become raw material for AI infrastructure, the incentives governing them change.
Their value is no longer primarily as objects that can be read, sold, collected or preserved.
Their value lies in being converted into tokens.
And conversion is fastest when the spine comes off.
That creates an asymmetry that should concern anyone who cares about the historical record.
An AI company can preserve the informational value it wants while destroying the physical object society might later discover it needed.
The company keeps the data.
Everyone else loses the artifact.
Perhaps the strangest part is that none of this requires a conspiracy.
It does not require secret orders to destroy culture or executives sitting around deciding which books should disappear.
It requires only optimization.
Buy at scale.
Minimize acquisition costs.
Maximize scanning throughput.
Reduce legal risk.
Keep the dataset proprietary.
Discard the physical input.
Every individual decision can be economically rational.
The result can still be culturally insane.
And that may be the defining characteristic of this stage of the AI boom: outcomes so strange that nobody needed to intend them.
The machines are hungry.
The internet is increasingly polluted with their own output.
And somewhere, sitting quietly on shelves, is one of the largest remaining collections of uncontaminated human thought on Earth.
We should probably decide how much of it we are willing to put through the shredder.
Sources and further reading
- Hedgie Markets — the post that prompted this article
- 404 Media — AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop
- Bartz et al. v. Anthropic PBC — federal court fair-use ruling
- The Guardian — reporting on unusual bulk purchases from secondhand booksellers
- Tom's Hardware — reporting on the tracked book shipment and Las Vegas scanning facility
- The Washington Post — background on Anthropic's book-scanning program
POPULAR ARTICLES
View all