For years, “offline” mostly meant “harder to find”. For AI labs, it is becoming a category of raw material.

In 2024 Anthropic bought millions of printed books, often second-hand. Its service providers removed bindings, cut pages to a workable size, scanned each volume and discarded the paper original. Judge William Alsup describes the chain directly in Bartz v. Anthropic.1

The company was not building a public library. Files entered an internal “research library” from which Anthropic could select sets of text for model development. A later-unsealed internal document gave the programme an almost implausibly neat name: Project Panama, described as an effort to destructively scan “all the books in the world”.23

That image invites the wrong argument. This is not Fahrenheit 451 with GPUs replacing the firemen. Books are not being destroyed to suppress their contents. They are destroyed because, in an industrial scanning line, a stack of loose sheets moves faster than a bound volume.

The more useful question is what happens when the text is extracted from an object and the object is then treated as disposable packaging.

Off the web

Why buy physical books in 2026 when billions of pages are already online?

The first answer is simple: not everything is online. Used-book catalogues contain local monographs, technical manuals, out-of-print editions, conference proceedings, obscure biographies and foreign-language works that never received a useful digital edition. For a model developer that has already consumed large amounts of reachable web text, this offline material is a different reservoir.

Booksellers in Britain and Ireland told The Guardian about unusual orders with little thematic coherence, opaque buyers, little price negotiation and freight destinations that reveal almost nothing about the final customer.4 An AI connection is plausible, but the evidence has a hard boundary: those anonymous orders do not prove that a particular lab sits behind every purchase.

The interest in pre-2022 books adds a second hypothesis. Buyers appear to value texts that predate the flood of machine-generated writing online and may therefore contain less synthetic material. That is a corpus-quality argument, not a magical property of paper. A 2014 print book can be relatively clean human text; so can a 2014 PDF.

Paper matters when it remains the practical route to the text.

Keeping these levels separate matters. It prevents a documented industrial practice from turning into a story in which every rare book sold online automatically vanishes inside an AI shredder.

Cut to scan

The June 2025 court order describes a straightforward process. Anthropic purchased print copies, had them dismantled, cut the pages to scanner-friendly dimensions, made PDFs containing machine-readable text and discarded the paper.1

The key variable is throughput.

A bound volume requires page turning, support for the opening angle, correction for curvature near the gutter and care around the spine. Loose sheets can move through automated equipment. Across a handful of books the difference is modest. Across hundreds of thousands or millions, it defines the industrial system.

Destruction is nevertheless not a technical requirement.

The Internet Archive describes its Table Top Scribe as a non-destructive colour-digitisation system with a V-shaped cradle and two cameras, rated at roughly 500 to 800 pages an hour depending on the job.6 The British Library likewise uses specialist rests, cradles, photography and handling procedures for rare and fragile material.7

An operator uses a V-shaped Table Top Scribe to digitise a bound book
The Internet Archive's Table Top Scribe photographs pages while preserving the binding. The trade-off is a more operated, slower pipeline than feeding loose sheets through a scanner.Internet Archive / EDUCAUSE Review

Google says that as early as 2003 its Books team bought volumes specifically to experiment with non-destructive scanning and developed a gentler method than common high-speed processes.8 Twenty-three years later, the underlying trade-off is familiar: preserving a binding costs time, mechanics and handling.

Anthropic's choice was therefore less about whether another method existed and more about what it was willing to spend per book to use one.

Two cases

The legal story contains a distinction that disappears in many summaries.

Anthropic had two very different sources of books. Before its large-scale print purchasing, it had downloaded more than seven million copies from Books3, Library Genesis and Pirate Library Mirror. Judge Alsup explicitly separates those pirated copies from the physical copies the company bought.1

For purchased books that were destroyed after scanning, the court held that replacing a lawfully acquired print copy with an internal digital copy was fair use in the circumstances of the case. One copy replaced another, and the digital library copy was not sold or distributed outside Anthropic. The order also treated the use of copies for LLM training as transformative.1

The pirate library was different. Those copies had not been purchased, and Anthropic retained some even after deciding they would not be used for model training. Their acquisition did not receive the same legal protection.1

The $1.5 billion settlement announced in 2025 relates to claims around pirated books, rather than being a penalty for cutting the bindings off legally purchased copies.9

That distinction also matters for preservation. An act can be lawful to copy and still be questionable to destroy. Copyright allocates rights around reproduction; it is not a preservation policy for physical artefacts.

Amazon too

In August 2026 the story stopped being only a retrospective look at Anthropic.

404 Media worked with a bookseller who had received an order of roughly a thousand books and placed an AirTag in one volume. The tracker eventually reached Amazon's LAS8 complex in Las Vegas, near an internal operation called VGT3. Worker accounts describe a line where some employees receive and scan barcodes while others cut bindings so books can be scanned.5

Amazon has not publicly identified every model or product that receives those scans. Its response to Ars was broader: the company purchases books through commercial channels to help develop and improve products and services for customers.5

This is stronger evidence than a bookseller's suspicion, but narrower than a full disclosure of an AI training corpus. We know where one bulk order ended up. We know an Amazon unit cuts and scans books. We do not have a public inventory of every downstream dataset, model and selection rule.

That asymmetry makes the subject unusually difficult to audit: the physical supply chain leaves visible traces while the data pipeline largely does not.

What disappears

A text and a particular copy of a book are not identical things.

For a recent paperback printed in hundreds of thousands of copies, sacrificing one unit may have almost no heritage impact. Libraries weed collections, unsold inventory is pulped and used books become paper fibre every day for reasons unrelated to AI.

For a low-circulation volume, the calculation changes. Binding, marginalia, stamps, paper, inscriptions, handwritten corrections or simply the scarcity of a particular edition can carry information that disappears when a book is reduced to page images and OCR text.

The British Library describes digitising chapbooks for which some examples may be the only known surviving copy, precisely so that those cultural assets can be preserved.7 Its logic is almost the reverse of the destructive pipeline: the digital object is made to protect the physical one, rather than to justify its removal.

The documented problem is therefore not that every destroyed book represents unique heritage. It is that there is no public filter showing when the copy is ordinary and when it is not.

A British Library technician digitises an older volume with a camera above the open book
In a preservation workflow, digitisation is organised around the material condition of the volume. The text is only one layer of what may need to survive.British Library

Anthropic showed that a company could buy millions of books and convert them into a private digital library. Amazon shows that a comparable physical pipeline is operating now. Booksellers, meanwhile, continue to see opaque bulk demand move through the second-hand market.145

At that point the sharpest issue may be less destruction than asymmetry.

The bookseller parts with an object without knowing whether it is going to a reader, a reseller or a data factory. The public may lose the copy. The buyer retains searchable text and a private corpus advantage.

Keep the object

A better rule need not be a dramatic ban on scanning books.

Before destructive scanning, check bibliographic scarcity. Pull out signed, annotated, unique or poorly represented editions. Use non-destructive capture when the value of the object warrants it. When industrial-scale destruction is chosen, record enough metadata to identify what disappeared. For public-domain works, opening the scans would prevent a publicly available physical object from being converted into a purely private digital asset.

None of these techniques is exotic. Libraries already choose digitisation methods according to the condition and significance of the material in front of them.67

What has changed is the scale and the beneficiary.

The old Google Books and preservation narrative was roughly: digitise so a work can be found. Project Panama almost reverses the sentence: make the text machine-usable, then remove the physical storage problem.

Both processes can produce a PDF. They do not produce the same world around it.