⏱️ Lectura: 10 min

A service called ISBNdb lets artificial intelligence companies order rare books in quantities of up to one million units per order, and part of those orders ends up shredded as soon as scanning is complete. The platform bills orders from anonymous buyers, and according to booksellers consulted by trade media, copies with very few surviving originals also enter that circuit.

📑 En este artículo
  1. TL;DR
  2. Introduction
  3. What happened to the rare books
  4. Context and history
  5. Technical details and performance
  6. How to verify it
  7. Impact and analysis
  8. What’s next
  9. Frequently Asked Questions
    1. What is Project Panama?
    2. Why do they destroy the book instead of keeping it after scanning?
    3. Is it legal to buy and shred rare books to train an AI model?
    4. Why are books published before 2022 worth more in this market?
    5. What is ISBNdb and what role does it play?
    6. What can a library do to protect a rare copy?
  10. References

This isn’t an isolated practice: it’s part of a race to secure training text that hasn’t been contaminated by AI-generated content. A federal judge already ruled the process to be fair use, clearing the legal path for companies like Anthropic to accelerate it.

TL;DR

  • ISBNdb enables orders of up to one million books and keeps buyers anonymous.
  • AI companies scan the books with high-speed machines that cut the spine, then destroy the original.
  • Books published before 2022 sell at a premium because they contain no AI-generated text.
  • A federal judge ruled the practice is fair use because, by eliminating the original, only one copy remains in circulation at a time.
  • Anthropic hired the former head of partnerships at Google Books for an internal project known as Project Panama.
  • According to court documents cited in the report, the project’s stated goal was to destructively scan every book in the world.
  • ISBNdb offers non-disclosure agreements (NDAs) to its clients and suggests describing the practice as digital preservation.

Introduction

The phenomenon of rare books shredded after scanning resurfaced this week from a thread by analyst Hedgie (@HedgieMarkets), who reconstructed how some training data providers for language models operate. The case connects three pieces: a brokering service (ISBNdb), a court ruling that enables the destruction of the original, and a company, Anthropic, which according to the thread hired staff specialized in mass library digitization.

The chain starts with the purchase. ISBNdb allows buyers to order entire catalog volumes, including out-of-print titles or small print runs, without the seller knowing who’s buying or why. That anonymity is part of the service, not a side effect: the platform itself offers it as another feature to attract institutional buyers.

What happened to the rare books

According to the account, AI companies buy rare books in bulk, process them in high-speed scanning machines that cut the spine so each sheet can be fed separately, and destroy what’s left of the physical volume once the process is complete. A federal judge determined that this practice is fair use: the central argument is that, by eliminating the original at the moment of digitization, two copies (one physical, one digital) never exist in circulation at the same time, which avoids the duplication problem that usually concerns copyright law.

The thread also notes that Anthropic hired the former head of partnerships at Google Books, the division that in the 2000s scanned millions of volumes from university libraries, with the stated goal, according to the internal proposal cited, of obtaining every book in the world. That effort is known internally as Project Panama, according to court documents mentioned in the thread, and reportedly involved spending tens of millions of dollars.

One detail explains the market’s urgency: books published before 2022 sell at a premium over more recent editions. The reason is that content published after that date is more likely to include text generated or rewritten by AI, something data teams want to avoid so as not to train a model on its own synthetic output.

Stack of rare books next to a destructive scanning machine
Copies with few surviving originals also enter the circuit, according to booksellers consulted. Foto de Cristina Gottardi en Unsplash

Context and history

Mass book digitization isn’t new. Google Books, launched in the mid-2000s, scanned millions of volumes from libraries like Harvard, Stanford, and Oxford’s Bodleian, and ended up in a lengthy lawsuit, Authors Guild v. Google, resolved in Google’s favor in 2015 on the grounds that showing book excerpts for search purposes is fair use. That precedent is the legal foundation AI labs now rely on to defend training on copyrighted works.

The difference with the current scenario is that the digital copy no longer coexists with the original: the physical book is destroyed. That variant was already evaluated in the litigation known as Bartz v. Anthropic, where a federal court distinguished between training on pirated books (which did create liability) and training on legally purchased books, even if the digitization process involves destroying the physical copy.

That nuance is key to understanding why the rare book market unintentionally became a strategic input for the AI industry: buying in bulk and destroying afterward is, according to that judicial reasoning, more legally defensible than scanning and keeping two copies. It’s also the first time a US court has explicitly endorsed destroying the original as part of a large-scale digitization practice.

Technical details and performance

The destructive scanning process isn’t exotic: it’s the same method archives and publishers use to digitize mass print runs when preserving the physical object doesn’t matter. An industrial guillotine cuts the book’s spine, the loose sheets pass through a high-speed automatic-feed scanner, and an OCR pipeline converts each page into plain text, which is then cleaned and structured for training.

flowchart TD
A["Anonymous buyer orders on ISBNdb"] --> B["Physical copies are acquired, up to 1,000,000 per order"]
B --> C["Guillotine cuts the spine of the book"]
C --> D["High-speed scanner digitizes each sheet"]
D --> E["OCR and text cleanup"]
E --> F[("AI training corpus")]
C --> G["The physical volume is shredded"]

The following table compares the destructive method described in Hedgie’s thread with the non-destructive scanning used by libraries and heritage archives when they do need to preserve the original object:

MethodWhen it’s usedAdvantageLimitation
Destructive scanning (guillotine + sheet feeder)High volume, no intention of preserving the physical copyHigher speed and uniform OCR quality thanks to flat pagesDestroys the original: irreversible
Non-destructive scanning (overhead scanner)Rare books, archives, and heritage collectionsPreserves the physical object intactSlower and requires careful manual handling

There’s no public figure in the available material for pages per hour or cost per volume in Anthropic’s pipeline: ISBNdb and the buying companies don’t publish those numbers. What is verifiable is the purchasing mechanism: anyone can check ISBNdb‘s catalog and confirm the service is designed for mass orders by ISBN, not individual collector purchases.

Loose pages of a rare book passing through a high-speed scanner
The spine is cut before scanning to feed loose sheets into the machine. Foto de QingYu en Unsplash

How to verify it

To confirm the court ruling that enables this practice, the most direct way is to check the public docket for the Bartz v. Anthropic case on CourtListener, where orders and motions from US federal courts are archived. That’s where the references to Project Panama cited in the original thread appear.

For a bookshop or independent seller, checking whether a given copy risks ending up in this circuit is simple: before selling a large lot to an anonymous buyer via ISBNdb or a similar intermediary, search the ISBN on WorldCat to see how many libraries worldwide still report holding that title in their catalog. If the number is low, in the single digits, that copy is a candidate for an irreplaceable piece and should be offered first to a library with a preservation mandate rather than to a buyer who won’t reveal their identity.

Impact and analysis

The legal reasoning behind the ruling (only one copy existing avoids the duplication problem) has a consequence that Hedgie’s own thread points out: the destruction is irreversible. A website can be re-uploaded. A bestseller can be reprinted. The last surviving copies of an 18th-century text cannot be recovered once shredded.

⚠️ Heads up: unlike data scraped from the web, a shredded book allows no later correction. If a higher court reverses the fair use ruling on appeal, the physical object that would have allowed a re-evaluation would no longer exist.

ISBNdb, according to the thread, is aware of the perception problem: its own site acknowledges that an AI company destroys two million books isn’t a headline that wins sympathy, and yet it built a business model around facilitating it discreetly, including non-disclosure agreements for its clients and the suggestion to describe the process as digital preservation.

The other effect is on the market: as books free of AI-generated text contamination become valuable, it creates a direct economic incentive for more booksellers to sell complete lots instead of offering individual titles to collectors or libraries, who typically pay less and buy one copy at a time. That shifts sales volume away from the traditional rare book circuit toward intermediaries oriented toward institutional data buyers.

What’s next

With legal backing already confirmed in court, the expectation, per the thread itself, is that the practice will accelerate rather than slow down. As of this writing, there’s no specific legislative proposal in the United States distinguishing between digitizing and preserving the original versus digitizing and destroying it. Until that distinction exists in law, the judicial standard that endorsed destruction as part of fair use remains the only reference framework.

Librarians and heritage preservation organizations, referenced indirectly in the conversation that sparked the thread, argue that the most realistic short-term response isn’t legal but logistical: identifying and buying the rarest titles first, before they reach an intermediary like ISBNdb.

Try it yourself: look up the ISBN of a rare or out-of-print book you have on hand at isbndb.com and compare how many libraries have it registered on WorldCat before deciding what to do with it.

📖 Summary on Telegram: View summary

Frequently Asked Questions

What is Project Panama?

It’s the internal name, cited in court documents according to the original thread, for an Anthropic project aimed at acquiring as many books as possible for training, including their destructive digitization.

Why do they destroy the book instead of keeping it after scanning?

According to the judicial standard cited, keeping only one copy (the digital one, after eliminating the physical one) avoids the legal problem of two versions of the same content being in circulation at once.

A federal judge ruled the practice to be fair use in the Bartz v. Anthropic litigation, as long as the books were legally acquired before digitization.

Why are books published before 2022 worth more in this market?

Because they’re less likely to contain text generated or rewritten by artificial intelligence, something data teams avoid so as not to train a model on its own synthetic output.

What is ISBNdb and what role does it play?

It’s a service that lets buyers order books by ISBN in bulk, up to one million units per order, while keeping the buyer’s identity anonymous.

What can a library do to protect a rare copy?

Check union catalogs like WorldCat to see how many copies survive worldwide, and prioritize preserving or acquiring the titles with the fewest registered copies before they’re sold to anonymous buyers.

References

📱 Enjoying this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Imagen destacada: Foto de Andrew V en Unsplash


Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

This site uses Akismet to reduce spam. Learn how your comment data is processed.