Amazon Is Buying and Destroying Rare Books to Train Its AI. Here's What That Means for Professionals.
A 404 Media investigation tracked rare books to an Amazon warehouse where workers cut off spines, scan the pages, and discard the originals. This is how AI companies are solving their training data problem — and it explains something real about what your AI tools know.
See Claude set up for your job
Skip the theory — pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
Models change every month.
One short weekly update that keeps this call current — free.
TL;DR. On August 17, 2026, investigative outlet 404 Media published the results of a tracking experiment: place a GPS device in a shipment of rare books sold through Biblio, and follow it. It ended up at VGT3, an Amazon warehouse in Las Vegas, where workers cut off spines, scan the pages, and discard what's left. Amazon is systematically acquiring physical rare books to train its AI models — acquiring knowledge that doesn't exist anywhere online. Here's why AI companies are doing this, what it reveals about how AI tools work, and what professionals should understand.
The Investigation
On August 17, 2026, Emanuel Maiberg at 404 Media published an investigation that put a tracking device into a shipment of rare books listed for sale on Biblio, a rare-book marketplace. The device led to VGT3, an Amazon fulfillment facility outside Las Vegas, Nevada. An anonymous Amazon employee confirmed the operation simply: "All we do is scan books."
TechCrunch corroborated the investigation in independent reporting. The process: workers at the facility cut off book spines, feed loose pages through high-speed scanners, and discard the originals after capture. Amazon's official statement offered no specifics: "Amazon purchases books through commercial channels to help develop and improve the products and services our customers use."
The investigation tracked roughly 1,000 books in one traced shipment. The systematic logging of ISBN barcodes at the facility suggests the actual scale is considerably larger — possibly a systematic attempt to digitize rare printed works across many categories.
Why Physical Books, and Why Rare Ones?
AI labs need training data — enormous quantities of text for language models to learn from. But the internet, once a seemingly inexhaustible source of human-written text, has a problem: since large language models went mainstream in 2022, it has been rapidly filling with AI-generated content. Articles, summaries, forum posts, documentation — an increasing share of what's published online was written by AI.
Training the next generation of AI models on that AI-generated content causes what researchers call model collapse: a gradual degradation where the model loses nuance, hallucinates more frequently, and converges on bland, generic-sounding outputs. The model's knowledge narrows toward whatever average patterns the prior AI had already baked in.
Rare physical books solve this problem cleanly. Books published before 2022, especially ones never digitized, are guaranteed to be:
- Human-written. No AI touched them.
- Historically deep. They contain knowledge, reasoning styles, and terminology that reflect decades or centuries of human thought in a domain.
- Not already in a competitor's training set. Common digital text (Wikipedia, web crawls, digitized public domain books) has already been consumed by everyone. Rare out-of-print books represent genuinely novel data.
For AI companies, rare physical books are one of the last reserves of high-quality, unduplicated, human-generated training material. The economics are simple: a rare book that costs $50 at auction may be worth far more as training data for a model trained at $100 million.
What This Means for AI Quality — and for You
Training data composition is a real differentiator in AI output quality. Providers rarely disclose what's in their training sets, but the pattern is consistent: AI tools are more reliable on topics that are heavily documented in digitized text, and weaker on topics that live in physical archives, specialist literature, and out-of-print sources.
Consider what that means for professional domains:
- A lawyer asking about obscure pre-digital case law gets a different quality answer depending on whether historical legal treatises and court reports were in the training data.
- A physician researching rare syndromes first described in pre-digital medical literature may find that AI knowledge becomes thin or approximate past a certain historical depth.
- An engineer or architect working with legacy systems or historical construction methods will encounter limits in AI knowledge when the reference material exists only in printed manuals never transferred to digital form.
This doesn't mean AI tools are unreliable — they're generally excellent for well-documented topics. It means the accuracy gradient matters: AI answers about specialized, historically rich, or niche professional topics deserve more verification than AI answers about mainstream subjects with abundant online coverage.
The Legal and Ethical Picture
The legal status of physical-book scanning for AI training sits in a gap that courts haven't fully addressed. Scanning and discarding a physical book — rather than copying a digitized version — is factually different from the digital piracy cases that have been litigated. A prior lawsuit against Anthropic (for training on pirated digital copies) established some related precedent, but digital piracy and physical-to-digital conversion are legally distinct. Amazon appears to be operating in this ambiguity deliberately.
What's harder to dispute is the practical effect on book communities. One rare-book dealer told investigators: "They just want the content as a bunch of words strung together." Rare books — historical records, specialized treatises, texts that document now-vanished fields of practice — are being acquired as consumable raw material, not preserved as cultural artifacts.
Publishers, libraries, and archivists have pushed back on AI companies' training data practices generally. Some are negotiating licensing agreements. Others are exploring legal challenges. But for now, the physical secondary market appears to remain an open channel.
What Professionals Should Do With This
Nothing urgent. This story doesn't require any immediate action. But three things are worth holding:
-
Calibrate your trust by domain. AI tools are not uniformly knowledgeable. For questions where the authoritative literature is mostly pre-digital, old, or specialized, verify AI answers against primary sources. This is especially true in law, medicine, engineering, and academic research.
-
Ask about training data when it matters. If you're evaluating an enterprise AI subscription for a domain with deep specialist literature, asking about training data sourcing is a fair question. Some providers disclose more than others, and for niche professional domains, it can meaningfully affect output quality.
-
Understand that this is what AI progress costs. Rare physical texts are being consumed to make the AI tools professionals use every day more capable. That's a real tradeoff — and understanding it helps you think more clearly about what AI systems are, what they cost, and who pays.
Sources
-
404 Media — "Amazon Buys and Destroys Rare Books to Train AI" — Emanuel Maiberg, August 17, 2026. Primary investigative source; includes the GPS tracking experiment.
-
TechCrunch — "Amazon, which started off selling books, is destroying rare texts to train AI models" — August 17, 2026. Independent corroboration with process details and scale context.
-
Futurism — "Amazon Is Destroying Rare Books to Train AI" — August 17, 2026. Additional reporting including scale estimates and bookseller reactions.
See Claude set up for your job
Skip the theory — pick your profession and get the real workflows, ready-to-use prompts, and exact setup for your work.
Set up AI for your job — free, in about 2 minutes
Pick your profession and get your first working AI tool, a step-by-step guide, and a $0 plugin to take home. No credit card.
Get my free setupSee Claude set up for your job
Real workflows and ready-to-use prompts, profession by profession.
Frequently asked questions
Is Amazon really destroying rare books to train AI?+
Yes, according to an August 2026 investigation by 404 Media. Journalists placed a GPS tracking device in a shipment of rare books listed on Biblio, an online rare-book marketplace. The device led them to VGT3, an Amazon fulfillment facility in Las Vegas. An anonymous worker at the facility confirmed: 'All we do is scan books.' TechCrunch corroborated the investigation, describing the process: workers cut off book spines, run loose pages through scanners, and discard the originals. Amazon's official response: 'Amazon purchases books through commercial channels to help develop and improve the products and services our customers use.'
Why would Amazon want physical rare books instead of digital text?+
Because pre-digital books solve a problem called 'model collapse.' Since large language models went mainstream in 2022, the web has been filling with AI-generated text — articles, summaries, documentation — all written by AI. If you train the next generation of AI on that content, the model gradually degrades: it gets blander, hallucinates more, and loses the precision that makes answers useful. Physical books published before 2022 — especially rare ones never digitized — are guaranteed to be human-written, historically deep, and not already in a competitor's training set. That makes them valuable raw material.
What kinds of books are affected?+
Rare and out-of-print books available through secondary markets like Biblio. These tend to be: specialized academic and technical texts no longer in print, historical documents and case reports, legal treatises, medical literature from earlier eras, and niche professional references that were never digitized. The 404 Media investigation tracked approximately 1,000 books in one shipment; the systematic barcode-scanning behavior at the facility suggests the operation is considerably larger.
Is this legal?+
Probably yes, in current U.S. law — though it's unsettled. A prior lawsuit against Anthropic over similar training practices (using pirated digital copies, a legally different situation) established some relevant precedent, and courts have generally not found that transforming a physical book into a digital file for training constitutes copyright infringement on its own. What's clear: Amazon isn't digitizing rare books to preserve them. It's acquiring them as consumable raw material. The books are discarded after scanning.
Does this affect the quality of AI answers I get?+
Yes, indirectly. Training data composition is a real differentiator in AI output quality, even if providers rarely disclose it. AI models are better-informed on topics that are heavily covered in digitized text, and weaker on topics that live in physical archives, unpublished reports, and out-of-print specialist literature. If your professional domain has a lot of pre-digital knowledge — older legal precedents, pre-EMR medical literature, historical engineering specifications — the depth of AI answers in those areas depends partly on whether that material was part of training.
Should I change which AI tool I use because of this?+
Not based on this story alone. Training data composition is one real input to AI quality, but it's hard to evaluate directly since no major provider fully discloses their training sets. A more practical approach: use your own professional experience to calibrate when AI answers are reliable in your domain, verify claims on specialized topics against primary sources, and treat AI as a research accelerator rather than a terminal authority on niche subjects.
What's the broader implication for AI training going forward?+
AI companies are running out of fresh, human-generated digital text to train on. Physical books — especially rare, pre-digital ones — represent one of the last large-scale reserves of text guaranteed to be human-written. As these archives are consumed and the supply of undigitized rare books dwindles, the training data problem intensifies. This is partly why there's growing pressure for AI labs to reach licensing agreements with publishers, libraries, and archives rather than acquiring physical copies through secondary markets.
Related Guides
A Third of New Web Pages Show AI Authorship, Pew Finds — What Professionals Need to Know About Credibility
Pew Research Center analyzed 490,000 web pages and found more than a third of content published since ChatGPT's launch shows signs of AI authorship. Commercial sites run ten times the AI-text rate of .edu or .gov pages. Here's what the shift means for professional credibility, disclosure, and standing out.
9 Workplace Monitoring Apps All Share Your Data. Here's Which Ones — and What Gets Sent.
A joint study from Columbia Law School, Northeastern, Vanderbilt, and UC Berkeley tested nine widely-used 'bossware' platforms — Hubstaff, Time Doctor 2, Deputy, and six others — and found all of them share worker names, emails, and employer info with Facebook, Google, Microsoft, and Yandex. Here's exactly what gets sent, where it goes, and what professionals can do today.
90% of Executives Say AI Hasn't Boosted Productivity — and AI Layoffs Are Making It Worse
A Federal Reserve survey of ~750 executives finds 90% say AI hasn't boosted productivity at their companies. A University of Pittsburgh study explains why: companies cutting jobs to offset AI costs are creating the employee resistance that kills the productivity gains they're chasing.