An investigation has revealed that Amazon is buying printed books in bulk, scanning them for artificial intelligence training data and destroying the physical copies in the process. The discovery was made after 404 Media placed an Apple AirTag inside a shipment of rare books and tracked the package to an Amazon facility in Las Vegas. The findings have raised fresh concerns about how technology companies obtain large volumes of copyrighted and potentially irreplaceable material for AI development.

The investigation points to a growing demand for physical books as AI companies seek training data that may not be readily available online. Printed books can contain older material published before the widespread use of generative AI, making them potentially valuable for building datasets that are less likely to contain AI-generated text. The controversy, however, goes beyond copyright questions because the scanning process can permanently destroy physical copies, including books that may have historical or cultural value.

AirTag Traced Rare Books to Amazon Facility

The investigation began when an AirTag was placed inside a shipment of rare books.

The tracking device followed the package across the United States before it eventually arrived at an Amazon warehouse in Las Vegas.

The facility is associated with a team known as VGT3, whose logo features a Tyrannosaurus rex holding a book.

According to the investigation, workers at the facility cut the spines off books to make them easier to process through high-speed scanning equipment.

Once the spine is removed, the original physical book can no longer be preserved in its original form.

The scanned material can then be converted into digital data that can be processed and used in AI development.

Amazon Uses Scanned Data for AI Models

The investigation reported that Amazon uses the scanned data to train its Nova family of AI models.

Amazon said it purchases books through commercial channels to improve its products.

The company did not characterize the practice as an effort to destroy rare books for its own sake. Instead, the scanning process is part of obtaining large quantities of text that can be processed digitally.

The distinction is important because the physical destruction appears to be a consequence of the scanning method rather than the primary objective.

However, for books that are rare or difficult to replace, the outcome is still significant.

Why Printed Books Are Valuable for AI Training

Printed books offer AI developers a potentially important source of high-quality text.

Large quantities of digital information are already available online, but the internet has several limitations as a training-data source.

Web content can be duplicated, modified or removed. It can also contain large amounts of low-quality material, spam and increasingly AI-generated text.

Older printed books provide a different type of dataset.

Many were published before generative AI became widespread and may never have been digitized or made freely available online.

Why AI Companies Want Physical Books

Printed books

Large amounts of structured text

Older information not always available online

Less exposure to AI-generated content

Potentially valuable training data

Digitization

AI training datasets

The result is increasing interest in physical books as a source of training material.

Booksellers Suspect Systematic Scanning

Booksellers have raised concerns that AI companies or intermediaries could be attempting to systematically acquire books based on their ISBN numbers.

An ISBN identifies a specific book edition, making it possible to search for particular titles and editions across booksellers and marketplaces.

If companies systematically purchase books according to ISBNs, they could potentially build large collections covering a broad range of published material.

That would allow them to create extensive digital datasets from books that may otherwise be difficult to access electronically.

For booksellers, unusually large purchases can therefore appear very different from normal consumer demand.

Rare Books Create a Bigger Ethical Problem

The destruction of ordinary mass-market books would already raise questions about waste and data acquisition.

The situation becomes more complicated when rare or collectible books are involved.

Some books have historical value beyond their text.

Their physical condition, illustrations, typography, annotations, bindings and provenance can make them valuable to collectors, researchers and libraries.

Once a rare book is cut apart for scanning, the physical artifact may be permanently lost.

This creates a tension between the value of the information contained in a book and the value of the physical object itself.

AI Training Can Create a Conflict Between Access and Preservation

AI companies need enormous quantities of data to train increasingly capable models.

Books offer structured, information-rich material.

But digitizing books at industrial scale can create difficult questions about preservation.

A company may view a book primarily as a source of text.

A collector, historian or library may view the same book as an irreplaceable cultural artifact.

The disagreement is therefore not simply about whether information can be copied.

It is also about whether the original physical object should be destroyed to create a digital training resource.

Amazon Is Not the Only AI Company Involved

The practice is not limited to Amazon.

Anthropic has also faced scrutiny over a similar book-digitization operation known as Project Panama.

A lawsuit brought by authors revealed that Anthropic purchased books through online marketplaces, removed their spines and digitized them.

The case raised questions about whether such copying was permissible under copyright law.

A judge ultimately ruled that the scanning process qualified as fair use and did not violate copyright in the circumstances at issue.

One factor considered was that the physical originals were destroyed rather than retained and redistributed.

The Legal Question Is Different From the Preservation Question

A legal ruling that a particular digitization process qualifies as fair use does not necessarily resolve the broader ethical debate.

Copyright law primarily addresses questions surrounding the use and reproduction of protected works.

Preservation raises a different set of concerns.

A book may legally be scanned under certain circumstances while its physical destruction still creates concerns for collectors, historians and libraries.

This distinction could become increasingly important as AI companies seek larger and more comprehensive training datasets.

AI Companies Are Looking for Data Beyond the Internet

The development of large language models has created enormous demand for training data.

Initially, much of the focus was on publicly accessible internet content.

As AI models have become more capable, however, companies have increasingly looked for additional sources of high-quality information.

Books, academic publications, specialized databases and other structured content can provide training material that is difficult to obtain from general web crawling alone.

This has made previously overlooked physical archives potentially valuable.

The Rise of AI Increases the Value of Old Information

The AI industry’s interest in older books also highlights an unexpected consequence of generative AI development.

Material created decades ago can become strategically valuable because it represents human-generated knowledge that predates the current wave of AI-generated content.

Older books can therefore provide models with language, ideas, historical information and specialized knowledge that may not be easily reproduced from modern web content.

The demand for such material could increase as AI companies compete to build more comprehensive datasets.

Publishers and Authors Could Face New Questions

The acquisition and scanning of physical books also raises questions for publishers and authors.

Even when a company purchases a physical copy legitimately, the subsequent digitization and use of its contents for AI training can create legal and commercial questions.

The central issue is whether purchasing a physical book gives a company sufficient rights to reproduce its contents digitally for training purposes.

Different jurisdictions and individual cases may produce different answers.

The rapid development of AI is forcing copyright law to confront these questions at unprecedented scale.

Physical Books Could Become Part of the AI Supply Chain

The investigation suggests that the AI industry’s supply chain may extend far beyond GPUs, data centers and cloud infrastructure.

Books and other physical information sources can also become inputs into AI systems.

A simplified supply chain can look like this:

Book acquisition

Physical shipment

Industrial scanning

Digital text extraction

Data processing

AI training

AI model

The process turns physical knowledge into machine-readable data.

The Destruction Problem Could Become Larger

If AI companies continue expanding their training datasets, demand for physical books could grow.

That could increase the number of books being purchased and processed at scale.

The key question will be whether companies develop methods that preserve physical copies while digitizing their contents.

Non-destructive scanning technologies already exist, but they can be slower or more expensive than cutting books apart and scanning individual pages.

For rare materials, however, preservation may justify the additional cost.

Libraries Could Become More Important

Libraries and archives contain enormous collections of books and other documents.

Many institutions have already undertaken digitization projects designed to preserve information while keeping original materials intact.

The growing demand for AI training data could make these collections even more valuable.

At the same time, libraries may face pressure over access rights, licensing and ownership of digitized collections.

The AI era could therefore create new debates around how cultural institutions manage digital access.

AI Training Is Becoming an Industrial Process

The scale of modern AI development is changing how training data is acquired.

Instead of researchers manually assembling datasets, companies can build industrial processes designed to acquire, scan, clean and organize information at enormous scale.

The book investigation illustrates how physical information can become part of that industrial pipeline.

What was once a one-off research task can now become a large-scale commercial operation.

The Debate Over Training Data Is Expanding

Disputes over AI training data have traditionally focused on web scraping, copyrighted images, books and code.

The Amazon investigation adds another dimension: what happens to the physical source material after it has been digitized?

This question could become increasingly relevant as AI companies look beyond easily accessible digital content.

The debate may eventually involve not only copyright but also cultural preservation, environmental costs and ownership of historical materials.

What This Means for AI Companies

For AI companies, physical books offer potentially valuable training material.

But acquiring them at scale creates several challenges.

Companies may need to consider:

  • Copyright compliance
  • Data provenance
  • Preservation
  • Acquisition costs
  • Scanning methods
  • Public perception
  • Training-data transparency

The controversy demonstrates that data acquisition strategies can create reputational risks even when companies believe their activities are legally defensible.

What This Means for Booksellers

Booksellers could increasingly become unexpected suppliers to the AI industry.

Large purchases of specific titles could create new demand for inventory.

However, sellers may also face ethical questions about where their books ultimately end up.

A rare book that would normally be sold to a collector or researcher could instead be purchased for industrial digitization.

That could change the economics of certain segments of the used-book market.

What This Means for Collectors

Collectors could face a different kind of competition if AI companies begin purchasing rare books at scale.

Higher demand could increase prices for certain editions.

At the same time, collectors may become more concerned about preserving scarce books before they disappear into industrial scanning operations.

The value of physical books could therefore increase precisely because digital AI systems are creating demand for their contents.

Key Facts at a Glance

MetricDetails
Company involvedAmazon
Investigation404 Media
Tracking deviceApple AirTag
DestinationAmazon facility in Las Vegas
Facility/teamVGT3
Reported purposeBook scanning for AI training
AI models mentionedAmazon Nova
Reported scanning methodRemoving book spines
Other AI company citedAnthropic
Anthropic projectProject Panama
Main issueDigitization and destruction of physical books
Broader concernCopyright, preservation and AI training data

Infographic: How Books Can Become AI Training Data

RARE OR PRINTED BOOKS

PURCHASED THROUGH COMMERCIAL CHANNELS

SHIPPED TO PROCESSING FACILITY

BOOK SPINE REMOVED

PAGES SCANNED

TEXT DIGITIZED

DATA PROCESSED

AI TRAINING DATA

AI MODEL

+

KNOWLEDGE RETAINED DIGITALLY

BUT

PHYSICAL ORIGINAL

MAY BE DESTROYED

KEY DEBATE

AI DEVELOPMENT

VS.

CULTURAL PRESERVATION

The Bigger Picture

The investigation into Amazon’s book-buying and scanning practices highlights a new dimension of the race to build increasingly capable AI systems. As companies search for large volumes of high-quality training data, physical books are becoming valuable sources of information that may not exist online and may predate the widespread creation of AI-generated content. The use of an AirTag to trace a shipment of rare books to an Amazon facility in Las Vegas has brought attention to how industrial-scale digitization can transform physical knowledge into private AI training material.

The controversy also demonstrates that the debate over AI training data extends beyond copyright. When books are cut apart for scanning, the original physical artifacts can be permanently lost, potentially affecting collectors, researchers and cultural institutions. Amazon is not the only company facing questions around this practice, with Anthropic’s Project Panama having involved a similar process of purchasing, cutting and digitizing books. As AI companies continue to seek larger datasets, the industry may face growing pressure to develop methods that preserve valuable physical works while still making their information available for machine learning.

Looking Ahead

The immediate focus will be on how Amazon and other AI companies respond to concerns about the acquisition and destruction of rare books. Greater transparency around where training data comes from, how physical materials are processed and what safeguards exist for historically significant works could help address some of the criticism. Publishers, authors, booksellers, collectors and technology companies are likely to remain involved in the debate as courts and regulators continue to examine AI training practices.

Over the longer term, the controversy could encourage the development of more preservation-friendly scanning systems and clearer licensing frameworks for AI training data. The growing value of physical books as AI inputs could also create unexpected changes in the publishing, used-book and archival markets. The central challenge will be finding a balance between the enormous demand for information needed to develop AI and the responsibility to preserve the physical record from which that information comes.

Get the day’s top stories in your inbox

One concise email. No spam, unsubscribe anytime.