An investigation has uncovered how Amazon allegedly destroyed rare and out-of-print books to use their contents as training data for artificial intelligence systems, with the process traced using an AirTag placed in a shipment.

What Happened

According to The Decoder's reporting, researchers or investigators placed an Apple AirTag inside a package of physical books being shipped to what they believed was a facility involved in AI training data collection. The tracking device allowed them to follow the books' journey and document their destruction rather than resale or preservation. The investigation reportedly found that physical volumes—including rare and out-of-print titles—were being destroyed so their text could be digitized and fed into machine learning models, raising questions about how the AI industry sources its training data.

Why It Matters

The report highlights an emerging ethical concern in the AI industry's appetite for training data. As frontier labs seek ever-larger text corpora to improve language models, physical books—many of which are rare or unavailable digitally—represent a potential source that may be accessed through destruction rather than licensing. For publishers, archivists, and authors' rights holders, the practice raises questions about intellectual property in an era when training data sourcing remains largely opaque.

The Bottom Line

The investigation adds to ongoing scrutiny of how AI companies obtain the vast amounts of text used to train large language models, a process that has already prompted lawsuits over copyright. The use of AirTag tracking provides a new investigative tool for documenting supply chains involved in AI data preparation.