Amazon is facing criticism over reports that it used rare and out-of-print books from library collections to train artificial intelligence systems, according to an investigation published by TechCrunch.
What Happened
The report alleges that Amazon digitized physical texts—including materials from specialized libraries and archives—to use as training data for its AI models. The company reportedly processed these documents without securing proper licensing or permissions from rights holders. The practice drew comparisons to a full-circle moment for the e-commerce giant, which originally built its business selling books online before expanding into cloud computing and AI services.
Why It Matters
The controversy highlights ongoing tensions in the AI industry over how training data is sourced. Publishers and authors have increasingly pushed back against tech companies using copyrighted materials without compensation or consent. For developers and enterprises building on AI platforms, questions about data provenance affect legal risk and long-term sustainability of AI applications. Libraries and archives that preserve rare collections also face pressure to prevent their holdings from being scraped for commercial AI development.
The Bottom Line
Amazon has not publicly detailed its training data sources in full. The company declined to comment on specific allegations in the TechCrunch report. The episode is likely to fuel ongoing policy debates about fair use, copyright, and how AI companies should compensate creators whose work helps train their systems.