Books traced to an Amazon AI training facility are said to be scanned and then destroyed, according to two outlets.
Amazon has reportedly acquired rare books for use in training its artificial intelligence systems, according to reports from CryptoBriefing and Decrypt. Both outlets say the books were traced to a facility connected to Amazon's AI training operations. Once there, the physical copies are said to be scanned and then destroyed.
Neither report specifies which titles were involved or how many books have gone through this process. The reports also do not detail the scale of the purchasing effort or how long it has been underway. Still, the underlying pattern described, buying physical books, digitizing their contents, then discarding the originals, points to a broader scramble across the AI industry for usable training data.
Large language models require enormous volumes of text to improve their outputs. Publicly available internet text has increasingly been exhausted or restricted by copyright holders. That has pushed AI developers toward alternative sources, including physical archives, out-of-print titles, and rare editions that may not exist in digital form anywhere else.
Rare books present a particular appeal for this purpose. Many predate modern copyright regimes or contain text unavailable through standard digital licensing deals. Scanning such volumes can yield unique training material. But destroying the physical originals afterward would remove any remaining scarcity or collectible value tied to those specific copies.
The reports do not indicate whether Amazon disclosed this destruction process to sellers at the time of purchase, or whether any legal or preservation obligations applied to the specific books involved. It is also unclear from the available reporting whether the books were sourced through public auctions, private dealers, or other channels.
The story touches on a wider tension in the AI industry. Companies are under pressure to find fresh, high-quality training data while facing growing scrutiny over data provenance, copyright, and preservation. Book collectors, librarians, and archivists have previously raised concerns when physical materials are treated as disposable inputs once their informational content is captured digitally.
Amazon has not been reported as issuing a detailed public statement addressing the scanning and destruction process described in these accounts. Additional details about which specific facility is involved, the volume of books processed, and Amazon's internal policies on the practice have not been made available in the current reporting.
The reports carry limited direct financial market impact, since they concern a physical book supply chain rather than publicly traded assets or crypto tokens. Any indirect effect would likely be reputational, touching Amazon's standing among archivists, rare book dealers, and AI ethics observers rather than its share price or cloud and AI business lines directly.
For the broader AI industry, the story underscores how aggressively companies are pursuing novel training data sources as easily accessible text becomes scarcer. Investors and analysts tracking AI infrastructure spending may view this as another data point in the ongoing search for proprietary or unique training material, a factor that can influence how AI firms are valued relative to their access to differentiated data.
The reports describe a niche but notable practice in AI data sourcing, one that intersects with questions about preservation, provenance, and the treatment of physical historical materials. Further details from Amazon or additional reporting would help clarify the scope and policies behind the described process.
According to CryptoBriefing and Decrypt, Amazon has been buying rare books, scanning their contents for AI training data, and then destroying the physical originals.
The reports do not explain Amazon's specific reasoning. In general, digitizing physical text for AI training does not require keeping the original copy once scanning is complete.
The available reporting does not indicate that Amazon has issued a detailed public statement addressing the scanning and destruction process described.
Rare or out-of-print books can contain text unavailable elsewhere online, making them a potentially unique source of training data as more common digital text becomes harder to license or access.
FalconX and Interstice Build Cross-Chain Bridge Linking Canton Network to Ethereum, Solana, Robinhood Chain
We measure how many people read this site. That is all it is used for — there is no ad network, no advertising cookie, and nothing sold to anyone. Decline and the site works exactly the same. What we collect