AI

Amazon Destroys Books for AI Training, Warehouse Worker Testifies to Inside Operations

404 Media reports on Amazon destroying books for AI training, the AI ghost issue in academia, and ICE's data procurement.

10 min read Reviewed & edited by the SINGULISM Editorial Team

Amazon Destroys Books for AI Training, Warehouse Worker Testifies to Inside Operations
Photo by Anirudh on Unsplash

404 Media’s weekly podcast published on September 2, 2026, highlighted the physical labor behind AI development and the social distortions that emerge as its byproducts. The program covered three topics: testimony from a worker who scanned and destroyed books at an Amazon warehouse, traces unique to AI-generated content in the form of specific personal names repeatedly appearing in academic papers, and spending by U.S. Immigration and Customs Enforcement (ICE) on voter data and robot dogs. All are cases that illustrate the pressure that the collection and use of AI and data place on existing institutional and legal frameworks.

In reporting by 404 Media’s Joseph Cox, details of Amazon’s book-scanning system were discussed in the first half of the program. The second half analyzed the phenomenon of the same names frequently appearing in large language model outputs, and the extended segment for paid subscribers explained ICE’s activities.

Inside the Warehouse Where Amazon Scans and Destroys Books for AI Training

The podcast is available on Apple Podcasts, Spotify, and YouTube, with the extended version offered to subscribers.

The Reality of Book Scanning at Amazon Warehouses

Opening the program was a follow-up report by Emanuel on Amazon’s book scanning. In response to the previous report that “Amazon scans and destroys books for use in training AI products,” this time a direct interview was conducted with a person who actually worked in that warehouse.

According to the worker’s testimony, inside the warehouse there is a process in which books are transported on conveyor belts and elsewhere, pages are digitized with high-speed scanners, and the physical books are then destroyed. The scanned data is said to be used for training Amazon’s AI products. The work was designed to process large volumes of books quickly, and workers carried out their duties aware that destruction was a prerequisite.

While this method is efficient from the standpoint of securing a large volume of training data, it raises issues surrounding the handling of publications as copyrighted works. A book is a bundle of rights involving the author, publisher, and distributor, and it is unclear how the act of scanning and destroying a purchased book is positioned within the scope of contracts and licenses. 404 Media’s reporting visualized that the front lines of AI development involve not only computational resources in the cloud, but also manual labor and physical destruction inside warehouses.

How data is collected in AI development also relates to the transparency of the development environment. As discussed in Desktop Commander MCP: Delegating Terminal Operations to AI, the importance of permission management and auditing when AI operates devices and the cloud is a factor that determines operational reliability. The book scanning process is also an area where management is questioned — under what authority what is recorded and where it is stored.

The AI Ghost Phenomenon Spreading in Academia

In the middle of the program, Emanuel and Sam addressed the phenomenon of the same small number of personal names repeatedly appearing in papers believed to be AI-generated. This issue, which 404 Media calls “AI ghosts,” is gaining attention as a new form of contamination in academic publishing.

Large language models tend to excessively reproduce specific personal names and citation patterns contained in their training data. In text generated to mimic the form of academic papers, non-existent citations and specific fictitious researcher names appear across numerous papers. If they are published without being detected during peer review or editing, the reliability of academic databases as a whole declines.

Behind this phenomenon is the spread of paper-writing assistance and automatic generation tools. While the use of AI by researchers to draft and summarize manuscripts is itself becoming more common, if erroneous citations and fictitious names contained in generated text circulate without verification, it will adversely affect subsequent research and the retraining of the AI itself. If a cycle emerges in which models relearn data that includes their own generated outputs, there is a risk that misinformation will be amplified.

Academia is already moving to strengthen peer review systems and develop methods to detect AI-generated content, but detection is not perfect. While the repeated appearance of specific names is a clue for detection, if the model or prompt changes, the form of the traces will also change. There may be growing moves on the part of publishing platforms to require disclosure at submission and to mandate explicit declaration of generative AI use.

ICE’s Strengthened Surveillance and Bulk Data

Procurement

In the paid subscriber-only segment of the program, Joseph reported on ICE’s spending plans. ICE is reportedly planning bulk purchases of data related to “voter fraud” and spending on the order of millions of dollars on Boston Dynamics robot dogs.

The former is said to be aimed at detecting fraud through the collection and analysis of voter data, but concerns have been raised about the source and scope of the data and its impact on privacy. Voter information includes sensitive personal information such as names, addresses, and voting histories. A move by a federal agency to aggregate such data on a nationwide scale would strengthen ties with the data broker market and increase the risk of misuse and harm from erroneous matching.

The latter, the robot dogs, are quadruped robots developed by Boston Dynamics. ICE is believed to be planning to use them for security, search, and reconnaissance in hazardous areas. The program reported that spending would amount to several million dollars. Equipping the robots with cameras and sensors would expand surveillance capabilities in places where human entry is difficult, but at the same time it invites public resistance to surveillance and debate over operational transparency.

ICE’s moves illustrate governance challenges when AI and robotics are introduced to law enforcement. With the parallel advance of data purchases and the introduction of physical surveillance equipment, the surveillance network is being strengthened in both digital and physical domains. Cost-effectiveness, operational standards, and the presence or absence of external audits will be the focus going forward.

The Structural Friction Between the

Publishing Industry and AI Training

The scanning involving the destruction of books by Amazon symbolizes the structural friction between the publishing industry and AI developers. Publishers and authors are increasingly concerned about their works being used without permission for AI training. In the United States, publishers and authors’ groups have filed a succession of lawsuits against AI companies, and the legality of training data is being contested in court.

Even if the scanning at the warehouse targets books legitimately purchased from publishers, whether purchase permits reproduction or derivative use for training purposes is a separate issue. Contractually, whether the act of digitizing book content and incorporating it into a model constitutes “use” is subject to differing interpretations. The worker’s testimony showed the reality that such legal gray zones are being processed as routine work on the ground.

This friction is also related to the development of technical verification environments. As the initiative presented in GNOME OS Test Center Inspired by Apple TestFlight shows, the importance of an isolated testing environment for safely trying out AI behavior is rooted in the principle of separating development and operations. For the process of collecting training data as well, a testing and auditing mechanism that can track what data was ingested and how is needed.

In addition, the problem of AI-generated content infiltrating academic publishing concerns the reliability of the publishing industry as a whole. As both books and papers — different forms of publications — are handled as training data for AI or as AI-generated products, the authenticity of primary information is shaken. Publishers have begun to put in place mechanisms for declaring refusal of use for AI training and for opting out, but the method of destroying and scanning physical books could become a route to circumvent such declarations.

What Is Being Questioned Between

Technological Innovation and Ethics

The three cases presented in this podcast all show a situation where technological progress outpaces existing systems. The scanning and destruction of books, the ghost contamination in academia, and the procurement of data and robots by law enforcement agencies differ in field, but share the common point that controls on the collection, generation, and use of data have not kept pace.

The aspect that hardware performance and price affect the speed of AI adoption cannot be ignored either. Even with familiar devices such as voice assistants and smart speakers, performance differences between generations and costs define the user experience. For example, the review New Google Home Speaker Falls Short of 6-Year-Old Nest Audio in Sound Quality shows that a new product does not necessarily mean across-the-board improvement. The quality and procurement method of training data for AI products likewise determine product value, but are difficult to see from the outside.

AI developers face pressure to increase the legality and transparency of training data, while publishers and research institutions are pressed to protect their own outputs and strengthen verification. Data purchases by law enforcement highlight the need to regulate the market for the distribution of personal information itself. In every area, not only technical feasibility but also procedural legitimacy — who uses data under what agreement — is being called into question.

404 Media’s podcast connected, in a single program, everything from the concrete現場 of work inside a warehouse to broad social institutions such as academia and law enforcement. Its significance lies in conveying, through the tangible material of a worker’s testimony, the current state in which a single technology, AI, is simultaneously affecting different layers — physical labor, knowledge production, and state surveillance.

Editorial Opinion

In the short term, we expect negotiations and disputes over licensing for training data between the publishing industry and AI companies to increase further. As illustrated by the Amazon case, as scanning involving the physical destruction of books comes to light, publishers are likely to strengthen demands for contract revisions and audits. In the next 3 to 6 months, we assess that major publishers will move to add clauses explicitly restricting use for AI training and to demand tracking of distribution channels.

In the long term, we believe the AI ghost problem in academic publishing will become a factor that shakes the foundations of research evaluation. On a 1- to 3-year horizon, we expect academic databases and peer review systems to be redesigned on the premise of detecting and eliminating AI-generated content, and that declaration of AI use at submission will become standardized. Researchers themselves will also be required to adopt new workflows to ensure the authenticity of citations, and the cost of restoring trust in academia is likely to increase.

The question from our editorial team is who and how will guarantee the legitimacy of data and surveillance. The purchase of voter data and the introduction of robot dogs by ICE raise the question of the gap between what is technically possible and what is socially acceptable.

References

Frequently Asked Questions

Why does Amazon go so far as to destroy books to scan them?
According to the testimony of the warehouse worker discussed on the 404 Media podcast, it is to digitize books at high speed and use them as training data for AI products. There is said to be a process that streamlines mass processing by destroying the physical books, and the operation is seen as prioritizing the securing of training data volume.
What is an AI ghost?
It refers to the phenomenon where the same small number of personal names and citations repeatedly appear in academic papers believed to be AI-generated. It is a term used by 404 Media, where large language models reproduce biases in training data, spreading fictitious researcher names and erroneous citations. It is attracting attention as a factor that undermines the reliability of academic databases.
What are the data and robot dogs that ICE plans to purchase?
According to 404 Media's reporting, ICE is planning to spend millions of dollars on nationwide voter data related to voter fraud and on quadruped robot dogs made by Boston Dynamics. The former involves the aggregation of data containing personal information, and the latter is envisioned for use in surveillance and searches, sparking debate from the perspectives of privacy and operational transparency.
Source: 404 Media

Comments

← Back to Home