The GIANT Paper Mill
Why tech giants are buying, scanning, and pulping physical books for AI data

Published: 25.8.2026 | photo / Video: AI generated, Magnific
The battle over artificial intelligence training data has extended from digital web scrapers into physical print archives. How publishers, courts, and legislators respond will determine not only the economic future of creators, but also how human knowledge is preserved and commercialized in the algorithmic age.
Across the United States, Europe, and Asia, an unsettling rush has taken hold of the antiquarian and second-hand book trade. Independent booksellers, antiquarians in Frankfurt and Berlin, and used-book wholesalers are suddenly receiving massive, targeted bulk orders. Thousands of obscure, out-of-print non-fiction books are being packaged into crates, shipped to automated logistics facilities, sliced open, scanned page by page, and subsequently destroyed.
The buyers behind these operations are automated purchasing agents acting on behalf of Silicon Valley's largest AI developers. Having exhausted easily available digital text datasets pre-dating the generative AI boom, tech giants have turned their focus to physical print runs. By acquiring and digitizing physical volumes, companies seek high-quality human language free of synthetic AI "slop" to train Large Language Models (LLMs).
Part I: The Anthropic case & the evolution from digital piracy to print scanning
To understand why tech giants are buying physical print stock today, one must examine how AI training began years ago—and the multi-billion-dollar legal trap that forced Silicon Valley to pivot its data acquisition strategies.
In the early stages of generative AI development, tech companies relied heavily on digital repositories. In the landmark lawsuit Bartz v. Anthropic, legal filings revealed that Anthropic's initial LLM models (such as Claude) were trained using millions of pirated digital e-books downloaded without authorization from illicit "shadow libraries" like Library Genesis (LibGen) and Pirate Library Mirror (PiLiMi).
When authors sued Anthropic for massive copyright infringement, U.S. District Judge William Alsup issued a historic split ruling that transformed AI copyright law:
Model training as Fair Use: Judge Alsup ruled that using text content to train LLMs is inherently "transformative" and constitutes legal fair use—provided the training source was lawfully acquired.
Piracy as straight infringement: However, downloading and storing over 7 million pirated e-books in a central corporate library was ruled unlawful copyright infringement. Facing statutory damages that could have reached tens of billions of dollars, Anthropic settled the lawsuit, agreeing to a historic $1.5 billion settlement to pay around 500,000 authors and publishers (~$3,000 per title) and certifying the total destruction of its pirated digital library.
The legal paradox & "Project Panama"
This split decision created a profound legal paradox that reshaped Silicon Valley's operational pipeline: Digitizing a lawfully purchased physical book and destroying the original is legally safer than storing a digital copy downloaded from the web.
Under the U.S. First Sale Doctrine, a buyer who lawfully purchases a physical copy owns the right to scan or destroy that specific physical unit. Because destroying the physical volume ensures no unauthorized "extra copy" remains in circulation, courts treated the resulting digital scan as a permissible replacement rather than piracy.
Internal court documents revealed that Anthropic anticipated this legal distinction by launching "Project Panama"—a multi-million-dollar initiative aimed at acquiring, scanning, and pulping physical books on a massive scale. As internal documents bluntly stated, Project Panama was designed to "scan and destroy all the books in the world" to secure clean, un-scraped human prose without digital copyright infringement liabilities.
AirTags and industry realities
To investigate this secretive supply chain, investigative outlet 404 Media placed an Apple AirTag inside a 1,000-book order placed with a cooperative bookseller. The tagged crate travelled to an Amazon logistics warehouse in Las Vegas equipped with high-speed scanning infrastructure—strikingly featuring a facility logo of a T-Rex chomping down on a book. When asked for comment by DER SPIEGEL, Amazon acknowledged buying books through commercial channels to improve customer services. Anthropic similarly confirmed that physical book acquisition is standard industry practice, though it maintained it does not target rare heritage items.
Trade reports from Germany, the UK, and the US confirm that purchasing agents focus on non-fiction with ISBN numbers assigned after 1972 (published between the 1970s and 1990s). Antiquarians in Frankfurt, Berlin, and Munich report orders for niche items such as decades-old regional dining guides, local village chronicles, geological studies of the Harz mountains, 18th-century African agricultural implement guides, and 1950s racing driver biographies.
The outrage machine vs. publishing realities
Viral sensationalism on TikTok: Videos on TikTok and Instagram garnering millions of views frame bulk scanning as modern "book burning" or invoke Orwellian "memory holes" from 1984 and Fahrenheit 451. Industry experts note these claims are heavily exaggerated for engagement.
Industry scrap scale: In the US alone, approximately 320 million books (around 1 million tons) are landfilled or pulped every year during routine inventory purges. As publisher Anne Trubek observed, disposing of unwanted stock is standard practice.
Unlicensed reproduction litigation: Authors and publishers continue to file copyright infringement suits over digital scraping. In Germany, author Marc-Uwe Kling and Carlsen Verlag launched legal proceedings against OpenAI after ChatGPT accurately reproduced substantial excerpts from Das NEINhorn.
Part II: How publishers and authors can fight back
As physical scanning expands into an industry-wide supply chain, publishers, author associations, and trade groups are organizing legal, commercial, and technical countermeasures:
Wholesale & distribution Opt-out mechanisms
In January 2026, wholesale giant Ingram Content Group issued a landmark directive allowing publishers to opt out of having their catalogs sold to AI companies. In Germany, the Börsenverein des Deutschen Buchhandels is actively advocating for technical and legal safeguards:
TDM Opt-out via ONIX: The Börsenverein has published detailed guidelines (2025/2026) instructing publishers to embed machine-readable TDM (Text and Data Mining) reservations in ONIX metadata and EPUB/PDF files. This opt-out, required under § 44b Abs. 3 UrhG, signals to AI crawlers that content may not be used for training.
Position statement: In response to inquiries about bulk antiquarian book purchases, Börsenverein spokesperson Thomas Koch stated: "In unseren Augen ist solch eine Trainingspraxis ein klarer Verstoß gegen das Urheberrecht." The association has called for regulatory transparency and stronger enforcement of existing copyright exceptions.
Fokustage & Industry education: In June 2026, the Börsenverein hosted "Fokustage" in Frankfurt, including sessions on implementing TDM reservations and protecting content from LLM scraping. The emphasis is on technical opt-out mechanisms, not distribution controls akin to Ingram's model.
EU AI Act compliance: The Börsenverein is preparing publishers for the August 2026 EU AI Act, which mandates labeling of AI-generated content. However, a February 2026 study found only 28% of German publishers currently label AI use, and 37% have not implemented TDM opt-outs.
The Börsenverein has not yet launched a collective licensing framework comparable to the UK's PLS/CLA model (which has >250 publishers opted in as of June 2026). Discussions are ongoing, but no operational German equivalent exists as of August 2026.
Contractual end-use restrictions
Publishers are incorporating restrictive covenants into direct sales, bulk invoice terms, and commercial distribution contracts. By explicitly stipulating that physical volumes are sold solely for human reading and prohibiting mechanical scanning or AI ingestion, publishers create actionable breach-of-contract claims that bypass First Sale protections.
Strategic ebook DRM & licensing enclosures
Unlike physical books, e-books and digital assets are governed by licensing agreements (EULAs). Shifting institutional, academic, and backlist titles toward tightly controlled digital platforms with robust Digital Rights Management (DRM) and dynamic watermarking prevents unauthorized automated extraction.
Collective licensing frameworks & data cooperatives
Trade organizations are developing collective AI licensing structures modeled on music rights societies. In the UK, Publishers' Licensing Services (PLS) and the Authors' Licensing and Collecting Society (ALCS) launched a pioneering collective AI licensing scheme in early 2026, with over 250 publishers opted in by June. Offering tech developers legal, high-resolution, pre-cleared text corpora at standard commercial rates eliminates the economic rationale for buying, shipping, and pulping physical books.
In Germany, the Börsenverein is discussing similar models but has not yet implemented an operational framework.
Legislative pressure & mandatory transparency
Publishers and authors are intensifying political pressure on European and U.S. lawmakers to enact strict transparency obligations. Proposals include requiring AI companies to publish complete logs of all training inputs and reforming First Sale doctrines so that converting physical works into machine-learning weights without authorization is explicitly classified as copyright infringement.
As of August 2026, no concrete legislative proposals have been introduced in the U.S. Congress or the European Parliament, though advocacy groups continue to lobby for reform.