The AI industry is buying up books and faces: why neural networks desperately need "human" raw materials

The artificial intelligence training market is undergoing a tectonic shift. Major labs and startups, facing a shortage of quality data, have begun hunting for "organic" raw materials—books, faces, and voices. This is not just about licensing, but about large-scale operations to buy up and destroy printed publications. This race for clean content is a direct consequence of a phenomenon known as "model collapse."
The Curse of Recursion
The term "model collapse" describes the degradation of a neural network trained on data generated by other AIs. Each cycle makes the model more averaged and detached from reality. This phenomenon was first formalized by a group of researchers led by Ilya Shumailov in May 2023, publishing the paper "The Curse of Recursion." They mathematically proved that mixing in synthetic data leads to the loss of rare knowledge and the accumulation of errors.
The problem is no longer theoretical. Based on web analytics data, I estimate that over 74% of new content on the internet is generated automatically. Meanwhile, according to Epoch AI calculations, reserves of "clean" text suitable for training could run out by 2026-2032. This is why books published before 2022 have become a strategic resource—they are almost guaranteed to contain neither synthetic junk nor the deliberate "traps" that authors have begun embedding to "poison" data.
The Bookish Ouroboros
The service ISBNdb, which holds metadata for 113 million publications, has offered AI companies a specific service: purchasing books in batches of up to a million copies, followed by digitization and disposal. When the media reported on this scheme, the company hurried to remove the page with the description and stated that no deals had been made, calling it a "test of market demand." However, independent used-book dealers confirm an anomalous surge in purchases since April 2026. Buyers are not interested in genre or value—only the presence of an ISBN. This reminds me of the precedent with Anthropic, which, under "Project Panama," anonymously bought up print runs, cut off bindings, and scanned books. Judge William Alsup ruled in June 2025 that digitizing legally purchased books falls under the fair use doctrine. Now it is enough to keep the receipt—and the origin of the data is legally confirmed.
The economics of this business are absurd: a book on the secondary market costs $2-5, while a model trained on millions of such copies is valued in the billions of dollars.
Renting Faces
In another market segment, platforms like ActID and New Claw pay people from $15 to $700 for using their faces as the basis for characters in AI series. Demand is enormous: over 95% of the 128,000 AI mini-dramas released in China in the first quarter of 2026 were created using neural networks. However, control over biometric data is illusory. There are already precedents where digital clones of British actors who sold rights to Synthesia were used in propaganda videos supporting dictatorial regimes. Beijing lawyer Ile Deng rightly notes: licensing terms are so vague that it is impossible to understand who ultimately uses your appearance and how.
My Verdict
We are witnessing the formation of a new data economy where human experience becomes a commodity. For now, the market is moving from theft to payment, but the deal structure remains exploitative: the asset irrevocably leaves the owner's control, just now with a receipt. This is a dangerous trend that will require strict regulation, otherwise we risk losing not only control over our data, but also the very diversity of information that makes AI truly useful.