The AI industry is buying up books and faces: why machines still need humans

The ISBNdb service, which holds metadata for 113 million publications, has offered AI labs an unusual service: bulk purchasing of books in batches of up to a million copies, followed by digitization and disposal. After the scheme came to journalists' attention, the company hastily removed the promo page and stated that no deals had been made, calling it merely a "test of market demand."
Meanwhile, another trend is gaining momentum in the industry: the ActID and New Claw platforms pay people from $15 to $700 for the right to use their faces in AI series and advertising. Let's figure out why neural networks critically need "human" content and what model collapse is.
The Curse of Recursion
The term "model collapse" describes the degradation of an AI system trained on data generated by other neural networks. With each cycle, results become increasingly averaged and predictable, losing touch with reality. This phenomenon was first formalized by an international group of researchers led by Ilya Shumailov in May 2023, publishing the paper "The Curse of Recursion."
The scale of the problem is confirmed by numbers: in April 2025, Ahrefs analysts examined 900,000 new pages in Google search and found that more than 74% of materials showed signs of auto-generation. According to Epoch AI estimates, all usable text for training on the internet could run out by 2026–2032. This is why books published before 2022 are becoming a gold mine for AI companies—they are one of the few sources guaranteed to be free of "synthetic junk."
Books Under the Knife
The practice of buying up and destroying printed publications has already moved beyond individual startups. In August 2024, three American authors filed a class-action lawsuit against Anthropic, accusing the company of copyright infringement. It emerged that the corporation had invested tens of millions of dollars in the "unofficial" Panama project: buying up print runs, cutting off bindings with hydraulic knives, and scanning books on an industrial scale. In June 2025, Judge William Alsup ruled that such digitization falls under the doctrine of fair use.
ISBNdb, which operated for over two decades as the largest metadata database, offered AI labs wholesale brokerage services: purchasing books in batches from 1,000 to 1 million copies, guarantees of "purity" from AI slop, and legal protection through NDAs. The platform's management openly stated: "An AI company destroyed two million books—that's not the headline that will win public sympathy." Used-book dealers confirmed anomalous purchases starting in April 2026—one trader's sales grew from twenty copies to several hundred per week.
The gap between the cost of raw material and the finished product is striking: a book destroyed immediately after scanning costs between $2 and $5 on the secondary market, while a model trained on millions of such copies is valued in the billions of dollars.
Renting Faces
In China, the AI content market is growing rapidly: more than 95% of the 128,000 AI mini-dramas released in the first three months of 2026 were created using artificial intelligence. Platforms like ActID and New Claw are structured as marketplaces: companies pay people from $15 to $700 for licensing their images. Users upload photos taken in special studios, and producers select faces by categories such as "girl next door" or "brutal."
However, control over one's own appearance proves fragile. British citizens who sold rights to digital avatars to the startup Synthesia found their faces in propaganda videos linked to Nicolas Maduro's government and Burkina Faso's dictator Ibrahim Traore. Beijing lawyer Ile Deng notes: licensing terms are often so vague that it's impossible to understand who will ultimately use the likeness and how.
Don't Forget the Receipt
The scheme is the same everywhere: texts, faces, voices, movements—all of it becomes raw material for AI. Previously, data was taken without permission; now people are paid for it. But the structure of the deal doesn't change: the asset irreversibly leaves the owner's control, just now with a receipt.
My conclusion: We are witnessing the formation of a new class of assets—"human data"—where value is determined not by content but by suitability for training. As long as the AI industry depends on "organic" content, this dependence will only intensify. But once models learn to self-generate quality data without collapse, the value of human contribution could rapidly plummet to zero. The only question is whether the legal framework will adapt in time before we finally lose control over our own digital essence.