Crypto news

08.08.2026
01:30

The AI industry is buying up books and faces: why machines still need people

img-bbaf321c89b70f44-1766402574726061

While AI labs report progress in model training, a race is unfolding in the background for a resource that cannot be synthesized. This is about "clean" data created by humans. And the demand for it is taking increasingly bizarre forms—from wholesale purchases of printed books with subsequent disposal to renting human faces for digital avatars.

The service ISBNdb, which holds metadata for 113 million publications, offered AI companies a specific service: purchasing books in batches of up to a million copies with subsequent digitization. When the scheme became public knowledge, the company hurried to remove the promotional page and stated that it was only a "test of market demand." In a parallel segment, platforms ActID and New Claw are already paying people from $15 to $700 for the right to use their faces as the basis for characters in AI-generated series.

Model Collapse: When AI Begins to Devour Itself

The phenomenon known as "model collapse" is the degradation of an AI system trained on data generated by other neural networks. With each cycle, the result becomes increasingly averaged out and loses connection with reality. The term was formalized by an international group of researchers led by Ilya Shumailov in May 2023, mathematically proving that when synthetic data is mixed in, models first lose the ability to reproduce rare events and then accumulate errors.

The scale of the problem is obvious. In April 2025, researchers at Ahrefs examined 900,000 newly indexed pages in Google—more than 74% showed signs of auto-generation. According to Epoch AI estimates, all usable text for training on the internet could run out by 2032. Therefore, books published before 2022 have become a golden asset: they are guaranteed to be free of synthetic junk and "poisoned" data.

Separately worth noting is the legal precedent of June 2025, when Judge William Alsup ruled that digitizing legally purchased printed books for training language models falls under the fair use doctrine. This decision opened the floodgates: now it is enough to keep a receipt for a paper book to document the origin of the data.

Faces as Commodities: The New Economy of Digital Avatars

The AI content market in China is showing explosive growth. More than 95% of the 128,000 AI micro-dramas released in the first three months of 2026 were created using artificial intelligence. Platforms like ActID and New Claw have turned appearance rental into a full-fledged business: from $15 to $700 for image licensing, with a 10% platform commission.

However, control over the use of digital clones remains fragile. Indicative are the cases of British citizens who sold rights to their avatars to the startup Synthesia, whose digital doubles then "starred" in propaganda videos supporting the dictators of Venezuela and Burkina Faso. The Equity union is already pushing for a legislative ban on transferring uncontrolled rights to digital likenesses.

The scheme is the same everywhere: texts, faces, voices, movements—all of this becomes raw material for AI. Previously, data was taken without asking; now people are paid for it. But the essence of the deal does not change: the asset irreversibly leaves the owner's control, only now with a receipt.

My comment: The AI industry faces a fundamental paradox: the more models are trained on synthetic data, the faster they degrade, and the more valuable "human" data becomes. But the market for this data is spontaneous and legally immature. Until a transparent licensing system with clear context-of-use restrictions is created, we will continue to see new scandals involving the loss of control over digital copies of identity. This is a systemic risk that the industry prefers to ignore.