Crypto news

08.08.2026
00:35

The AI industry is buying up books and faces: why machines still need humans

img-bbaf321c89b70f44-1766402574726061

The ISBNdb service, which aggregates metadata for 113 million printed publications, has offered AI labs an unusual service: purchasing books in batches of up to a million copies, followed by digitization and disposal. After the scheme came to the attention of journalists, the company hastily removed the page with the description and stated that it was merely a "test of market demand," and that no deals had been made.

Meanwhile, in another segment of the platform, services like ActID and New Claw offer people from $15 to $700 for licensing their faces for AI characters in series and advertising. Let's figure out why labs are forced to hunt for "organic" data and what else human-like machines will require.

The Curse of Synthetics

The term "model collapse" describes the degradation of an AI system trained on data generated by other neural networks. With each cycle, the results become increasingly averaged, predictable, and lose touch with the reality that the LLM is supposed to reflect. This phenomenon was first formalized by an international group of researchers led by Ilya Shumailov in May 2023, publishing the paper "The Curse of Recursion" in Nature. They mathematically proved that with regular mixing in of synthetic data, the model first forgets rare events and subtle patterns, then accumulates errors and reduces the diversity of responses.

The scale of the problem is confirmed by numbers. In April 2025, Ahrefs analysts checked 900,000 new pages in Google search and found that more than 74% of content shows signs of auto-generation. According to Epoch AI estimates, all usable training text on the internet could run out by 2026–2032. Books published before 2022 are becoming one of the few sources for AI companies that are guaranteed to be free of synthetic junk and embedded traps. Some authors even resort to "poisoning" data by hiding hidden commands in the text. Research by the UK AI Safety Institute, Anthropic, and Oxford showed that LLMs can be covertly trained on malicious patterns using just 250 "poisoned" documents.

Books Under the Knife

In August 2024, three American authors filed a class-action lawsuit against Anthropic, accusing the company of copyright infringement in training Claude models. According to court documents, Anthropic invested tens of millions of dollars in the Panama project: anonymously buying up print runs of books, cutting off bindings, and scanning them on an industrial scale, followed by disposal. In June 2025, Judge William Alsup ruled that digitizing legally purchased books for training language models falls under the fair use doctrine. Now it's enough to keep the receipt—and the origin of the data is documented.

This practice has already extended beyond individual startups. ISBNdb, which has operated for over two decades as the largest metadata database for printed books, has turned the buying up and destruction of print runs into a turnkey service. On July 21, 2026, journalist Emanuel Maiberg discovered a promo section on the company's website offering AI labs:

  • purchase of books published up to 2022 in batches from 1,000 to 1 million copies;
  • guarantees of "purity" from AI slop and poisoning algorithms;
  • legal protection through NDAs and the legality of purchases on the secondary market.

The platform's management openly stated: "The perception problem is real. 'AI company destroyed two million books' is not the kind of headline that will win public sympathy." A survey of independent used-book dealers confirmed anomalous purchases since April 2026: one dealer's sales grew from twenty copies to several hundred per week. Buyers were interested in neither genre nor value—only the presence of an ISBN. Unique editions are meanwhile lost irreversibly, going under the scanner instead of onto a bookshelf.

After the investigation was published, ISBNdb removed the promo page and issued a statement denying any deals. However, the gap between the cost of raw material and the finished product is telling: a book worth $2–5 on the secondary market turns into a model valued at billions of dollars.

Faces as Commodities

In China, over the past three years, courts have reviewed about 700 cases of illegal use of appearance via AI. In March 2026, a Beijing court issued an unprecedented ruling: using someone else's face in deepfakes without permission is illegal, even if the image is modified. This reflects the boom in AI content: more than 95% of the 128,000 AI mini-dramas released in the first three months of 2026 were created using artificial intelligence. Platforms like ActID and New Claw are structured as marketplaces: companies pay people from $15 to $700 for licensing images, and users themselves upload photos taken in special studios. Categories range from "girl next door" to "brutal" and "supermodel."

The stated goal is to protect people from what happened with ByteDance's Hongguo platform, when two influencers accused the company of stealing faces, and a series with 40 million views had to be removed. Since the beginning of 2026, ByteDance has destroyed over 85,000 videos with unauthorized use of other people's faces and voices. However, control remains fragile. British citizens who sold rights to their avatars to the startup Synthesia discovered that their digital clones were used in propaganda videos supporting dictators in Venezuela and Burkina Faso. The actors' union Equity is now pushing for a legislative ban on transferring uncontrolled rights to digital likenesses.

Not Just Books and Faces

Robots need examples of natural movements. Modern physics simulations are imperfect, and synthetic data for robots faces the same collapse problem. In 2025, over $6 billion was invested in humanoid development, and companies like Scale AI and Encord are hiring operators to record everyday movements. Tesla pays up to $48 per hour to people who, in motion capture suits, repeat routine actions to train Optimus.

Setting aside the difference in medium—texts, faces, voices, movements—the pattern is the same everywhere. Previously, data was taken without asking: scanning archives, scraping photos, training on voices from open access. Now people are paid for it. But the structure of the deal doesn't change: the asset irreversibly leaves the owner's control, just now with a receipt.

My comment: We are witnessing a fundamental shift: the AI industry is moving from parasitic data consumption to legal, but no less predatory, purchasing. Until the market develops mechanisms for long-term control over the use of licensed assets, owners—whether authors or models—remain hostages to the same deal: a one-time payment against indefinite loss of control.