Crypto news

08.08.2026
04:52

The AI industry is buying up books and faces: why neural networks need "organic" data and what comes next

img-bbaf321c89b70f44-1766402574726061

The ISBNdb service, which aggregates metadata for 113 million publications, has offered AI companies a specific service: purchasing books in batches of up to one million copies, followed by digitization and disposal. When this scheme became public knowledge, the company hurried to remove the page describing the service and stated that no deals had been made, calling it merely a "test of market demand."

In another market segment, the platforms ActID and New Claw are ready to pay people from $15 to $700 for the right to use their faces as the basis for characters in AI series and advertising.

Let's figure out why laboratories are forced to buy up "organic" data, what "model collapse" is, and what other human resources machines will need.

Ouroboros: When AI Devours Itself

The term "model collapse" describes the degradation of an AI system as it trains on data generated by other neural networks. With each cycle, the result becomes more averaged, predictable, and loses touch with the reality that a large language model is supposed to reflect.

This term was first formalized by an international group of researchers led by Ilya Shumailov. In May 2023, they published the resonant paper "The Curse of Recursion: Training on Generated Data Makes Models Forget," where they mathematically proved that with regular mixing of synthetic data, models stop correctly reproducing information about rare events and then begin accumulating errors.

The scale of the problem is confirmed by numbers. In April 2025, researchers at Ahrefs checked 900,000 newly indexed pages in Google search — the share of materials with signs of auto-generation exceeded 74%. According to Epoch AI estimates, all suitable training text on the internet could run out between 2026 and 2032.

This is precisely why books published before 2022 hold special value — they are one of the few sources guaranteed to be free of machine-generated junk and deliberate "traps." Some authors have already resorted to a countermeasure — "poisoning" data by hiding symbols in the text that the model interprets as a hidden command. Research by a joint team from the UK AI Safety Institute, Anthropic, and Oxford proved that LLMs can be covertly trained on malicious patterns using just 250 "poisoned" documents.

Fahrenheit 451 by Anthropic

The fate of books in this race is not obvious to all sellers. In August 2024, three well-known American authors filed a class-action lawsuit against Anthropic, accusing the management of copyright infringement when training Claude models. According to case materials, the company invested tens of millions of dollars in the "unofficial" Panama project: anonymously buying up print runs of physical books, cutting off bindings, and scanning them on an industrial scale, followed by disposal.

In June 2025, Judge William Alsup ruled that digitizing legally purchased printed books and using digital copies for training language models falls under the fair use doctrine. Now it is enough to buy a paper book and keep the receipt — the origin of the data is documented.

According to 404 Media, this practice has long extended beyond individual startups. ISBNdb, which has operated for over two decades as the largest international metadata database of printed books, has turned the buying up and destruction of books into a turnkey service. On July 21, 2026, journalist Emanuel Maiberg discovered a promotional section where the company offered AI laboratories wholesale brokerage services:

  • organizing the buyout of printed books published before 2022 in batches from 1,000 to 1 million copies;
  • guarantees of book "purity" from AI slop and poisoning algorithms;
  • legal protection through non-disclosure agreements and the legality of acquisition on the secondary market.

The platform's management directly stated: "The perception problem is real. 'AI company destroyed two million books' is not the kind of headline that will win public sympathy."

Independent used-book dealers confirmed anomalous purchases starting in April 2026: sales grew from twenty copies to several hundred per week. Buyers were not interested in genre or value — the main criterion was the ISBN code. Even rare foreign editions were purchased, ignoring the cost. Unique copies are lost irretrievably — the only remaining copy may end its journey under a scanner.

After the investigation was published, ISBNdb removed the promotional page and issued a statement: "The facts: ISBNdb has never bought, scanned, or sold books. This page was merely a test of market demand."

The gap between the cost of raw material and the finished product is hardly proportional. A book destroyed immediately after scanning costs between $2 and $5 on the secondary market. A model trained on millions of such copies is valued in the billions of dollars.

Need More: The Market for Faces and Voices

Over the past three years, the Guangzhou Internet Court has reviewed about 700 cases of illegal use of others' appearances through AI. In March 2026, a Beijing court issued an unprecedented ruling: using someone else's appearance in deepfakes without permission is illegal, even if the image was modified.

This practice reflects the growth of the AI content market. In China, more than 95% of the 128,000 AI mini-dramas released in the first three months of 2026 were produced using artificial intelligence. The demand for real faces has given rise to a new type of transaction — appearance rental. The platforms ActID and New Claw are structured like a marketplace: companies pay from $15 to $700 for image licensing. Users upload photos taken in special studios, and producers select faces by gender, age, and categories such as "girl next door" or "rugged."

New Claw's head of operations, Long Liu, explained the logic simply: "Selling the rights to use photos gives people the opportunity to earn money while continuing their offline careers." ActID, founded in Shenzhen in March 2026, registered about 800 users, of whom 300 agreed to license their images. Prices range from $15 to $74 per episode, with a platform commission of 10%.

The stated goal of the platforms is to protect people from what happened with the AI drama platform Hongguo, owned by ByteDance. In April 2026, two influencers accused the company of stealing faces. A series with 40 million views had to be removed. Since the beginning of 2026, ByteDance has destroyed over 85,000 videos with unauthorized use of others' faces and voices.

Beijing lawyer Yile Deng considers the licensing market a positive step but warns: "Many agencies pay people mere tens of dollars for data collection. Licensing terms are often so vague that it's impossible to understand who ultimately uses this appearance and for what purpose."

In practice, control has proven fragile. People who sold a license for their face have found their digital clones used in dishonest advertising or political propaganda. Two high-profile cases involved Britons who sold avatar rights to the startup Synthesia: actor Dan Dewhirst discovered that his "clone" had become the host of a fake news channel linked to Nicolás Maduro's government, and model Mark Torres found his likeness used in propaganda videos supporting the dictator of Burkina Faso. Both cases became precedents for the British actors' union Equity, which is pushing for a legislative ban on transferring uncontrolled rights to digital likenesses.

Don't Forget the Receipt

Books and faces are not all that AI companies need. A similar situation is developing in the voice data market: in June 2026, members of the SAG-AFTRA union ratified an agreement expanding protection against the use of synthetic voice copies without consent.

Robots need examples of natural movements. Modern physics simulations do not accurately reproduce how a hand wraps around a mug, so synthetic data for robots faces a problem similar to model collapse. According to MIT, more than $6 billion was invested in humanoid development in 2025, and Tesla pays up to $48 per hour to people who, for hours in motion capture suits, repeat routine actions to train Optimus machines.

Setting aside the difference in medium — texts, faces, voices, movements — the scheme is the same everywhere. Previously, data was taken without permission: archives were scanned, photos were scraped. Now people are paid for it. Only the structure of the deal does not change: the asset still irretrievably leaves the owner's control, just now with a receipt.

Books and human appearance seem like different categories, but for the AI industry they are one type of resource — training data. Companies are willing to buy, but the market is structured in such a way that the deal does not imply protection. After transferring rights, a person cannot fully control who will use their voice, face, or created content and how.

My conclusion: we are witnessing a fundamental shift in the data economy. The AI industry, having reached the ceiling of "free" content, is now forced to pay for organic data, but does so without transparent rights protection mechanisms. While courts and unions catch up with reality, the market will remain a "Wild West," where the stakes are not just money, but also control over one's own identity.