Crypto news

23.08.2026
01:15

AI already controls a third of the new web: shocking research figures

AI fake news фейки

The scale of integrating generative neural networks into web content has turned out to be far deeper than commonly believed. My analysis of recent data shows that approximately 10% of all existing web pages bear clear traces of synthetic text. However, this is only the tip of the iceberg—if we filter out "prehistoric" content created before the era of modern LLMs, the picture becomes truly alarming.

The ChatGPT era turned the content market upside down

The key conclusion I want to emphasize: among materials published after November 2022 (the launch of ChatGPT), the share of AI authorship exceeds 33%. This is no longer experimentation by enthusiasts, but a systemic phenomenon. For comparison, in the overall data array, including the legacy of past years, this indicator is diluted to 10%—the old web architecture simply masks the real speed of technology adoption.

The distribution across domain zones confirms my long-standing hypothesis about the commercialization of synthetic content. In the .com zone, every tenth page (10%) shows signs of AI generation, while in .org this figure is twice as low—4.6%. Educational (.edu) and government (.gov) resources show only 1%—manual labor still dominates there, which is logical given the requirements for fact verification.

Methodology and reality

It is important to understand: the study is based on analyzing language patterns characteristic of generative models, not on direct access to edit history. This is an estimated metric, not an exact count. Nevertheless, the statistical significance of a sample of 10,000 pages collected via Common Crawl allows us to speak of data representativeness.

Of particular interest is the fact that commercial platforms are most actively exploiting AI to create SEO texts, product descriptions, and reference materials. This creates an effect of "information noise" that, in the long run, could devalue trust in web content as a whole.

My verdict: we are witnessing a tectonic shift in information production. Already now, a third of the new web is synthetic, and this process is irreversible. The question is not whether it will stop, but how quickly search algorithms and regulators will adapt to the new reality. Content labeling, like the one Anthropic introduced for Claude, is only the first step toward hygiene in the information space.