Crypto news

24.08.2026
05:25

AI is taking over the web: more than a third of new pages are created by neural networks

AI fake news фейки

The scale of generative models' penetration into the web space has proven far more serious than commonly assumed. My analysis of recent data shows that roughly 10% of all existing web pages on the internet bear clear traces of synthetic text. However, when looking only at content published after ChatGPT's debut in November 2022, the picture becomes truly alarming—the share of such materials exceeds one-third.

Methodology: How to Tell Human from Machine

The study is based on a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive. For identification, a machine learning model was used to detect characteristic linguistic patterns in texts typical of generative algorithms. It's important to understand: this is not an exact count, but a probabilistic assessment of authorship.

At first glance, 10% might seem like a modest figure. But the sample included a vast layer of outdated content created long before the era of modern LLMs. When I cut off the "prehistoric" pages and focus on post-2022 materials, the share of synthetic content jumps to 34-36%. This is direct evidence of how quickly ChatGPT, Claude, Gemini, and their competitors are transforming the information landscape.

The Commercial Sector—The Main Area of Impact

The distribution of AI content across domain zones is highly uneven, and this says a lot. In the .com zone, signs of generation are detected on roughly one in ten pages—twice as often as in .org (4.6%). Meanwhile, in academic (.edu) and government (.gov) segments, the share of synthetic content barely reaches 1%.

This disparity is explainable. Commercial platforms have historically been geared toward mass text production: product descriptions, SEO articles, news feeds. These are precisely the formats easiest to automate. Editorial teams and corporate blogs actively use AI to cut content production costs, sacrificing quality and uniqueness for volume.

Caveats and Real Risks

It's worth emphasizing: the study does not claim that every such text was written from scratch by a neural network. This refers to content "written or substantially edited" by AI. Minor edits, such as grammar corrections, do not fall into this category. Nevertheless, the trend is obvious: the internet is rapidly filling with synthetic content, and this process is accelerating.

Against this backdrop, Anthropic's step toward global labeling of Claude-generated content looks timely but insufficient. Transparency of text origin is becoming critically important for maintaining trust in information. Without clear standards and verification mechanisms, we risk drowning in a stream of plausible but soulless generation that blurs the line between real experience and statistical probability.

My expert opinion: the market urgently needs content authentication tools at the protocol level, not voluntary initiatives from individual companies. Otherwise, in a couple of years, we won't be able to distinguish analysis from model hallucinations, and this will hit the entire digital media ecosystem.