AI is taking over the internet: more than a third of new web pages are created by neural networks.

The scale of generative models' penetration into web content has proven far more serious than commonly assumed. My analysis of recent data shows that approximately 10% of all existing web pages on the internet bear clear traces of machine generation or deep editing. However, when looking at content created after ChatGPT's debut in November 2022, the picture becomes radically different—more than a third of such pages show signs of synthetic origin.
Methodology and Key Figures
During the study, a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive was analyzed. A machine learning model tuned to detect characteristic linguistic patterns typical of generative algorithms was used for identification. It is important to understand: the sample included a significant layer of the "old" web, created before the era of modern LLMs, which explains the relatively modest 10% in the overall mass.
But if outdated pages are filtered out and the focus is placed exclusively on publications that appeared after ChatGPT's launch, the share of synthetic content soars to 33% and above. This is direct evidence that the spread of Claude, Gemini, and other models is not just accelerating content production but fundamentally changing its structure.
The Commercial Sector—The Main Driver
The distribution of AI content across domain zones is extremely uneven. In the .com zone, signs of generation were detected on roughly one in ten pages—twice as much as in .org (4.6%). On academic and government resources (.edu and .gov), the share drops to 1%, which is ten times lower. This disparity is simply explained: commercial platforms are mass-automating routine texts—product descriptions, SEO articles, reference materials, and news briefs, where publication speed is critical and the uniqueness of the authorial voice is not a priority.
Caveats and Reality
It should be emphasized: the study is not an exact count. This is a probabilistic estimate based on statistical linguistic markers, not verification of edit histories. Minor proofreading by a neural network does not fall into the category of "written by AI," but deep rewriting does. This is an important nuance that, however, does not negate the main conclusion: we are witnessing a tectonic shift in web content production, and its consequences for SEO, media, and audience trust have yet to be fully understood.
My expert opinion: these figures are merely the tip of the iceberg. While detection algorithms rely on linguistic patterns, generative models are quickly learning to mask them. It is already clear that without the implementation of cryptographic labeling of content origin (as Anthropic does for Claude), we risk drowning in a stream of high-quality synthetic content indistinguishable from human work, which would devalue the labor of living authors and undermine the foundations of information hygiene.