AI is taking over the web: more than a third of new pages are created by neural networks

The scale of generative models' penetration into the web space has proven far more serious than commonly assumed. A fresh analysis, based on a sample of 10,000 English-language pages collected in July 2026 via Common Crawl, demonstrates that about 10% of all web content bears clear traces of synthetic generation. However, this is merely the tip of the iceberg.
Key shift after the launch of ChatGPT
The figure of 10% may seem moderate, but it is deceptive. After all, the overall sample includes a vast layer of archival materials created long before the era of modern LLMs. If we isolate only those pages published after November 2022 — the moment of ChatGPT's debut — the picture changes radically. Here, the share of content with signs of AI authorship exceeds one-third (33%+). This is direct evidence that generative algorithms have become the primary tool for text production for a significant portion of web publishers.
Where AI feels at home
The distribution of synthetic content across domain zones is highly uneven and follows purely economic logic. In the .com zone, signs of AI generation are detected on roughly one in ten pages (about 10%), which is twice the rate of .org (4.6%). Educational and government resources (.edu and .gov) remain the "cleanest" — there, the share barely reaches 1%. This is explainable: the commercial sector actively automates routine text streams — product descriptions, SEO articles, news briefs — where speed matters more than a unique authorial voice.
Methodology and nuances
It is important to emphasize: this is not about precise detection, but a probabilistic assessment. The analysis relied on linguistic markers statistically correlated with generative models. Researchers did not check the edit history of each page or interview authors. Therefore, it is more accurate to speak of a "significant probability" of AI involvement rather than a one hundred percent fact. Additionally, the category includes only content written or radically reworked by a neural network; minor proofreading with the help of assistants is not taken into account.
My view: These data are merely a starting point. The market is moving toward total content automation, and the current 33% among new pages is a minimum that will only grow. The question is not how to distinguish AI text from human text, but how search algorithms and readers will adapt to a new reality where synthetic content becomes the norm rather than the exception. Upcoming changes in ranking and labeling are just the first steps in this direction.