The scale of generative neural networks' penetration into web content has turned out to be significantly deeper than commonly believed. My analysis of fresh data shows: approximately 10% of all existing web pages bear clear traces of synthetic generation or substantial editing by algorithms. However, the most alarming indicator is that among materials published after ChatGPT's debut, the share of such content exceeds one-third.
Methodology and Key Figures
During the study, a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive was analyzed. A machine learning model tuned to detect specific linguistic patterns characteristic of modern LLMs was used for identification.
The fact that only every tenth page in the overall corpus shows signs of AI may be misleading. This is explained by the huge amount of "legacy" content created before the generative AI era that ended up in the sample. If we filter out the historical layer and focus exclusively on publications released after November 2022, the picture changes radically: more than 33% of new pages are in one way or another created or edited by neural networks.
Where AI Feels at Home
The distribution of synthetic content across domain zones is extremely uneven, and this speaks volumes. In the .com zone, signs of AI were detected on approximately every tenth page. For comparison, in .org this figure stands at 4.6%, while in academic (.edu) and government (.gov) domains it is around 1%.
Such a disparity is logical. The commercial sector has historically been tied to large volumes of content—product descriptions, SEO texts, news feeds. These are precisely the formats easiest to automate, which is what we observe in practice. At the start of the ChatGPT boom, the gap between zones was less pronounced, indicating an accelerating commercialization of synthetic content.
Important Caveats
It is worth emphasizing that the study is not a "lie detector" for each specific page. The methodology assesses probability rather than the exact fact of AI usage. We are talking about texts where the share of machine intervention is significant; minor edits or proofreading by a neural network most likely remained out of scope.
My commentary: These figures are merely the tip of the iceberg. By my estimates, the share of synthetic content in new material has already crossed the 50% mark, especially in news aggregators and niche blogs. We are moving toward a point of no return, where search algorithms and recommendation systems will be forced to completely rethink their ranking criteria, or they will simply drown in a flood of generated noise. The question is not whether the internet will become synthetic, but how we learn to separate the wheat from the chaff.