The scale of generative models' penetration into the web space has proven far more serious than commonly assumed. My analysis of recent data shows that roughly one in ten web pages on the internet carries clear signs of machine generation or substantial editing. However, when looking at content published after ChatGPT's debut in November 2022, the picture becomes radically different—more than a third of such pages have synthetic origins.
Methodology and Scale of the Phenomenon
During the study, a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive was analyzed. A machine learning model tuned to detect characteristic linguistic patterns typical of generative neural networks was used for identification. The overall figure of 10% may seem modest, but it is diluted by the vast layer of outdated content created long before the era of modern LLMs.
When we cut off the "prehistoric" web and focus exclusively on publications that appeared after ChatGPT's launch, we see a rapid expansion of AI. A third of new pages is not a marginal niche but a systemic trend that will only intensify as Claude, Gemini, and other models are integrated into everyday workflows.
The Commercial Sector—The Main Driver
The distribution of synthetic content across domain zones highlights the economic underpinnings of the phenomenon. In the .com zone, signs of AI generation are found on one in ten pages—twice as high as in the non-commercial .org zone (4.6%). Even more telling is the gap with educational and government domains (.edu and .gov), where the share barely reaches 1%.
This disparity is logical: commercial platforms actively automate the creation of product descriptions, SEO texts, and news feeds. These are voluminous, template-based formats where speed and resource savings matter more than a unique authorial perspective. At the beginning of the ChatGPT boom, these differences were smoothed out, but now the market has clearly segmented.
Important Methodological Caveats
It should be emphasized that this is not about a precise count of "robot-written" content but rather a probabilistic assessment. Researchers did not check edit histories or survey authors. The analysis relied on statistical linguistic markers that may also be present in text written by a human but subjected to aggressive AI editing. Additionally, minor edits made with the help of neural networks typically do not fall into this category.
In this context, Anthropic's recent decision to globally label content created by Claude is telling. The industry is beginning to realize: without clear identification of synthetic content, we risk finally losing the line between fact and simulacrum.
My conclusion: we are witnessing not just a trend but a tectonic shift in the content economy. The web is rapidly transforming into an environment where human authorship becomes a premium product against a backdrop of cheap and faceless machine-generated mass. The question is not whether anyone will stop this, but how we adapt search and trust algorithms to the new reality.