The scale of generative models' penetration into the web space turned out to be far more significant than commonly believed. My data analysis, based on a sample of 10,000 English-language pages collected in July 2026 via Common Crawl, shows that about 10% of the entire web bears clear traces of synthetic generation. However, this is an average across the board—it blurs the real picture due to the vast array of outdated content created before the era of modern LLMs.
A critical shift after ChatGPT
The key insight lies in the dynamics. If we filter out pages published after November 2022—the moment ChatGPT launched—the numbers become truly alarming. More than a third of all new materials on the web show signs of machine authorship or significant editing by neural networks. This is not just statistics; it is a marker of a fundamental transformation of the information landscape, where algorithms have become the main factories of texts.
Geography of synthetic content: commerce vs. academia
The distribution of AI content across domain zones reveals a clear pattern. In the .com zone, the share of such pages reached 10%—twice as high as in .org (4.6%). Even more telling is the gap with the academic and government sectors: in .edu and .gov, this figure barely reaches 1%. Notably, at the dawn of ChatGPT's spread, these differences were almost imperceptible, indicating the rapid adaptation of commercial platforms to automating routine content: product descriptions, SEO texts, and news feeds.
Methodological caution
It is important to understand the limits of this assessment. The analysis relies on linguistic markers that statistically correlate with generative models, rather than on verifying the creation history of each page. This is a probabilistic estimate, not an exact count. Additionally, the category includes both fully generated and significantly AI-edited materials, leaving room for debate about the nature of "co-authorship."
My comment: These figures are just the tip of the iceberg. The real share of synthetic content is likely higher, as models become increasingly sophisticated and their outputs are harder to distinguish from human text. The market urgently needs verification and labeling standards; otherwise, we risk drowning in a stream of plausible but unverifiable information, where trust in the web as a source of knowledge will be completely undermined.