The scale of generative neural networks' penetration into the web space has proven far more significant than commonly assumed. My analysis of recent data shows: roughly every tenth web page on the internet bears clear traces of machine generation or deep editing. However, this is merely the tip of the iceberg—if we consider only content published after ChatGPT's debut in November 2022, the picture becomes truly alarming: already more than 33% of such pages show signs of synthetic origin.
Methodology and Scale of the Phenomenon
These conclusions are based on an analysis of a representative sample of 10,000 English-language pages collected in July 2026. For identification, a machine learning model was used, calibrated to detect specific linguistic patterns characteristic of modern LLMs. It is important to understand: the overall statistics also include the "old" web, created long before the era of generative AI, which explains the relatively modest 10% in absolute figures.
The dynamics are evident: the growth in the share of synthetic content directly correlates with the expansion of ChatGPT, Claude, Gemini, and their counterparts. The wider these tools are adopted, the more content they produce.
Where Does AI Feel at Ease?
The distribution of machine-generated text across domain zones is extremely uneven and telling. The commercial sector (.com) leads—about 10% of pages. This is twice as high as in the non-commercial .org zone (4.6%). Meanwhile, educational (.edu) and governmental (.gov) resources demonstrate almost complete "purity"—only about 1% synthetic content.
Such polarization is explainable: the commercial web is geared toward volume and SEO. Product descriptions, reference articles, news feeds—all of this is an ideal testing ground for automation. In contrast, academic and governmental spheres impose higher requirements for authorship and verification, which curbs the indiscriminate use of neural networks.
Important Caveats
It should be emphasized: this is not about an exact count, but a probabilistic assessment. The methodology does not involve checking edit histories or surveying authors. The analysis relies on statistical linguistic markers, which may be either a clear sign of generation or the result of stylization. Additionally, light proofreading by humans does not fall into the category of "written or substantially edited."
My comment: We stand on the threshold of a fundamental shift. The web is rapidly transforming into an environment where synthetic content becomes the norm, not the exception. This calls into question not only the reliability of information but also the very content economy: search algorithms and advertising networks are already failing to filter out the noise. Labeling initiatives, like the one introduced by Anthropic, are merely a first, tentative step. Without global standards for verification and identification, we risk drowning in a sea of plausible but meaningless information, where the human voice will be finally drowned out by machines.