The scale of generative artificial intelligence's penetration into the web space has proven far more serious than commonly believed. My analysis of recent data shows that approximately 10% of all existing web pages bear clear traces of machine generation or significant editing. However, the most alarming indicator is that among content published after ChatGPT's debut, the share of synthetic texts exceeds one-third.

Methodology and Key Figures

During the study, a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive was analyzed. A machine learning model identifying characteristic linguistic patterns inherent to generative models was used for detection. The fact that the sample included a significant volume of "old" content created before the era of modern neural networks somewhat smooths the overall picture. If outdated materials are filtered out and the focus is placed exclusively on pages published after November 2022, the picture changes radically: signs of AI authorship are found in more than 33% of such documents.

The Commercial Sector — the Main Affected Area

The distribution of synthetic content across domain zones is extremely uneven. In the .com zone, signs of AI are detected in roughly one in ten pages, which is twice the rate of .org (4.6%). Educational (.edu) and government (.gov) resources show minimal presence—around 1%. This disparity is explainable: commercial platforms actively use automation to generate product descriptions, SEO texts, and news briefs, where speed matters more than uniqueness.

Important Caveats

It should be emphasized that the study does not claim every detected text was written by AI from scratch. The methodology relies on statistical analysis of linguistic structures rather than verification of a document's creation history. This is an assessment of probable authorship, not an exact count. Furthermore, the category of "significantly edited" content excludes cases of minor text revision by a neural network.

My comment: these figures are merely the tip of the iceberg. It is already clear that we are moving toward a web where synthetic content will become the norm rather than the exception. The question is not how to stop this process, but how to build systems of trust and information verification in the new reality. Content labeling, similar to what Anthropic implemented for Claude, is just the first step on this path.