The scale of generative AI's penetration into the web space turns out to be far more serious than commonly believed. My analysis of fresh data shows: roughly one in ten web pages on the internet bears clear traces of machine generation or significant editing. However, when looking at content created after the launch of ChatGPT in November 2022, the picture becomes radically different — more than a third of such pages show signs of AI authorship.

Methodology and scale of the phenomenon

A sample of 10,000 random English-language pages collected in July 2026 via Common Crawl was analyzed using a machine learning model. The algorithm identified linguistic patterns statistically characteristic of generative models. It is important to emphasize: this is not an exact count, but an estimate of probable authorship based on language markers.

The fact that among "fresh" publications the share of synthetic content exceeds 33% speaks to a tectonic shift in information production. ChatGPT, Claude, Gemini, and their competitors have become not just tools, but the primary "authors" of a significant portion of web content.

The commercial sector — the main driver

The distribution of AI content across domain zones is extremely uneven. In the .com zone, signs of generation were detected in about 10% of pages — twice as much as in .org (4.6%). Even more telling is the gap with the academic and governmental segments: in the .edu and .gov zones, the share of AI texts is only about 1%.

This disparity is explainable. Commercial platforms actively automate the production of product descriptions, SEO articles, and news briefs — formats that are easily scalable by neural networks. In contrast, sites with high requirements for reliability and expertise (education, the public sector) retain "human" authorship.

Important caveats and conclusions

A significant remark must be made: the study did not check the history of page creation nor survey the authors. This is about a probabilistic estimate, not a proven fact of generation. Furthermore, the category of "written or significantly edited" by AI excludes cases of light proofreading, making the estimate conservative.

Against this backdrop, Anthropic's recent step of introducing global labeling for content created by Claude looks logical. However, in my view, this is only the tip of the iceberg. We are moving toward a world where verifying the origin of information will become no less important than its content. The market needs new standards of authenticity, otherwise trust in the web as a source of knowledge will rapidly erode.