The digital landscape is transforming rapidly: my analysis of fresh data shows that approximately 10% of all web pages on the internet bear clear traces of machine generation or deep editing. However, a far more alarming figure lies in the dynamics—among materials published after ChatGPT's debut, the share of synthetic content exceeds 33%.

The Scale of Expansion: From 10% to a Third

During the study, a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive was analyzed. A machine learning model tuned to detect linguistic patterns characteristic of generative neural networks was used for identification. At first glance, 10% is a moderate figure, but it includes a vast layer of "prehistoric" content created before the era of advanced LLMs.

If we filter out this legacy content and focus exclusively on publications that appeared after November 2022, the picture changes radically. More than one-third of such pages demonstrate a high probability of AI authorship. This is direct evidence that generative models have become the primary tool for text production across a significant portion of the web.

The Geography of Synthetic Content: Commerce Leads

The distribution of AI content across domain zones is highly uneven, pointing to a pragmatic use of the technology. In the .com zone, signs of generation are detected on approximately one in ten pages. By comparison, in the non-commercial .org zone, this figure is twice as low—4.6%. Educational (.edu) and governmental (.gov) resources remain the "cleanest," with the share of synthetic content barely reaching 1%.

This differentiation is explainable: commercial platforms actively automate the creation of SEO texts, product cards, and news feeds. At the early stage of ChatGPT's spread, the gap between zones was minimal, but now the market has clearly segmented by the degree of AI adoption.

Methodological Caution

It is important to understand the limitations of this assessment. The study does not check the edit history of each page or survey authors. The conclusions are based on a statistical analysis of linguistic markers that correlate with AI generation. This is a probabilistic estimate, not an exact count. Additionally, the category of "substantially edited" content excludes light proofreading, making the estimates even somewhat conservative.

My comment: The market has already passed the point of no return. We are witnessing not just a trend, but a fundamental shift in the content economy. However, the race for volume using AI creates a risk of information devaluation and a rise in "digital noise." In the coming years, the key asset will not be generation speed, but trust in the source and the quality of data verification. Players who can build hybrid editorial processes, where AI enhances rather than replaces analytics, will gain a decisive advantage.