The digital space is changing rapidly, and judging by the latest data, humanity has ceded a significant portion of the "creative kitchen" on the internet to machines. My analysis of fresh statistical data shows that about 10% of all web pages on the global network bear a clear imprint of generative models. However, the most alarming indicator lies in the dynamics—if we consider only content published after November 2022, that is, after the debut of ChatGPT, the share of synthetic texts exceeds one third.

During the study, a representative sample of 10,000 English-language pages collected in July 2026 was analyzed. To identify the "digital fingerprint," a machine learning model tuned to search for specific linguistic patterns characteristic of AI generation was used. The fact that archives created long before the era of modern neural networks were included in the overall statistics only underscores the scale of the current expansion.

The commercial sector is the main area of impact

The distribution of synthetic content across domain zones looks extremely telling. While in the .com zone, signs of AI are detected in roughly one in ten pages, in the non-commercial .org segment this figure drops to 4.6%. Educational and government resources (.edu and .gov) demonstrate an even more striking contrast—there, the share of such content barely reaches 1%.

This disparity is explainable: commercial platforms most often depend on large volumes of routine textual content—product descriptions, SEO articles, and news feeds—which are extremely easy to automate. At the beginning of ChatGPT's spread, the gap between zones was not as critical, which points to the accelerating commercialization of AI tools.

Methodological caveat: accuracy versus probability

It is important to understand: this study is not a "lie detector" for each individual text. The methodology is based on statistical analysis of linguistic structures, not on checking a document's edit history. Therefore, it is more correct to speak of "probable authorship" rather than a proven fact. Moreover, minor editing of a text by a neural network most likely remained outside the sample—we are talking specifically about fully generated or radically reworked material.

My comment: We are witnessing not just a trend, but a tectonic shift in the content economy. The fact that a third of the new internet is created without direct human involvement calls into question not only copyright and SEO optimization, but also the very concept of "information noise." In the coming years, we will have to develop new data validation mechanisms, otherwise trust in the web as a source of knowledge will be completely undermined.