The scale of generative neural networks' penetration into the web space has proven significantly deeper than commonly assumed. My analysis of fresh data shows: roughly every tenth web page on the internet bears traces of synthetic authorship. However, when looking at content published after the launch of ChatGPT, the situation appears far more radical—more than a third of such pages exhibit clear signs of generation or substantial editing by artificial intelligence.
Methodology and Key Figures
During the study, a random sample of 10,000 English-language pages collected in July 2026 was analyzed. For identification, a machine learning model trained to detect linguistic patterns characteristic of modern LLMs was used. At first glance, 10% of the total volume is a moderate figure. But one should not forget: the sample included a huge amount of "old" web content created long before the era of generative AI.
If we cut off the legacy of the past and focus exclusively on materials published after November 2022, the picture changes dramatically. Synthetic traces are found in more than a third of new pages. This is direct evidence that ChatGPT, Claude, Gemini, and their counterparts have become the main "authors" of modern content.
Who Leads in Generation
The distribution of synthetic content across domain zones is extremely uneven, and this reflects economic logic. In the .com zone, signs of AI are found on every tenth page (about 10%), which is twice the rate of .org (4.6%). Meanwhile, educational (.edu) and government (.gov) resources show only about 1%—ten times less.
This disparity is explained by the nature of the content. The commercial sector floods the web with mass-produced texts: product descriptions, SEO articles, and news briefs. These are precisely the formats easiest to automate. At the start of the boom, the gap between zones was less noticeable, but now the market has fully adapted to AI capabilities.
Important Caveats
It should be emphasized: the study does not claim that every such text was written by a bot. The methodology is based on statistical analysis of linguistic markers, not on checking file creation history. This is more an estimate of probable authorship rather than an exact count. Furthermore, the category of "written or edited" by AI implies significant processing, not merely the use of autocorrect or minor tweaks.
My comment: These data are a wake-up call for the content market. We stand on the brink of an inflation of information noise, where search engines and aggregators are drowning in synthetic content. The industry needs new standards for verification and labeling; otherwise, trust in the web as a source of knowledge will be completely undermined. Initiatives like content labeling from Anthropic are only a first step in the right direction, but the industry has a long road ahead.