AI is taking over the web: more than a third of new pages are created by neural networks

A large-scale study covering a sample of 10,000 random English-language web pages collected in July 2026 sheds light on the rapid expansion of generative models. The analysis, conducted using a specialized machine learning model, revealed that approximately 10% of all indexed pages bear clear traces of synthetic authorship.
However, this average figure is deceptive. The key insight lies in the dynamics: if we consider only content published after the debut of ChatGPT in November 2022, the picture changes radically. Among "fresh" materials, the share of pages created or significantly edited by AI exceeds one-third. This is direct evidence that generative models have become not just a tool, but the primary "author" of a significant portion of new web content.
The commercial sector is the main driver of synthetic content
The distribution of AI generation across domain zones is extremely uneven and reflects economic incentives. The .com zone leads, where signs of AI are detected on approximately one in ten pages—twice as many as in .org (4.6%). Educational (.edu) and government (.gov) resources remain the "cleanest," with a rate of about 1%.
This polarization is explainable: commercial platforms actively automate the production of SEO texts, product descriptions, and news feeds, where speed and volume are critically important. At the dawn of the ChatGPT era, the gap between zones was less noticeable, but now the trend is obvious.
Methodology and caveats
It is important to emphasize: the study is not an exact count, but rather an estimate of probable authorship. The methodology is based on the analysis of linguistic markers statistically correlated with generative models, not on verifying the creation history of each page. Furthermore, the "AI content" category includes only texts written or substantially reworked by a neural network; minor edits are not taken into account.
This study is another wake-up call for the market. We are witnessing not just growth, but a structural transformation of the web, where synthetic content is becoming the "new normal" in the commercial segment. This poses new challenges for search algorithms and marketers: trust in information and the fight for organic traffic will increasingly depend on the ability to distinguish human text from machine-generated text. Labeling initiatives, such as the one implemented by Anthropic for Claude, are only the first step in this arms race.