AI is taking over the web: more than a third of new pages are created by neural networks — my analysis of Pew Research data

The scale of generative models' penetration into the web space turned out to be far more serious than commonly believed. My analysis of fresh data from a research center shows that about 10% of all indexed web pages in the English-language segment bear clear traces of synthetic text. However, the key figure is hidden deeper—among materials published after ChatGPT's debut in November 2022, the share of AI content exceeds 33%.
Methodology and Real Numbers
The study is based on a random sample of 10,000 English-language pages collected in July 2026 via the Common Crawl archive. For identification, a machine learning model trained to detect characteristic linguistic patterns of generative algorithms was used. It is important to emphasize: this is not a binary detector, but a probabilistic assessment based on statistical anomalies in the text.
At first glance, 10% may seem like a modest figure. But the sample includes a vast layer of legacy content created before the era of modern LLMs. When I filter out outdated pages and focus exclusively on publications from recent years, the picture changes radically: every third new page on the web shows signs of AI authorship or significant neural network editing.
The Commercial Sector—The Main Driver of Synthetic Content
The distribution of AI content across domain zones demonstrates a clear pattern. In the .com zone, signs of generation are found in about 10% of pages—twice as high as in .org (4.6%). Most indicative are educational (.edu) and government (.gov) resources, where the share of synthetic content does not exceed 1%.
This disparity is explainable by the economics of content. Commercial platforms actively automate the production of product descriptions, SEO texts, and news briefs, where speed matters more than uniqueness. At the early stage of ChatGPT's spread, the gap between zones was less pronounced, but now the trend is obvious: the higher the traffic monetization, the more aggressive the adoption of AI.
A Critical Caveat on the Data
Despite the impressive figures, I must note the limitations of the methodology. The study does not analyze edit history nor survey authors. It concerns linguistic markers that statistically correlate with AI generation but are not absolute proof. Additionally, the category of "significantly edited" content excludes minor edits, making the assessment conservative.
In this context, let me remind you: in August, Anthropic introduced global labeling for content created by Claude models. However, without a unified standard for all LLMs, the market remains fragmented.
My verdict: these data are an alarming signal for the web ecosystem. If a third of new content is generated by algorithms without transparent labeling, we are moving toward an information environment where trust in text will require cryptographic verification. For analysts and investors, this means a growing value of platforms with proven human authorship—and an inevitable tightening of regulatory requirements for AI publications.