The scale of generative models' penetration into web content turned out to be far more serious than commonly believed. My analysis of fresh Pew Research Center data shows that roughly 10% of all web pages on the internet bear traces of text created or deeply edited by artificial intelligence. However, the key figure is hidden deeper—among materials published after ChatGPT's debut, the share of synthetic content exceeds one-third.
AI is capturing new pages
Researchers examined a random sample of 10,000 English-language pages collected in July 2026 via Common Crawl. To identify authorship, a machine learning model trained to detect characteristic language patterns of generative neural networks was used. At first glance, 10% is a modest figure, but it includes a vast layer of historical content created before the era of modern LLMs.
When I filter out old artifacts and focus only on pages that appeared after November 2022, the picture changes radically. More than a third of such pages show clear signs of AI generation. This confirms that the growth of synthetic content directly correlates with the expansion of ChatGPT, Claude, Gemini, and their competitors.
The commercial sector is AI's main testing ground
The distribution of AI content across domain zones is extremely uneven and reflects economic incentives. In 2026, every tenth page in the .com zone bears traces of neural networks—twice as high as in .org (4.6%). On educational (.edu) and government (.gov) resources, the share drops to 1%, ten times lower than in the commercial sector. The gap is explained simply: business sites mass-publish product descriptions, SEO articles, and news that are easy to automate. At the dawn of the ChatGPT era, these differences were smoothed out, but now the market has clearly segmented.
Methodological nuances
It is important to emphasize: Pew did not verify the creation history of each page or survey authors. The analysis is based on statistical language markers that correlate with AI generation. Therefore, the results should be interpreted as an estimate of probable authorship, not an exact count. Moreover, this refers to content that is "written or substantially edited"—minor AI edits do not fall into this category.
Notably, Anthropic already introduced global labeling of Claude-generated content in August, which only confirms the systemic nature of the problem.
My conclusion: these figures are not just statistics but a signal of a tectonic shift in information production. The market is rapidly moving toward hybrid content, where the line between human and machine authorship is blurring. For investors and analysts, this means the need to reassess the quality of web assets and the risks of disinformation in the coming years.