A Third of the New Web Shows Signs of Being Written by AI, Pew Finds
Pew Research scanned roughly half a million web pages and found AI authorship signals in 35 percent of those published after ChatGPT launched. Commercial sites score ten times higher than universities and government.
Pew Research published a study on 20 August that puts a number on something most of us have only suspected while reading. Of the web pages published since ChatGPT’s launch in November 2022, roughly 35 percent show significant signs of having been written or heavily edited by AI. In a plain random sample of 10,000 pages collected in July 2026, the figure was around 10 percent, but that sample inevitably included old pages written before AI writing tools existed. Filter those out and the share jumps to more than a third.
The method matters here. Pew pulled nearly half a million English-language pages from Common Crawl, a free public archive of the web, covering roughly five years and starting a couple of years before ChatGPT existed. It then ran them through Open Pangram, an AI text detector. The breakdown by domain is the most striking part: .com addresses showed AI authorship at about ten times the rate of .edu and .gov addresses, which both landed around 1 percent. The .org domains sat at 4.6 percent. Pew also tracked the supposed tells that people love to argue about, and found rising use of em dashes, Oxford commas and the “it’s not X, it’s Y” construction.
Two caveats belong right next to that headline number. First, detection is not proof. Pangram and every tool like it will sometimes flag human writing as machine-made, especially formal or formulaic prose, so the figure is directionally useful rather than exact. Second, “written with AI” covers an enormous range, from a marketing page generated wholesale in one prompt to a human draft that someone ran through a model for grammar. Those are very different things and the study cannot cleanly separate them. What the domain breakdown does suggest is that the pressure is commercial: pages that exist to rank in search and sell something are where the volume is, while universities and public bodies have far less reason to churn out text.
What this means for you: if you search for something and the top results all feel oddly interchangeable, you are not imagining it. Treat generic .com content pages with more caution than you used to, and lean harder on sources with a name and a track record attached. If you write for the web yourself, the useful takeaway is not “never use AI” but that the tells are measurable and the detectors are getting better, so passing off unedited output as your own is a shrinking bet. And if you are curious about your own habits, the em dash finding is a good reminder that style quirks travel: the models learned them from us first.
Sources
Source: https://www.pewresearch.org/data-labs/2026/08/20/how-much-of-the-internet-is-written-with-ai/
Reinforcement Learning Pioneer Rich Sutton Calls Synthetic Data a Big Mistake
Sutton argues that a world of endless complexity cannot be learned from data a model generates for itself. It is a direct challenge to one of the industry's favourite shortcuts.