
- Pew Research pulled nearly 500,000 English-language pages from the Common Crawl archive covering roughly five years, starting two years before ChatGPT shipped.
- Once pages published before November 2022 are filtered out, 35 percent of the remaining pages show significant signs of AI authorship or heavy AI editing.
- Domain type predicts almost everything: .com sits at roughly 10x the rate of .edu and .gov, which both land near 1 percent, while .org comes in at 4.6 percent.
- On the same day the study landed, Google shipped an embeddable Preferred Sources button, with over 345,000 unique sources already selected by readers.
Ten percent. That is the share of a random 10,000-page sample of the web that Pew Research flagged as showing significant signs of AI authorship in July 2026. Then Pew removed every page published before ChatGPT existed, because those pages could not possibly have been machine-written, and the number moved to 35 percent. Two numbers, one dataset, and a fairly uncomfortable conclusion about what the open web is turning into.
A Third of the New Web Has a Machine in the Byline
What Pew Actually Counted
The study, released on Thursday, August 20, 2026, is built on two pieces of public infrastructure. The first is Common Crawl, the nonprofit web archive that most large language models were trained on in the first place. Pew drew close to half a million English-language pages from it, spanning roughly five years and deliberately starting a couple of years before ChatGPT’s November 2022 release so there would be a clean pre-AI baseline.
The second piece is Open Pangram, a detection system built to flag text that was either written by a model or substantially edited by one. That second category matters more than the headline suggests. Pew is not only counting pages a bot generated from nothing. It is also counting pages a human drafted and then handed to a model for a rewrite, which is much closer to how most content teams actually operate in 2026.
Why 10 Percent and 35 Percent Are Both True
A random sample of the web is mostly old web. Pages from 2014 sit in the index next to pages from last Tuesday, and no amount of prompting could have made a 2014 page synthetic. So the 10 percent figure is a measure of the whole archive, diluted by a decade of human writing that is still sitting there being indexed.
The 35 percent figure answers the question a publisher actually cares about: of the pages entering the web right now, how many have a model in the production chain. That is the number to plan around. It also implies the ratio is still climbing, because the denominator keeps refreshing while the pre-2022 human corpus stays fixed in size.
Trend Insight — The dilution effect cuts both ways. Every quarter that passes, the pre-ChatGPT human archive becomes a smaller share of what crawlers see. If you are training a model on fresh crawl data in 2027, roughly a third of your new corpus was authored by an earlier generation of models. That is the recursive training problem moving from a thought experiment to a measured input.
The Real Split Is Not Human Versus Machine. It Is .com Versus .edu
Where the Synthetic Text Concentrates
When Pew broke the results down by top-level domain, the pattern was stark. Commercial domains showed signs of AI authorship at roughly ten times the rate of .edu and .gov domains, both of which came in around 1 percent. Nonprofit .org domains landed at 4.6 percent, sitting between the two poles.
This is not a story about AI flooding the internet uniformly. It is a story about AI flooding the parts of the internet where publishing volume converts to revenue. Universities and government agencies publish because they are obligated to. Commercial sites publish because each additional indexed page is a lottery ticket in the search results, and generation cost just collapsed to near zero. The incentive structure explains the entire gap.
The Detection Caveat Nobody Should Skip
Pew is explicit that the analysis is imperfect. Pangram, like every AI-detection tool shipped so far, can misclassify human writing as machine writing. Anyone who has watched a detector flag a careful non-native English speaker’s essay knows the failure mode. At the scale of half a million pages, Pew argues the data is at least directionally correct, which is a reasonable claim and also a real limitation worth stating out loud before anyone quotes 35 percent as gospel.
There is a secondary finding that deserves a warning label of its own. Pew observed that supposed tells of AI writing have risen over time: em dashes, Oxford commas, and the “it’s not X, it’s Y” construction. Those are stylistic correlations, not evidence. Plenty of human editors have used em dashes for two centuries. Treating punctuation as a fingerprint is how false accusations start.
Trend Insight — The .edu and .gov floor near 1 percent is the most useful number in the report. It establishes what a low-AI baseline looks like under the same detector, which makes the .com figure much harder to dismiss as detector noise. If Pangram were simply over-flagging, government pages would be inflated too. They are not.
Bots Are Now Reading Pages Written by Bots
Two Curves That Crossed in the Same Season
The Pew report arrives shortly after Cloudflare reported that bot web traffic has overtaken human web traffic, a milestone the company says arrived sooner than its own projections. Pew measures what is being published. Cloudflare measures who is doing the reading. Put the two together and the picture is a web where the majority of requests come from automated clients, fetching a growing share of pages that automated writers produced.
For anyone running a content business, this reframes the analytics problem entirely. Traffic that used to be a proxy for attention is now partly a proxy for crawl frequency. A page that gets fetched a thousand times by retrieval agents and read by four humans is a very different asset than the same number would have described in 2021, and most dashboards still report both as a single line.
Trend Insight — Content strategy is quietly splitting into two products with different buyers. One is written for retrieval systems that need structured, citable, verifiable facts. The other is written for humans who will only show up if they already trust the byline. Optimizing a single page for both is getting harder, and the teams that separate them first will have cleaner metrics than the ones that do not.
Google’s Answer Is to Let Readers Pick Sides
The Preferred Sources Button
Hours after the Pew data went out, Google announced an interactive Preferred Sources button that publishers can embed directly on their own websites. A reader clicks it once, and that publisher gets surfaced more often for them across Search, Discover, and Google News. It extends a feature Google rolled out in May 2026 to AI Mode and AI Overviews, which had itself grown out of an option previously limited to Top Stories.
Two figures make the launch worth paying attention to. Google says people across the web have already selected over 345,000 unique sources through this mechanism. And in its own earlier studies, Google found that readers are twice as likely to click through to a preferred source when one is available. Alongside the button, Google is adding natural-language Discover feed tuning, where a reader taps the three-dot menu and simply tells the app what to show more or less of, plus customizable audio briefings in the Google News app on Android.
What Publishers Should Actually Do This Week
Read the timing honestly. Google’s AI Overviews are a major reason publisher traffic fell, and Preferred Sources is a partial remedy offered by the same party. It does not restore the old click economy. What it does is convert one-time readers into a persistent signal, which is a genuinely valuable asset if you can get people to press the button.
The practical sequence is short. Embed the button where an engaged reader finishes something, not in a header nobody looks at. Ask for it in newsletters, where the audience already opted in once. And accept the strategic consequence: when a third of new pages are synthetic and a machine is doing most of the reading, the only durable moat is being the specific source a human deliberately chose. Volume was the old game. Being picked is the new one.
Trend Insight — A 2x click-through lift on a shrinking base of organic clicks is still a smaller number than 2023 traffic. Preferred Sources should be treated as damage mitigation, not recovery. The publishers who benefit most will be the ones with a direct relationship to ask through, which means the button rewards exactly the audiences that were never dependent on search in the first place.
Related
- Anthropic Hid Something in Every Word Claude Writes
- OpenAI Found a Way to Watch You Without Looking
- GitHub Went Dark. Cursor Was Already Waiting.
- 60 AI Agents Just Attacked Math’s 150-Year Mystery
- Meta Just Made the Cloud Optional for AI Agents
Sources
- TechCrunch, “A third of web pages published since ChatGPT’s launch show signs of AI authorship, study finds” (Aug 20, 2026)
- Pew Research Center, “How much of the internet is written with AI?” (Aug 20, 2026)
- TechCrunch, “Google gives publishers a new way to fight AI-driven traffic losses” (Aug 20, 2026)
- Google, “Preferred Sources and original high-quality content in Search”
- Common Crawl, open web archive used as the study corpus
AI Biz Insider · AI Trends EN · aibizinsider.com

댓글 남기기