Pew Study Finds Over a Third of New Web Content Is AI-Generated
Analysis of 500,000 webpages shows synthetic text surging on commercial domains.
An examination of roughly 500,000 English-language webpages published over the past five years reveals that artificial intelligence has taken over a massive footprint of online text following OpenAI’s launch of ChatGPT, according to findings published Thursday by Pew Research Center. The research organization noted that its data aligns with previous industry investigations tracking how recent web material is increasingly authored or “substantially edited” by machine-learning models.
These findings follow a recent report from web infrastructure provider Cloudflare showing automated bot traffic surpassing human navigation across the internet ahead of internal projections.
Rather than tracking web visitors, Pew analyzed the actual text hosted on the web. The combined datasets depict an ecosystem where automated crawlers frequently consume material produced by synthetic systems.
Pew compiled its dataset by extracting nearly 500,000 English-language webpages from the Common Crawl archive, spanning roughly five years beginning prior to ChatGPT’s debut in November 2022. Investigators then applied detection software from Open Pangram to evaluate the probability of AI authorship or heavy editing across the sample.
Within an unbiased sample of 10,000 webpages archived in July 2026, Pew identified “significant signs of AI authorship” in approximately 10% of the entries.
Researchers acknowledged that a broad cross-section inherently contains legacy webpages created prior to the availability of consumer AI generators, which could not have utilized automated writing platforms.
% of webpages showing significant AI editing or authorship
Image Credits:Pew Research
To isolate recent trends, Pew adjusted its methodology to exclude pre-2022 content and examine exclusively pages created after ChatGPT’s public rollout.
Once pre-existing pages were removed from the dataset, indicators of AI composition jumped to 35%, accounting for more than one in three newly published webpages.
ScreenshotImage Credits:Pew Research
Categorizing the data by top-level domain revealed that commercial .com addresses exhibited AI content traits at roughly ten times the frequency of .edu and .gov domains, both of which registered near 1%. Meanwhile, non-profit .org domains demonstrated a 4.6% prevalence of AI-generated text.
The sharp variance across domain extensions highlights the economic drivers powering automated text creation. While academic and governmental institutions maintain strict publishing controls and institutional standards, commercial domains frequently leverage automated content generation at scale to capture search traffic and monetize digital advertising.
Detection frameworks retain technical limitations, as Pangram and similar auditing software can produce false positives on human-authored text. However, statistical researchers consider large-scale sampling of this volume reliable for identifying macro-level trends.
Pew’s analysis also documented a measurable rise in specific stylistic markers frequently associated with large language models, including increased usage of em dashes, Oxford commas, and antithetical framing phrases such as “it’s not X, it’s Y.”
Reader purchases made through affiliate links on this site may yield a small commission, which operates independently from our newsroom coverage.








