Machine-Generated Text Dominates Commercial Web, Pew Research Analysis Finds
Pew study reveals automated text is multiplying rapidly across .com domains as publishers adopt large language models
Artificial intelligence detectors applied to nearly half a million public webpages reveal a rapid spread of machine-generated text across the internet. A new study by the Pew Research Center shows that automated content is heavily concentrated within commercial top-level domains.
Researchers examined a representative sample of approximately 500,000 public webpages to measure the footprint of generative artificial intelligence online. The analysis utilized advanced classifier tools designed to identify linguistic patterns typical of large language models.
Non-profit, educational, and governmental sites registered vastly lower levels of synthetic text across evaluated samples, a contrast to commercial .com websites, which hold the highest concentration of AI-generated material compared to other domain extensions, the results indicate.
The surge in automated text corresponds with the widespread rollout of consumer-facing generative AI tools over the past two years. Millions of websites have increasingly deployed large language models to automate content creation, search engine optimization, and marketing copy.
This rapid proliferation raises significant concerns among digital researchers regarding the overall quality and reliability of web information. Automated content farms and synthetic text hubs have multiplied, creating vast networks of low-cost pages designed primarily to capture search engine traffic.
Search engines have responded by updating ranking algorithms to combat automated spam and unoriginal text. Major platform operators continue to adjust indexing guidelines to penalize domains that produce automated content at scale without human oversight.
The study also highlights technical challenges in tracking synthetic text as generative models become increasingly sophisticated. Detecting AI output requires continuous calibration of classification models, as newer language models produce text that closely mimics human writing styles.
Domain-level differences show that regulatory and organizational structures play a key role in limiting automated content spread. Educational institutions and government agencies maintain strict publishing protocols, which significantly reduces the presence of unvetted synthetic text on .edu and .gov domains.
The Pew findings align with broader industry metrics tracking the shift toward automated web publishing. Data from web crawling agencies shows that synthetic material now accounts for a growing percentage of newly indexed pages across global commercial networks.
Researchers noted that commercial publishers face increasing economic pressures to cut costs through automation. As artificial intelligence models become cheaper to deploy, the volume of automated content on commercial domains is projected to increase further in coming reporting cycles.








