TL;DR
European publishers get scraped by AI bots far more than North American ones, receive fewer readers back, and have their robots.txt instructions disregarded more often — according to TollBit, which studied bots from 40 vendors across 3,906 sites. Cloudflare and DataDome measure the same phenomenon and do not find the same gap. The scraping is real; the explanation is contested.
What TollBit found
The median European site is scraped four times as often as its North American counterpart. It gets back one human visit from an AI app for every 179 bot visits, roughly three times worse than the North American rate. Across the first half of 2026, AI apps sent European sites 0.05% of their external referrals, against 0.16% the other side of the Atlantic.
The trend is deteriorating. Scraping of European sites in June ran almost 20% above January, with local news and sports growing fastest, and the ratio of scrapes to referrals widened from 150:1 to 227:1 between the first and second quarters. Instructions in robots.txt telling crawlers to stay away are ignored on European sites close to three times as often.
Two explanations, and a caveat
TollBit’s cofounder Olivia Joslin points to language: Europe has many, and models training across them pull from more sources. Grzegorz Piechota of INMA reaches the same conclusion from usage data — American users made up 21.6% of Claude usage and English-speaking countries under a third of its user base, on Anthropic figures from last September, implying most demand is for non-English material. His Common Crawl analysis supports it: European domains were 4.75% of captured pages in 2009 and are 29.98% now.
Piechota also supplies the caveat, noting the finding may partly reflect which publishers happen to be TollBit customers.
Where the datasets diverge
Cloudflare’s Lai Yi Ohlsen says absolute bot requests are consistently higher for North American sites, while European sites see bots make up a larger share of their traffic. DataDome’s Jérôme Segura finds no consistent regional gap at all, describing enormous variance between individual publishers — some seeing one AI visit for every ten humans.
Both can be true. A higher bot share with lower absolute volume is what you would expect for smaller sites, and regional medians hide exactly the publisher-level variance Segura describes.
Looking forward
For UK publishers — the Telegraph sits in TollBit’s sample — the number that matters is the referral ratio, not the scrape count. Traffic once paid for content. At 227 scrapes per referral, that exchange has effectively stopped, which is the commercial injury underneath the text-and-data-mining argument, and it will not be settled by better robots.txt files.