Key takeaways
- A new 1,100-person field experiment found that removing Google’s AI features increased clicks to outside publishers, while an AI Mode-only experience reduced them.
- Blocking every AI crawler is not a clean answer: separate research found lower human traffic among large publishers that blocked generative-AI bots.
- Cloudflare now divides bots into search, agent and training categories, giving publishers a more precise choice than a blanket allow-or-block rule.
- The business goal is not maximum crawling. It is profitable discovery: allow uses that produce reach or revenue, restrict uses that extract value without a return, and measure both.
AI website traffic is creating a two-sided squeeze for publishers: AI answers reduce the number of people who click through, while crawler requests consume the very content that makes those answers useful. The sensible response is not to shut every bot out. It is to distinguish discovery bots from training bots and user-triggered agents, then set access according to what each one returns.
That distinction matters now because two developments have changed the debate. Researchers reported fresh experimental evidence in August 2026 that AI search suppresses publisher referrals. At the same time, Cloudflare is moving toward category-based bot controls, including new defaults for ad-supported pages from September 15. Together, they turn a vague fight over “AI scraping” into a measurable product and revenue decision.
The wider AI market is still growing quickly. Lapaas Voice has tracked how ChatGPT, Claude and Gemini are competing for AI users and how Google is packaging AI into paid products. The missing piece is the supplier side: publishers create the reporting, product information and explanations that these systems summarize, but often receive neither an equivalent visit nor a direct payment.
Why AI website traffic is falling
Traditional web search usually offered a list of links. A user selected one, visited the publisher and created an opportunity for advertising, subscriptions, affiliate revenue or a direct customer relationship. AI search changes the transaction. It can combine several pages into one response and satisfy the user before a visit occurs.
A preregistered field experiment by Stephanie T. Wang, Jeffrey Gleason, Yakov Bart, Christo Wilson and Danaé Metaxa tested this effect with 1,100 US-based Google users. Participants were assigned to normal Google Search, a version without AI Overviews and AI Mode, or a version that routed searches into AI Mode. The experiment happened in ordinary browsing rather than in a one-off lab task.
The August 2026 research paper reported that removing AI Overviews increased external click-through by 8.8 percentage points. Forcing AI Mode reduced external click-through by 18.8 percentage points compared with current Google Search. It also reduced the share of users clicking news sites, Wikipedia and Reddit.
Just as important, the researchers did not find a matching improvement in user experience. AI Mode reduced reported trust, usefulness, satisfaction, agency and relevance. Some users liked getting answers faster, but others reported difficulty reaching the sites they wanted and seeing fewer diverse sources.
Earlier observational evidence pointed in the same direction. A Pew Research Center analysis found users clicked a traditional result in 8% of visits when an AI summary appeared, compared with 15% when it did not. The new experiment strengthens the case because it manipulated access to AI features and measured resulting behavior.
Why blocking bots can also reduce AI website traffic
If AI systems take value without returning visitors, blocking them appears rational. A publisher can place instructions in robots.txt, use a content-delivery network or web application firewall, or deny requests at the server. But “AI bot” is no longer one job description.
A training crawler collects material that may help build a future model. A search crawler indexes pages so they can appear in results. An agent or user-triggered fetcher retrieves a page because somebody asked a live question. Blocking all three may protect content from one use while removing the publisher from another discovery channel.
A separate working paper, The Impact of LLMs on Online News Consumption and Production, analyzed high-frequency publisher data and staggered crawler blocks. Its authors found a moderate decline in publisher traffic after August 2024. More surprisingly, large publishers that blocked generative-AI bots experienced 23% lower total traffic and 14% lower real consumer traffic compared with publishers that had not yet blocked or never blocked.
That result does not prove every block causes a loss for every site. Blocking decisions may differ across publishers, and the paper is a working study rather than a final rule for operations. Still, it identifies the strategic danger: a publisher can prevent one form of extraction and unintentionally surrender future visibility.
| Bot use | What it does | Publisher value | Likely policy question |
|---|---|---|---|
| Search | Indexes pages for discovery and citations | Potential referrals and brand visibility | Does it send measurable users? |
| Agent | Fetches content for a live user task | Possible high-intent visit or citation | Is the request user-triggered and attributed? |
| Training | Collects material for model development | Little immediate referral value | Is there consent, licensing or an opt-out? |
| Mixed use | Combines several purposes under one identity | Hard to measure or control | Can the operator prove which use occurred? |
Cloudflare’s three-way control changes the decision
Cloudflare introduced separate controls for Search, Agent and Training traffic in July 2026. Its Bot Preference Sync announcement in August explained how a site’s dashboard choice can be reflected in robots.txt and enforced at the network edge. That closes a common gap in which a site politely disallowed a bot but did not technically stop it.
For Search and Agent traffic, site owners can allow requests, block them on pages that serve ads, or block them everywhere. For Training, publishers can express a “no training” preference while allowing cooperating mixed-use crawlers to retain search access—provided those operators meet transparency requirements.
Those requirements include respecting the no-training signal, offering an opt-out from AI summaries, giving publishers URL-level visibility into use, and showing that declining training does not harm traditional search results. Cloudflare says bots that cannot provide that transparency will remain blocked when training is disallowed.
From September 15, Cloudflare’s updated onboarding defaults will allow search but block training and agent use on ad-supported pages for relevant new sites, while leaving owners free to change the setting. That policy encodes an economic idea: a page funded by human attention should not automatically subsidize machine activity that bypasses the ad impression.
What publishers should measure before they block
The first metric is the crawl-to-referral ratio: how many requests does a bot make for each human visit it produces? The second is page-level value. A crawler repeatedly fetching expensive, frequently updated pages may create more infrastructure cost than a bot that checks a static archive once.
Publishers should also separate citations from sessions. An AI answer may mention a brand without creating a click, and that visibility can still influence later searches or purchases. Conversely, a crawler can generate enormous request volume without a single useful mention. Server logs, analytics and referral headers need to be joined before policy is set.
The third metric is substitutability. A commodity answer—weather, a definition or a short calculation—is easy for an AI system to reproduce. Original interviews, proprietary datasets, local reporting and continuously updated tools are harder to replace. Investing in distinctive information makes a publisher more likely to earn a citation, licensing deal or direct audience relationship.
Finally, publishers need channels that do not depend on an algorithmic middleman. Email newsletters, memberships, mobile apps, communities and events create first-party relationships. This is the same strategic logic behind businesses reducing dependence on one platform or supplier: distribution risk falls when the audience can return directly.
What the AI website traffic fight means for readers
Traffic is not only an advertising statistic. It helps pay for the reporting, reviews, databases and expert explanations that AI products use. If those economics weaken, the web may contain less original work and more recycled summaries. AI systems would then have a poorer information base.
Readers also lose something when a synthesized answer hides source differences. A regulator’s filing, a company statement and an anonymous social-media post do not carry equal weight, even if an AI interface presents them in the same tone. Clicking through reveals dates, methods, caveats and corrections that a compact answer may omit.
That is why attribution design matters. AI products should name sources clearly, link at the claim level and make it easy to inspect the original. Publishers, meanwhile, should write self-contained passages, label evidence and maintain stable URLs so both people and machines can understand what is authoritative.
The next phase of AI website traffic will therefore be shaped less by a single robots.txt line and more by negotiated access. Search discovery, live agents, training rights, citations and payments are separate products. Treating them separately gives publishers a chance to protect their work without making themselves invisible.
FAQs
What is AI website traffic?
AI website traffic includes human visits referred by AI search or chat products and automated requests made by AI crawlers or agents. Publishers should measure those two flows separately because a crawler request is not a reader.
Why does AI search reduce publisher clicks?
AI search can answer a question inside the search interface by synthesizing several sources. When the answer is sufficient, the user has less reason to visit the original pages.
Should a publisher block all AI bots?
Not automatically. Search bots, user-triggered agents and training crawlers serve different purposes. A selective policy can preserve useful discovery while restricting unlicensed training or expensive automated use.
What changes on September 15, 2026?
Cloudflare says relevant new ad-supported sites will receive defaults that preserve search while restricting training and agent activity on pages with ads, with site owners able to change those settings.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



