Home/Insights/Which AI crawlers to allow, which to block, and what each choice costs

Which AI crawlers to allow, which to block, and what each choice costs

AI crawler access is decided in three places, and robots.txt is the weakest. A bot by bot decision table and a seven step audit.

BY Esh, FOUNDER, ETHOS ENGINEPUBLISHED OCTOBER 4, 20268 MIN READ
KEY TAKEAWAYS
  1. AI bots fall into three classes, training, search and user-triggered. Only the last two put you in answers.
  2. Google-Extended and Applebot-Extended are permission tokens, not crawlers. They never appear in your logs.
  3. Since 15 September 2026, Cloudflare's Block setting for AI training also stops Googlebot, Bingbot and Applebot.
  4. A robots.txt check is not an audit. Check the edge, the logs, the IPs and the raw HTML as well.

AI crawler access is decided in three places: your robots.txt file, your CDN or firewall, and the HTML your server sends. Most guides stop at the first. It is the weakest of the three. Several of the bots that matter most are documented as not bound by it, two of the best known names in it are not crawlers at all, and since 15 September 2026 one setting in Cloudflare can shut out Googlebot along with the AI bots. The useful question is not which bots to list. It is which bots can reach and read your pages, and what you lose for each one that cannot.

Three kinds of bot do three different jobs

Every major AI company now runs separate agents for separate purposes.

Training crawlers collect pages for future models. GPTBot and ClaudeBot are in this class. So is Meta-ExternalAgent, though Meta says it may also index content for its products. Blocking them is an opt-out from training. It does not touch live answers.

Search crawlers build the index an assistant searches when it answers. OAI-SearchBot, Claude-SearchBot and PerplexityBot are in this class. OpenAI's documentation says sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers. Perplexity states that PerplexityBot is not used to gather content for foundation models, so blocking it buys no training protection.

User-triggered fetchers open one page because one person asked. ChatGPT-User, Claude-User, Perplexity-User and Meta-ExternalFetcher are in this class. Anthropic says disabling Claude-User stops Claude retrieving your content in response to a user query.

Site owners have mostly understood the split. Chris Humphrey's June 2026 scan of 7,870 sites from the Majestic top 10,000 found 17.6% blocking GPTBot and 6.6% blocking OAI-SearchBot. The search and fetch classes are the ones that feed the retrieval step that decides which brands get named.

Two names in your robots.txt are not crawlers

Google-Extended and Applebot-Extended are permission tokens. Google says Google-Extended has no user agent string of its own, and that it does not affect inclusion or ranking in Google Search. Apple says Applebot-Extended does not crawl webpages. Neither will ever appear in your logs.

The consequence for Google is larger than it looks. Google's guidance says a page needs only to be indexed and eligible for a snippet to appear in AI Overviews or AI Mode. That crawl is Googlebot's. So blocking Google-Extended does not take you out of AI Overviews, and the only crawler-level way out is to block Googlebot, which takes you out of Search.

Bing has no AI token in robots.txt at all. Microsoft's controls, announced in September 2023, are meta tags. NOARCHIVE keeps a page out of Bing's chat answers and out of training. NOCACHE limits both to the URL, title and snippet. Cloudflare reports that Microsoft is targeting early 2027 for a robots.txt training preference.

Robots.txt does not bind the fetchers

OpenAI's page on ChatGPT-User says that because a user starts the request, "robots.txt rules may not apply." Perplexity says Perplexity-User generally ignores robots.txt. Google says the same of its user-triggered fetchers, which include Google-Agent. Meta says Meta-ExternalFetcher may bypass it. Anthropic is the exception among the large operators. It states that its bots honour robots.txt, and it lists Claude-User as something a site can disable.

The field data matches the documentation. Search Engine Journal reported in August 2026 on TollBit's report for the first half of the year, which found ChatGPT-User reaching pages on European sites that had disallowed it in robots.txt. A year earlier, in August 2025, Cloudflare published tests in which Perplexity fetched pages from new domains that disallowed all bots, using an undeclared browser user agent. In those same tests ChatGPT-User stopped when disallowed, so behaviour has not been constant over time.

A block is also less than total for the bots that do comply. OpenAI says a disallowed page may still surface as a link and title if the URL arrives from a third-party search provider. Perplexity says it may still index the domain, headline and a brief factual summary of a blocked page.

The edge now decides more than the file does

A firewall rule is enforced. A robots.txt line is a request. That makes the CDN the real control point, and Cloudflare changed its rules twice this year.

On 1 July 2026 it replaced its single AI bot toggle with three settings: Search, Agent and Training. On 15 September it added a fourth choice inside Training, called Disallow AI Training. This publishes a no-training preference in robots.txt and leaves Googlebot, Bingbot and Applebot free to index for search. The stricter Block setting, and Block on pages with ads, now apply to those three crawlers too. Cloudflare's own wording is that either setting affects search as well as training.

Existing settings carry over, according to Cloudflare. New domains are offered a preset. For a site that earns money from advertising, the preset allows Search, disallows Training and blocks Agent traffic on pages with ads. Agent is Cloudflare's label for bots acting in real time for a user. So a new ad-supported site on Cloudflare starts with user-triggered fetchers shut out of its monetised pages, and here the fetcher's view of robots.txt is irrelevant.

The transition was not clean. On 4 August 2026 Search Engine Journal covered a site owner's report that Googlebot and Bingbot were receiving 403 responses with the training block switched on. The report was anecdotal, but it is a reason to look at your own settings.

Cloudflare is not the only edge. Vercel's AI bot ruleset, launched in May 2025 and off by default, blocks GPTBot, ClaudeBot and PerplexityBot together. One click removes a training crawler and a search crawler alike.

Does blocking cost citations? The studies disagree

BuzzStream's March 2026 study, built on 4 million citations from 3,600 prompts, found that 82.4% of the news sites blocking OAI-SearchBot still appeared in AI citations. Goodie's study of 31 million citations and 105 news publishers concluded the opposite: blocking works against the labs that honour it.

Both can be true. BuzzStream's citation set spans ChatGPT, Gemini, AI Overviews and AI Mode. Three of those four are fed by Google's crawl, which an OpenAI block does not touch. That is our reading of the method, not BuzzStream's own conclusion. Large publishers also have routes a typical brand does not. Goodie notes that the Associated Press blocks OpenAI's crawlers and still draws most of its citations from ChatGPT through a licence.

The traffic evidence points the same way. Zhao and Berman's working paper, revised in April 2026, found that large news publishers lost about 7% of weekly traffic within six weeks of blocking AI crawlers. The data ends in May 2024 and the effect reversed for publishers ranked 101 to 500, so treat it as indicative. For a brand without a licence deal, the search crawler and the fetcher are the only routes it controls.

A page the bot reaches but cannot read

Access is not the same as legibility. Vercel and MERJ's analysis, published in December 2024, found that the crawlers from OpenAI, Anthropic, Meta, ByteDance and Perplexity did not execute JavaScript. Google's and Apple's did. Glenn Gabe's tests in August 2025 showed the result from the user's side: ChatGPT, Perplexity and Claude could not read content on a client-side rendered site, while Google's and Bing's AI surfaces could.

Both sources are more than a year old, and we found no newer published test that overturns them. The operators' crawler pages we read do not address rendering, apart from Apple's. Until someone shows otherwise, assume that text which exists only after JavaScript runs is invisible to most AI bots.

The decision table

Bot or tokenWhat it feedsWhat blocking costsRecommendation
GPTBot, ClaudeBotModel trainingNo documented loss in search answersYour policy choice
OAI-SearchBotChatGPT search indexNot shown in ChatGPT search answersAllow
Claude-SearchBotClaude's search indexLower visibility and accuracy in Claude search resultsAllow
PerplexityBotPerplexity's index, not trainingReduced to domain, headline and brief summaryAllow
ChatGPT-User, Claude-User, Perplexity-UserA live fetch for one userThe assistant cannot read the page it was asked aboutAllow, at the edge too
GooglebotSearch, AI Overviews, AI ModeRemoval from Google SearchNever block
Google-ExtendedGemini training and grounding outside SearchNothing in SearchYour policy choice
BingbotBing and CopilotRemoval from Bing. Use meta tags for narrower controlAllow
ApplebotSiri, Spotlight, SafariRemoval from Apple search featuresAllow
Applebot-ExtendedApple model trainingNothing in searchYour policy choice

How to audit access in seven steps

  1. Read robots.txt against the table. Check each token above, and check what the wildcard group disallows. Changes take about 24 hours to register at OpenAI and up to 24 hours at Perplexity.
  2. Read the edge settings. In Cloudflare, look at Search, Agent and Training. Anything set to Block under Training now stops Googlebot. Check any other firewall or bot ruleset in front of the site.
  3. Probe from outside. Request a page with each bot's user agent and once with a neutral one. LovedByAI did this across 314 sites in August 2026 and found 135 refused at least one AI crawler, while only 23 had written a rule about it. The author runs a GEO platform, and a probe from your own IP is indicative only, because firewalls can treat a claimed bot from an unlisted address differently.
  4. Read the logs at the edge. A request refused by the CDN never reaches the origin, so origin logs cannot show it. Filter by user agent and look at status codes. On Cloudflare, the AI Crawl Control crawler table lists requests and robots.txt violations per bot.
  5. Verify the IPs. OpenAI, Anthropic and Perplexity publish their IP ranges. Google documents a reverse DNS check. A user agent without a matching IP is not evidence.
  6. Check what the bot receives. Fetch the page without JavaScript. If the main text, prices or product names are missing, the bot does not have them.
  7. Confirm the outcome. ChatGPT adds utm_source=chatgpt.com to its referral links. Recheck after every CDN or hosting change.

A hypothetical example shows why the order matters. A software company allows every bot in robots.txt and assumes it is done. Step 2 finds Agent traffic blocked on its ad-carrying blog. Step 4 shows ChatGPT-User receiving 403 on those posts and OAI-SearchBot receiving 200 everywhere. Step 6 shows the pricing page returns an empty shell without JavaScript. The file was correct and two of three layers were closed.

Crawler access sits underneath AI visibility, one of the five pillars of digital authority alongside press coverage, search presence, brand entity and organic authority. It earns nothing by itself, but everything in the citation checklist assumes it. If you want a baseline across the five pillars, the free Digital Authority report is open by waitlist.

Questions

Does blocking GPTBot remove my site from ChatGPT answers?

No. OpenAI's documentation says GPTBot collects content that may be used for model training, and that ChatGPT search results come from a separate crawler, OAI-SearchBot. Blocking GPTBot is a training opt-out. Blocking OAI-SearchBot is what removes a site from ChatGPT search answers.

Does blocking Google-Extended affect AI Overviews or AI Mode?

No. Google states that Google-Extended does not affect inclusion or ranking in Google Search. AI Overviews and AI Mode use pages that are indexed and eligible for a snippet, and that crawl is controlled through Googlebot. Google-Extended only governs Gemini training and grounding outside Search.

Do AI assistants obey robots.txt when a user asks them to open a page?

Not reliably. OpenAI, Perplexity, Google and Meta each document that user-triggered fetches may not follow robots.txt, because a person requested them. Anthropic says its bots, including Claude-User, honour robots.txt. To stop a user-triggered fetcher with certainty you need a rule at the CDN or firewall.

How do I know whether AI crawlers actually reach my pages?

Read the logs at your CDN or edge, filter by user agent, and check the status codes. Then match the IP addresses against the ranges each operator publishes. A request blocked at the edge never appears in origin logs, and a user agent string on its own can be faked.

Sources

  1. OpenAI, Overview of OpenAI Crawlers developers.openai.com
  2. OpenAI Help Center, Publishers and Developers FAQ help.openai.com
  3. Anthropic, Does Anthropic crawl data from the web, and how can site owners block the crawler support.claude.com
  4. Perplexity, Perplexity Crawlers docs.perplexity.ai
  5. Perplexity Help Center, How does Perplexity follow robots.txt perplexity.ai
  6. Google, Google's common crawlers developers.google.com
  7. Google, Google User-Triggered Fetchers developers.google.com
  8. Google, Verify Requests from Google Crawlers and Fetchers developers.google.com
  9. Google Search Central, AI Features and Your Website developers.google.com
  10. Apple Support, About Applebot support.apple.com
  11. Meta for Developers, Meta Web Crawlers developers.facebook.com
  12. Microsoft Bing Webmaster Blog, Announcing new options for webmasters to control usage of their content in Bing Chat blogs.bing.com
  13. Cloudflare Blog, Have it both ways, stay discoverable in search while disallowing AI training blog.cloudflare.com
  14. Cloudflare Changelog, New options to manage AI traffic developers.cloudflare.com
  15. Cloudflare Docs, Manage AI crawlers developers.cloudflare.com
  16. Cloudflare Blog, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives blog.cloudflare.com
  17. Search Engine Journal, Cloudflare Lets Sites Disallow AI Training Without Blocking Googlebot searchenginejournal.com
  18. Search Engine Journal, Report That Cloudflare AI Bot Blocking Prevents Googlebot From Indexing Sites searchenginejournal.com
  19. Search Engine Journal, OpenAI Says Robots.txt May Not Apply To ChatGPT's Fetch Bot searchenginejournal.com
  20. Vercel, The rise of the AI crawler vercel.com
  21. Vercel Changelog, New one-click AI bot managed ruleset vercel.com
  22. GSQi (Glenn Gabe), AI Search and JavaScript Rendering gsqi.com
  23. BuzzStream, Do News Publishers That Block AI Crawlers Get Cited Less Often by AI buzzstream.com
  24. Goodie, AI Citations and News Publishers, 2026 Study higoodie.com
  25. arXiv (Zhao, Berman), Strategic Response of News Publishers to Generative AI arxiv.org
  26. Chris Humphrey, Who Blocks AI Crawlers, robots.txt in the Top 10,000 Sites chrishumphrey.ai
  27. LovedByAI via DEV Community, How to check if your site is blocking AI crawlers dev.to
E
Esh

Founder, Ethos Engine. Builds the Digital Authority Score and runs the placement marketplace. esh@ethosengine.io

NEXT STEP

See where your brand stands.

Your Digital Authority Score, who AI recommends for your target questions, and the placements that close each gap. Your first report is free.

Get your free report →
MOREINSIGHTS