Learn · AI crawlers

Should a local business block AI crawlers?

By Cloute · Published October 1, 2026

The short answer

Usually not the ones that matter for customers. Blocking training crawlers such as GPTBot or ClaudeBot is a preference and does not hide a business from AI answers. Blocking search crawlers such as OAI-SearchBot, Claude-SearchBot or PerplexityBot, or the live fetchers, removes the business from the pages those engines read when a customer asks, and AI engines named a local business 84% of the time when they had read its own page versus 16% when they had not.

84%

named when the engine had read the business's own page

16%

named when it had not

7,582

live ChatGPT page visits in Cloute's server logs of local business pages in 2026

What are the three kinds of AI crawler?

KindExamplesWhat blocking it does
TrainingGPTBot, ClaudeBot, meta-externalagent, Google-Extended, Applebot-ExtendedKeeps future content out of model training. Does not remove you from AI search answers.
SearchOAI-SearchBot, Claude-SearchBot, PerplexityBot, GooglebotRemoves your pages from the index that engine searches when it answers. OpenAI says blocked sites will not be shown in ChatGPT search answers.
Live fetchChatGPT-User, Claude-User, Perplexity-UserStops (where honored) the engine opening your page while answering a customer. OpenAI and Perplexity say robots.txt may not or generally does not apply to theirs.

Does blocking AI crawlers hurt a local business?

Blocking search crawlers and live fetchers does. When a customer asks an AI engine who to call, the engine searches, opens a few pages, and writes its answer from them. Across 64,441 AI answers about local businesses, the business was named 83.7% of the time when the engine had read one of its own pages and 16.2% when it had not. A blocked page cannot be read.

Blocking training crawlers is different. It decides whether your content may shape future models. It does not decide whether today's ChatGPT, Claude or Perplexity can find and quote your page.

What robots.txt should a local business use?

To allow everything, which is what most local businesses want (an empty robots.txt does the same):

  • User-agent: *
  • Allow: /

To opt out of AI training but stay visible in AI search and live answers. One caution: Google says Google-Extended also controls whether its content may be used for grounding Gemini's answers, so blocking it is not purely a training choice. Common Crawl's CCBot is left out here: Common Crawl describes its archive as an open repository, and blocking it is a separate decision.

  • User-agent: GPTBot
  • Disallow: /
  • User-agent: ClaudeBot
  • Disallow: /
  • User-agent: Google-Extended
  • Disallow: /
  • User-agent: Applebot-Extended
  • Disallow: /
  • User-agent: meta-externalagent
  • Disallow: /
  • User-agent: *
  • Allow: /

Do not do this unless you mean to leave AI search, because it removes the site from ChatGPT, Claude and Perplexity search:

  • User-agent: OAI-SearchBot
  • Disallow: /
  • User-agent: Claude-SearchBot
  • Disallow: /
  • User-agent: PerplexityBot
  • Disallow: /

Does blocking Google-Extended remove me from AI Overviews?

No. Google says Google-Extended does not affect inclusion in Google Search. AI Overviews and AI Mode draw on pages Googlebot crawls for Search; Google says the way to limit what they show is snippet controls such as nosnippet, data-nosnippet, max-snippet or noindex, which also affect regular search results.

What else stops AI crawlers reading a site?

  • JavaScript-only pages. AI crawlers do not run scripts, so text that appears only after JavaScript loads is invisible to them. Put the content in the HTML the server sends. Cloute found its own website was invisible to AI crawlers for this reason in 2026, and fixed it.
  • Firewall or CDN bot blocking. Some hosting and security settings block AI crawlers by default, regardless of robots.txt. Check the setting, and your server logs, for the crawlers above.
  • Old rules. Anthropic's current tokens are ClaudeBot, Claude-SearchBot and Claude-User; rules written only for older names may not apply.

Questions people ask

Will blocking GPTBot stop ChatGPT recommending my business?

No. GPTBot is OpenAI's training crawler. ChatGPT search uses OAI-SearchBot, and pages opened during a conversation use ChatGPT-User.

Should I block all AI crawlers to protect my content?

Only if you are willing to disappear from AI answers. A local business's hours, services and location are what it wants AI engines to repeat. Blocking training crawlers alone is the narrower option.

Do AI crawlers slow down my website?

Rarely for a small site. AI crawlers made 80,420 visits in Cloute's server logs of local business pages in 2026, spread across many sites and months. If one is too busy, Anthropic supports Crawl-delay, and a firewall rule can rate-limit the rest.

How do I know if AI crawlers can read my site?

Check server logs for their user agents, check robots.txt and any firewall bot settings, and view the page source to confirm the text is in the HTML rather than added by JavaScript.

About Cloute. Cloute is an AI visibility company for local businesses. It measures how often ChatGPT, Gemini, Claude, Perplexity and Google AI Overviews recommend a business against its local competitors, and works every month to move it to the top.

See how often AI names your business.