Learn · AI crawlers

AI crawlers: the complete list for 2026

By Cloute · Published October 1, 2026

The short answer

The main AI crawlers in 2026 are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic's ClaudeBot, Claude-SearchBot and Claude-User, PerplexityBot and Perplexity-User, Amazonbot, Meta's meta-externalagent, ByteDance's Bytespider, Common Crawl's CCBot and MistralAI-User. Google-Extended and Applebot-Extended are robots.txt tokens, not separate crawlers. They fall into three jobs: collecting training data, building a search index, and fetching a page live for a user.

80,000+

AI crawler visits in Cloute's server logs of local business pages in 2026

19

different AI crawlers identified

Which AI crawlers are there?

NameCompanyPurpose (as the company states it)robots.txt tokenFetches live for a user?
GPTBotOpenAIContent that may be used to train OpenAI's modelsGPTBotNo
OAI-SearchBotOpenAISurfacing websites in ChatGPT search resultsOAI-SearchBotNo
ChatGPT-UserOpenAIVisits pages for user actions in ChatGPT and custom GPTs; robots.txt may not applyChatGPT-UserYes
ClaudeBotAnthropicWeb content that could contribute to training ClaudeClaudeBotNo
Claude-SearchBotAnthropicImproving the relevance and accuracy of Claude's search responsesClaude-SearchBotNo
Claude-UserAnthropicVisits websites when people ask Claude questions; honors robots.txtClaude-UserYes
PerplexityBotPerplexitySurfacing and linking websites in Perplexity search; not for trainingPerplexityBotNo
Perplexity-UserPerplexityVisits pages to answer users' questions; generally ignores robots.txtPerplexity-UserYes
Google-ExtendedGoogleNot a crawler. A token controlling whether content Google crawls may be used to train Gemini models and for groundingGoogle-ExtendedNo (no requests of its own)
GooglebotGoogleGoogle Search crawler; also the crawler behind AI Overviews and AI ModeGooglebotNo
Applebot-ExtendedAppleNot a crawler. A token controlling whether Applebot's data trains Apple's foundation modelsApplebot-ExtendedNo (no requests of its own)
AmazonbotAmazonImproving Amazon products and services; may be used to train AI modelsAmazonbotNo
Amzn-UserAmazonUser actions such as Alexa queries needing current information; may not follow all robots.txt rulesAmzn-UserYes
meta-externalagentMetaUses such as training foundation AI models or indexing content for productsmeta-externalagentNo
meta-externalfetcherMetaFetches individual links at a user's request; may bypass robots.txtmeta-externalfetcherYes
BytespiderByteDanceByteDance publishes no documentation for itBytespiderNot documented
CCBotCommon CrawlBuilds Common Crawl's free, open repository of web crawl dataCCBotNo
MistralAI-UserMistral AIVisits a page when a user asks Mistral's assistant a question; not for trainingMistralAI-UserYes
Purposes from each company's own crawler documentation, checked September 2026. Bytespider has no official documentation.

Which AI crawlers visit local business websites most?

CrawlerVisits
ClaudeBot17,471
GPTBot13,991
Amazonbot10,389
OAI-SearchBot8,861
ChatGPT-User7,582
Meta's AI crawlers5,120
PerplexityBot4,160
Visits in Cloute's server logs of local business pages in 2026. Only crawlers that identify themselves in their user agent are counted.

Counted by company, OpenAI's three crawlers made the most visits, 30,434 of 80,420.

What is the difference between training, search and live-fetch crawlers?

  • Training crawlers (GPTBot, ClaudeBot, meta-externalagent, and the Google-Extended and Applebot-Extended tokens) gather content that may shape a future model. Blocking them is a content-use choice.
  • Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) build the index an engine searches when it answers with live results. Blocking them removes the site from those answers.
  • Live fetchers (ChatGPT-User, Claude-User, Perplexity-User, Amzn-User, meta-externalfetcher, MistralAI-User) open a page at the moment someone's question leads there. Only Anthropic says its live fetcher honors robots.txt.

Is Google AI Overviews a separate crawler?

No. Google says robots.txt rules for Googlebot are the control for how a site is crawled for Search, including AI Overviews and AI Mode, and that a page needs to be indexed and eligible for a snippet to appear as a supporting link. Google-Extended has no user agent of its own and, Google says, does not affect inclusion in Google Search.

Do AI crawlers run JavaScript?

Generally no. A page that builds its text with JavaScript after loading looks empty to them. Content should be in the HTML the server sends. Cloute found its own website was invisible to AI crawlers for this reason in 2026, and fixed it.

Questions people ask

What is the most active AI crawler?

In Cloute's server logs of local business pages in 2026, ClaudeBot made the most visits (17,471), then GPTBot (13,991). By company, OpenAI's crawlers combined made the most.

Is Google-Extended a crawler?

No. Google says it has no separate user agent; crawling is done by Google's existing crawlers, and the Google-Extended token only controls whether that content may be used for Gemini training and grounding.

How do I verify an AI crawler is genuine?

Check the IP address. OpenAI, Anthropic, Mistral and Common Crawl publish IP lists for their crawlers. A user agent string alone can be faked.

Which AI crawlers ignore robots.txt?

By their own documentation, user-initiated fetchers may not follow it: OpenAI's ChatGPT-User, Perplexity-User, Amazon's Amzn-User and Meta's meta-externalfetcher. Anthropic says Claude-User does honor it.

About Cloute. Cloute is an AI visibility company for local businesses. It measures how often ChatGPT, Gemini, Claude, Perplexity and Google AI Overviews recommend a business against its local competitors, and works every month to move it to the top.

See how often AI names your business.