Learn · AI crawlers
AI crawlers: the complete list for 2026
By Cloute · Published October 1, 2026
The short answer
The main AI crawlers in 2026 are OpenAI's GPTBot, OAI-SearchBot and ChatGPT-User, Anthropic's ClaudeBot, Claude-SearchBot and Claude-User, PerplexityBot and Perplexity-User, Amazonbot, Meta's meta-externalagent, ByteDance's Bytespider, Common Crawl's CCBot and MistralAI-User. Google-Extended and Applebot-Extended are robots.txt tokens, not separate crawlers. They fall into three jobs: collecting training data, building a search index, and fetching a page live for a user.
AI crawler visits in Cloute's server logs of local business pages in 2026
different AI crawlers identified
Which AI crawlers are there?
| Name | Company | Purpose (as the company states it) | robots.txt token | Fetches live for a user? |
|---|---|---|---|---|
| GPTBot | OpenAI | Content that may be used to train OpenAI's models | GPTBot | No |
| OAI-SearchBot | OpenAI | Surfacing websites in ChatGPT search results | OAI-SearchBot | No |
| ChatGPT-User | OpenAI | Visits pages for user actions in ChatGPT and custom GPTs; robots.txt may not apply | ChatGPT-User | Yes |
| ClaudeBot | Anthropic | Web content that could contribute to training Claude | ClaudeBot | No |
| Claude-SearchBot | Anthropic | Improving the relevance and accuracy of Claude's search responses | Claude-SearchBot | No |
| Claude-User | Anthropic | Visits websites when people ask Claude questions; honors robots.txt | Claude-User | Yes |
| PerplexityBot | Perplexity | Surfacing and linking websites in Perplexity search; not for training | PerplexityBot | No |
| Perplexity-User | Perplexity | Visits pages to answer users' questions; generally ignores robots.txt | Perplexity-User | Yes |
| Google-Extended | Not a crawler. A token controlling whether content Google crawls may be used to train Gemini models and for grounding | Google-Extended | No (no requests of its own) | |
| Googlebot | Google Search crawler; also the crawler behind AI Overviews and AI Mode | Googlebot | No | |
| Applebot-Extended | Apple | Not a crawler. A token controlling whether Applebot's data trains Apple's foundation models | Applebot-Extended | No (no requests of its own) |
| Amazonbot | Amazon | Improving Amazon products and services; may be used to train AI models | Amazonbot | No |
| Amzn-User | Amazon | User actions such as Alexa queries needing current information; may not follow all robots.txt rules | Amzn-User | Yes |
| meta-externalagent | Meta | Uses such as training foundation AI models or indexing content for products | meta-externalagent | No |
| meta-externalfetcher | Meta | Fetches individual links at a user's request; may bypass robots.txt | meta-externalfetcher | Yes |
| Bytespider | ByteDance | ByteDance publishes no documentation for it | Bytespider | Not documented |
| CCBot | Common Crawl | Builds Common Crawl's free, open repository of web crawl data | CCBot | No |
| MistralAI-User | Mistral AI | Visits a page when a user asks Mistral's assistant a question; not for training | MistralAI-User | Yes |
Which AI crawlers visit local business websites most?
| Crawler | Visits |
|---|---|
| ClaudeBot | 17,471 |
| GPTBot | 13,991 |
| Amazonbot | 10,389 |
| OAI-SearchBot | 8,861 |
| ChatGPT-User | 7,582 |
| Meta's AI crawlers | 5,120 |
| PerplexityBot | 4,160 |
Counted by company, OpenAI's three crawlers made the most visits, 30,434 of 80,420.
What is the difference between training, search and live-fetch crawlers?
- Training crawlers (GPTBot, ClaudeBot, meta-externalagent, and the Google-Extended and Applebot-Extended tokens) gather content that may shape a future model. Blocking them is a content-use choice.
- Search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot) build the index an engine searches when it answers with live results. Blocking them removes the site from those answers.
- Live fetchers (ChatGPT-User, Claude-User, Perplexity-User, Amzn-User, meta-externalfetcher, MistralAI-User) open a page at the moment someone's question leads there. Only Anthropic says its live fetcher honors robots.txt.
Is Google AI Overviews a separate crawler?
No. Google says robots.txt rules for Googlebot are the control for how a site is crawled for Search, including AI Overviews and AI Mode, and that a page needs to be indexed and eligible for a snippet to appear as a supporting link. Google-Extended has no user agent of its own and, Google says, does not affect inclusion in Google Search.
Do AI crawlers run JavaScript?
Generally no. A page that builds its text with JavaScript after loading looks empty to them. Content should be in the HTML the server sends. Cloute found its own website was invisible to AI crawlers for this reason in 2026, and fixed it.
Read next
Research: 80,000+ AI crawler visits to local business websites What is GPTBot? What is OAI-SearchBot? What is ChatGPT-User? What is ClaudeBot? (and Claude-User, Claude-SearchBot) What is PerplexityBot? (and Perplexity-User) Should a local business block AI crawlers?Questions people ask
What is the most active AI crawler?
In Cloute's server logs of local business pages in 2026, ClaudeBot made the most visits (17,471), then GPTBot (13,991). By company, OpenAI's crawlers combined made the most.
Is Google-Extended a crawler?
No. Google says it has no separate user agent; crawling is done by Google's existing crawlers, and the Google-Extended token only controls whether that content may be used for Gemini training and grounding.
How do I verify an AI crawler is genuine?
Check the IP address. OpenAI, Anthropic, Mistral and Common Crawl publish IP lists for their crawlers. A user agent string alone can be faked.
Which AI crawlers ignore robots.txt?
By their own documentation, user-initiated fetchers may not follow it: OpenAI's ChatGPT-User, Perplexity-User, Amazon's Amzn-User and Meta's meta-externalfetcher. Anthropic says Claude-User does honor it.
About Cloute. Cloute is an AI visibility company for local businesses. It measures how often ChatGPT, Gemini, Claude, Perplexity and Google AI Overviews recommend a business against its local competitors, and works every month to move it to the top.
