crawlindex

Registry

The crawlers we track

23 AI crawler tokens, checked against every site's robots.txt on each crawl. The registry is version-pinned so the index stays comparable over time.

Answer-surface crawlers

These put content in front of a person the same day. Blocking one removes the site from answers users are reading right now, which is why they carry the most weight in the score.

Answer-surface crawlers
TokenOperatorWhat it doesBlocked by
GPTBotOpenAICollects training data for OpenAI models.549 (14.9%)
OAI-SearchBotOpenAIBuilds the index behind ChatGPT search results.262 (7.1%)
ChatGPT-UserOpenAIFetches a page live when a ChatGPT user asks about it.324 (8.8%)
ClaudeBotAnthropicCollects training data for Anthropic models.529 (14.4%)
Claude-UserAnthropicFetches a page live on behalf of a Claude user.266 (7.2%)
Claude-SearchBotAnthropicBuilds the index behind Claude search results.262 (7.1%)
PerplexityBotPerplexityBuilds the Perplexity answer index.383 (10.4%)
Perplexity-UserPerplexityFetches a page live for a Perplexity user query.274 (7.5%)
Google-ExtendedGoogleControls use in Gemini training and grounding. Does not affect Google Search ranking.478 (13.0%)
Applebot-ExtendedAppleControls use in Apple Intelligence training.456 (12.4%)
meta-externalagentMetaCollects training data for Meta AI.493 (13.4%)

Index and training breadth

These shape which models know the site exists at all. The effect is real but slower, so they are weighted lower.

Index and training breadth
TokenOperatorWhat it doesBlocked by
CCBotCommon CrawlFeeds the Common Crawl corpus, which most open models train on.585 (15.9%)
AmazonbotAmazonPowers Alexa and Rufus answers.445 (12.1%)
BytespiderByteDanceCollects training data for ByteDance models.548 (14.9%)
cohere-aiCohereRetrieval for Cohere assistants.386 (10.5%)
MistralAI-UserMistralFetches a page live for a Le Chat user.258 (7.0%)
YouBotYou.comBuilds the You.com answer index.331 (9.0%)
DuckAssistBotDuckDuckGoPowers DuckAssist summaries.276 (7.5%)
kagi-fetcherKagiFetches pages for Kagi assistant features.169 (4.6%)
DiffbotDiffbotBuilds a commercial knowledge graph.387 (10.5%)
Google-NotebookLMGoogleFetches sources for NotebookLM notebooks.191 (5.2%)
TavilyBotTavilyRetrieval API used by agent frameworks.199 (5.4%)
FirecrawlAgentFirecrawlRetrieval API used by agent frameworks.210 (5.7%)

Using these figures

Free to reuse in research, journalism or a product under CC BY 4.0, with credit to Fidget Labs BV. Quote the measurement date so the claim stays checkable as the index moves.

CrawlIndex by Fidget Labs BV. "AI crawler registry and blocking rates." https://crawlindex.org (measured 2026-08-09). Licensed CC BY 4.0.