Tous les crawlers IA, et ce que coûte leur blocage
37 crawlers, agents et assistants qui lisent le web en 2026. Ce que chacun alimente, le jeton robots.txt exact, et si le bloquer vous coûte quelque chose.
Gratuit, sans compte, environ 5 secondes.
Recherche IA
Le bloquer vous coûte de la visibilitéOAI-SearchBot
OpenAIOAI-SearchBotBuilds the index ChatGPT searches when it answers. Block it and you cannot be cited in ChatGPT search results at all.
Googlebot
GoogleGooglebotThe same crawl feeds Google Search, AI Overviews and AI Mode. Still the largest single source of AI answers.
Bingbot
MicrosoftbingbotMicrosoft Copilot answers are grounded in the Bing index. No Bingbot, no Copilot citations.
PerplexityBot
PerplexityPerplexityBotBuilds the index behind Perplexity answers. Perplexity sends real referral traffic, so this one converts.
Claude-SearchBot
AnthropicClaude-SearchBotIndexes pages so Claude can find and cite them when it searches the web.
Applebot
AppleApplebotFeeds Siri, Spotlight and Apple Intelligence answers across every Apple device.
Amazonbot
AmazonAmazonbotFeeds Alexa and the Rufus shopping assistant. Matters for retail.
ExaSearchBot
ExaExaSearchBotExa is the retrieval layer behind a long tail of AI products and agents.
Bravebot
BraveBravebotBrave Search sells its index to AI products, so one crawl reaches several assistants.
YouBot
You.comYouBotIndexes pages for You.com answers.
PetalBot
HuaweiPetalBotRelevant if you care about Huawei devices and Asian markets.
Andibot
AndiAndibotSmall but growing generative search engine.
Assistant
Le bloquer vous coûte de la visibilitéChatGPT-User
OpenAIChatGPT-UserFetches your page live when a user asks ChatGPT about it. Blocked means ChatGPT tells the user it cannot open your site.
Perplexity-User
PerplexityPerplexity-UserFetches a page a Perplexity user explicitly pointed at.
Claude-User
AnthropicClaude-UserOpens your page when a Claude user asks about it.
MistralAI-User
Mistral AIMistralAI-UserFetches pages for Le Chat. Worth allowing if French and EU audiences matter to you.
DuckAssistBot
DuckDuckGoDuckAssistBotPowers the AI answers shown above DuckDuckGo results.
Google-NotebookLM
GoogleGoogle-NotebookLMPulls a page a NotebookLM user added as a source.
Meta-ExternalFetcher
Metameta-externalfetcherOpens a link a Meta AI user shared in a conversation.
Agent
Le bloquer vous coûte de la visibilitéChatGPT Agent
OpenAIChatGPT AgentBrowses and acts for a user — filling a form, comparing products, starting a checkout. Blocking it removes you from agent-driven journeys.
GoogleAgent-URLContext
GoogleGoogleAgent-URLContextRetrieves your page when a Gemini API caller passes its URL as context.
FirecrawlAgent
FirecrawlFirecrawlAgentExtraction layer used by thousands of AI apps to read your pages on demand.
Claude-Code
AnthropicClaude-CodeCoding agents read docs live. Relevant if developers are your audience.
Entraînement
À vous de voir — non compté dans le scoreGPTBot
OpenAIGPTBotCollects content to train future OpenAI models. Separate from search: blocking it does not remove you from ChatGPT search.
ClaudeBot
AnthropicClaudeBotCollects content to train Anthropic models. Independent from Claude-SearchBot.
Google-Extended
GoogleGoogle-ExtendedA robots.txt control token, not a crawler. Disallowing it keeps your content out of Gemini training and grounding without touching Search.
Applebot-Extended
AppleApplebot-ExtendedA control token, not a crawler. Disallow it to stay out of Apple model training while keeping Siri and Spotlight visibility.
Meta-ExternalAgent
Metameta-externalagentCollects content for Meta AI products and Llama training.
Bytespider
ByteDanceBytespiderAggressive crawler; widely reported to ignore robots.txt. Block at the edge if you mean it.
DeepSeekBot
DeepSeekDeepSeekBotCollects training data for DeepSeek models.
cohere-ai
Coherecohere-aiRetrieves pages for Cohere model training and grounding.
ERNIEBot
BaiduERNIEBotCollects public content for Baidu ERNIE models.
Jeu de données
À vous de voir — non compté dans le scoreDiffbot
DiffbotDiffbotBuilds a commercial knowledge graph resold to AI companies.
CCBot
Common CrawlCCBotOne crawl, many models: Common Crawl is the training corpus behind a large share of open-weight LLMs.
AI2Bot
Allen Institute for AIAI2BotNon-profit research crawler feeding open, documented datasets.
Timpibot
TimpiTimpibotBuilds a decentralised search index sold on to AI consumers.
bedrockbot
AmazonbedrockbotCrawls pages into customer-built Bedrock knowledge bases.
registry 2026.09.1 · 37 crawlers