# Crawlable > Crawlable vérifie que les assistants IA peuvent explorer, lire et citer > un site web, puis revérifie chaque semaine et prévient son propriétaire > par email le jour où l’accès casse. > Le scan est gratuit ; la surveillance démarre à 4,99 € par mois. Crawlable envoie trois types de requêtes depuis une seule IP — un navigateur de référence, un bot de contrôle anonyme, et une requête par crawler IA — afin de pouvoir imputer un blocage au nom d’un crawler au lieu de le supposer. Bloquer les crawlers d’entraînement de modèles ne fait jamais baisser le score : c’est un choix éditorial, pas une erreur. ## Pages principales - [Accueil](https://crawlable.fr/fr) : lancer un scan gratuit de n’importe quel site public. - [Méthodologie](https://crawlable.fr/fr/methodology) : ce que nous mesurons exactement, et ce que nos résultats ne prouvent pas. - [Tarifs](https://crawlable.fr/fr/pricing) : scan ponctuel gratuit, surveillance à 4,99 €/mois, 14,99 €/mois pour les agences. - [Annuaire des crawlers IA](https://crawlable.fr/fr/bots) : les 37 crawlers, ce que chacun alimente, et le jeton robots.txt. ## Référence des crawlers IA - [OAI-SearchBot](https://crawlable.fr/fr/bots/oai-searchbot) : Builds the index ChatGPT searches when it answers. Block it and you cannot be cited in ChatGPT search results at all. - [ChatGPT-User](https://crawlable.fr/fr/bots/chatgpt-user) : Fetches your page live when a user asks ChatGPT about it. Blocked means ChatGPT tells the user it cannot open your site. - [ChatGPT Agent](https://crawlable.fr/fr/bots/chatgpt-agent) : Browses and acts for a user — filling a form, comparing products, starting a checkout. Blocking it removes you from agent-driven journeys. - [GPTBot](https://crawlable.fr/fr/bots/gptbot) : Collects content to train future OpenAI models. Separate from search: blocking it does not remove you from ChatGPT search. - [Claude-SearchBot](https://crawlable.fr/fr/bots/claude-searchbot) : Indexes pages so Claude can find and cite them when it searches the web. - [Claude-User](https://crawlable.fr/fr/bots/claude-user) : Opens your page when a Claude user asks about it. - [ClaudeBot](https://crawlable.fr/fr/bots/claudebot) : Collects content to train Anthropic models. Independent from Claude-SearchBot. - [PerplexityBot](https://crawlable.fr/fr/bots/perplexitybot) : Builds the index behind Perplexity answers. Perplexity sends real referral traffic, so this one converts. - [Perplexity-User](https://crawlable.fr/fr/bots/perplexity-user) : Fetches a page a Perplexity user explicitly pointed at. - [Googlebot](https://crawlable.fr/fr/bots/googlebot) : The same crawl feeds Google Search, AI Overviews and AI Mode. Still the largest single source of AI answers. - [Google-Extended](https://crawlable.fr/fr/bots/google-extended) : A robots.txt control token, not a crawler. Disallowing it keeps your content out of Gemini training and grounding without touching Search. - [Google-NotebookLM](https://crawlable.fr/fr/bots/google-notebooklm) : Pulls a page a NotebookLM user added as a source. - [GoogleAgent-URLContext](https://crawlable.fr/fr/bots/googleagent-urlcontext) : Retrieves your page when a Gemini API caller passes its URL as context. - [Bingbot](https://crawlable.fr/fr/bots/bingbot) : Microsoft Copilot answers are grounded in the Bing index. No Bingbot, no Copilot citations. - [Applebot](https://crawlable.fr/fr/bots/applebot) : Feeds Siri, Spotlight and Apple Intelligence answers across every Apple device. - [Applebot-Extended](https://crawlable.fr/fr/bots/applebot-extended) : A control token, not a crawler. Disallow it to stay out of Apple model training while keeping Siri and Spotlight visibility. - [MistralAI-User](https://crawlable.fr/fr/bots/mistralai-user) : Fetches pages for Le Chat. Worth allowing if French and EU audiences matter to you. - [DuckAssistBot](https://crawlable.fr/fr/bots/duckassistbot) : Powers the AI answers shown above DuckDuckGo results. - [Amazonbot](https://crawlable.fr/fr/bots/amazonbot) : Feeds Alexa and the Rufus shopping assistant. Matters for retail. - [YouBot](https://crawlable.fr/fr/bots/youbot) : Indexes pages for You.com answers. - [ExaSearchBot](https://crawlable.fr/fr/bots/exasearchbot) : Exa is the retrieval layer behind a long tail of AI products and agents. - [Bravebot](https://crawlable.fr/fr/bots/bravebot) : Brave Search sells its index to AI products, so one crawl reaches several assistants. - [FirecrawlAgent](https://crawlable.fr/fr/bots/firecrawlagent) : Extraction layer used by thousands of AI apps to read your pages on demand. - [Diffbot](https://crawlable.fr/fr/bots/diffbot) : Builds a commercial knowledge graph resold to AI companies. - [CCBot](https://crawlable.fr/fr/bots/ccbot) : One crawl, many models: Common Crawl is the training corpus behind a large share of open-weight LLMs. - [Meta-ExternalAgent](https://crawlable.fr/fr/bots/meta-externalagent) : Collects content for Meta AI products and Llama training. - [Meta-ExternalFetcher](https://crawlable.fr/fr/bots/meta-externalfetcher) : Opens a link a Meta AI user shared in a conversation. - [Bytespider](https://crawlable.fr/fr/bots/bytespider) : Aggressive crawler; widely reported to ignore robots.txt. Block at the edge if you mean it. - [DeepSeekBot](https://crawlable.fr/fr/bots/deepseekbot) : Collects training data for DeepSeek models. - [cohere-ai](https://crawlable.fr/fr/bots/cohere-ai) : Retrieves pages for Cohere model training and grounding. - [AI2Bot](https://crawlable.fr/fr/bots/ai2bot) : Non-profit research crawler feeding open, documented datasets. - [Timpibot](https://crawlable.fr/fr/bots/timpibot) : Builds a decentralised search index sold on to AI consumers. - [PetalBot](https://crawlable.fr/fr/bots/petalbot) : Relevant if you care about Huawei devices and Asian markets. - [ERNIEBot](https://crawlable.fr/fr/bots/erniebot) : Collects public content for Baidu ERNIE models. - [Andibot](https://crawlable.fr/fr/bots/andibot) : Small but growing generative search engine. - [bedrockbot](https://crawlable.fr/fr/bots/bedrockbot) : Crawls pages into customer-built Bedrock knowledge bases. - [Claude-Code](https://crawlable.fr/fr/bots/claude-code) : Coding agents read docs live. Relevant if developers are your audience. ## Autres - [Sitemap](https://crawlable.fr/sitemap.xml) : toutes les URLs indexables, en français et en anglais. - [Version anglaise](https://crawlable.fr/en) : le même produit, en anglais.