Model trainingYour call — not scored
cohere-ai
Retrieves pages for Cohere model training and grounding.
- robots.txt token
cohere-aicohere-training-data-crawler- Operator
- Cohere ↗
- Powers
- Cohere enterprise models
- Honours robots.txt
- Not documented
Allow it
User-agent: cohere-ai
Allow: /Block it
User-agent: cohere-ai
User-agent: cohere-training-data-crawler
Disallow: /Need the whole file, with every crawler and your private paths — and the llms.txt to go with it? The generator writes them. AI config generator →
User-Agent
cohere-ai/1.0 (+https://cohere.com) Crawlable/1.0; +https://crawlable.fr/probeThis is the string we send when probing. The Crawlable suffix identifies us in your logs.
Check your own site
robots.txt is only half the answer: your CDN can refuse cohere-ai before it ever reads the file. A scan checks both.
Free, no account, about 5 seconds.