The bot directory · 28 crawlers · refreshed monthly
Every AI crawler, explained
This is a maintained directory of the AI bots and user agents that crawl the web — the crawlers behind ChatGPT, Claude, Perplexity, Gemini, and the training datasets feeding every major model. Each profile explains what the bot does, whether it actually respects robots.txt, and gives copy-paste rules to allow or block it. To check what your own robots.txt currently permits, run the AI Crawler Access Checker.
- Respects robots.txt
- Partially respects robots.txt
- Ignores robots.txt
- Compliance unknown
OpenAI / 3
Anthropic / 3
Perplexity / 2
Google / 2
Common Crawl / 1
ByteDance / 1
Apple / 2
Amazon / 1
Meta / 2
Cohere / 1
Mistral / 1
Allen Institute for AI / 1
DuckDuckGo / 1
You.com / 1
Huawei / 2
Diffbot / 1
The Hive / 1
Webz.io / 1
Timpi / 1
Work with these bots, not just read about them:
- robots.txt Generator — build per-bot allow/block rules for all 28 crawlers at once.
- AI Bot Log Analyzer — see which of these bots actually visit your site.
- The complete guide to AI crawlers — background reading on how these bots differ.
Use the data in your own project:
This directory is published as an open dataset under CC BY 4.0 — grab it as JSON or CSV. Attribution appreciated; a link back to this page is plenty.