AI crawler comparison
GPTBot vs CCBot: what's the difference?
Blocking GPTBot does not stop your content reaching OpenAI, because Common Crawl is a separate upstream source.
Short answer
GPTBot and CCBot do the same job — AI model training — so the choice between them is not either/or. If you want to opt out of that use entirely, you need a rule for each of them; blocking one leaves the other free to crawl.
Side by side
| GPTBot | CCBot | |
|---|---|---|
| Operator | OpenAI | Common Crawl |
| Purpose | AI model training | AI model training |
| robots.txt | Respects robots.txt | Respects robots.txt |
| Published IP ranges | Yes — verifiable | None published |
| Official docs | Yes | Yes |
| Blocking costs you | No traffic — training only | No traffic — training only |
What GPTBot does
GPTBot is OpenAI's web crawler for gathering training data. Content it collects may be used to improve future GPT models. It is OpenAI's broadest crawler, and the one most site owners mean when they talk about "blocking ChatGPT" — though blocking it does not remove you from ChatGPT's live search answers, which use OAI-SearchBot instead.
Full profile, with every robots.txt, nginx, Apache and Cloudflare rule: GPTBot.
What CCBot does
CCBot builds Common Crawl, the open web archive that has served as foundational training data for many large language models, including early GPT models. Blocking CCBot is the single broadest AI-training opt-out available, since dozens of model builders consume the Common Crawl dataset downstream.
Full profile, with every robots.txt, nginx, Apache and Cloudflare rule: CCBot.
Block both
User-agent: GPTBot Disallow: / User-agent: CCBot Disallow: /
Block GPTBot, allow CCBot
The split most sites want when the two agents do different jobs: opt out of one without giving up the other.
User-agent: GPTBot Disallow: / User-agent: CCBot Allow: /
Check your own site
- AI Crawler Access Checker — see which of the two your live robots.txt currently allows.
- robots.txt Generator for AI Bots — build per-bot rules for both agents and the other 26.
- AI Bot Log Analyzer — paste your access log and see which of them is actually visiting.
- Should you block AI bots? — the decision framework behind the split above.
Frequently asked questions
What is the difference between GPTBot and CCBot?
Both are AI model training agents, but they belong to different operators — GPTBot to OpenAI, CCBot to Common Crawl. GPTBot: Crawls content to train OpenAI's models. CCBot: Common Crawl's crawler; its dataset is widely used to train LLMs.
Does blocking GPTBot also block CCBot?
No. robots.txt matches on the user-agent token, so a group naming GPTBot applies only to GPTBot. CCBot reads the group that names it, or the wildcard group if none does. To stop both you need a rule for each — or a wildcard rule, which would also affect every other crawler that reads it.
Should I block GPTBot, CCBot, or both?
It depends which outcome you want. Blocking GPTBot opts you out of OpenAI's model training and costs you no traffic today, because training crawlers send no visitors. Blocking CCBot opts you out of Common Crawl's training data on the same terms. The common choice is to block training agents and allow search agents, so your content stays citable without feeding model training.
Do GPTBot and CCBot both respect robots.txt?
GPTBot: OpenAI documents GPTBot's IP ranges and states it honors robots.txt disallow rules. CCBot: Common Crawl is a nonprofit with a long public record of honoring robots.txt and crawl-delay directives.
Can I tell GPTBot and CCBot apart in my server logs?
Yes — they send different user-agent strings, so grep for each token separately: grep -i "GPTBot" and grep -i "CCBot". Be aware that a user-agent header is self-declared and anything can send either string. For a check that cannot be spoofed, match the request IP against the published ranges: GPTBot at https://openai.com/gptbot.json.
Other comparisons
Part of our directory of every known AI crawler. Last verified: 2026-08-14.