GeoPromptTracker

AI crawler comparison

GPTBot vs CCBot: what's the difference?

Blocking GPTBot does not stop your content reaching OpenAI, because Common Crawl is a separate upstream source.

Short answer

GPTBot and CCBot do the same job — AI model training — so the choice between them is not either/or. If you want to opt out of that use entirely, you need a rule for each of them; blocking one leaves the other free to crawl.

Side by side

 GPTBotCCBot
OperatorOpenAICommon Crawl
PurposeAI model trainingAI model training
robots.txtRespects robots.txtRespects robots.txt
Published IP rangesYes — verifiableNone published
Official docsYesYes
Blocking costs youNo traffic — training onlyNo traffic — training only

What GPTBot does

GPTBot is OpenAI's web crawler for gathering training data. Content it collects may be used to improve future GPT models. It is OpenAI's broadest crawler, and the one most site owners mean when they talk about "blocking ChatGPT" — though blocking it does not remove you from ChatGPT's live search answers, which use OAI-SearchBot instead.

Full profile, with every robots.txt, nginx, Apache and Cloudflare rule: GPTBot.

What CCBot does

CCBot builds Common Crawl, the open web archive that has served as foundational training data for many large language models, including early GPT models. Blocking CCBot is the single broadest AI-training opt-out available, since dozens of model builders consume the Common Crawl dataset downstream.

Full profile, with every robots.txt, nginx, Apache and Cloudflare rule: CCBot.

Block both

robots.txt
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

Block GPTBot, allow CCBot

The split most sites want when the two agents do different jobs: opt out of one without giving up the other.

robots.txt
User-agent: GPTBot
Disallow: /

User-agent: CCBot
Allow: /

Check your own site

Frequently asked questions

What is the difference between GPTBot and CCBot?

Both are AI model training agents, but they belong to different operators — GPTBot to OpenAI, CCBot to Common Crawl. GPTBot: Crawls content to train OpenAI's models. CCBot: Common Crawl's crawler; its dataset is widely used to train LLMs.

Does blocking GPTBot also block CCBot?

No. robots.txt matches on the user-agent token, so a group naming GPTBot applies only to GPTBot. CCBot reads the group that names it, or the wildcard group if none does. To stop both you need a rule for each — or a wildcard rule, which would also affect every other crawler that reads it.

Should I block GPTBot, CCBot, or both?

It depends which outcome you want. Blocking GPTBot opts you out of OpenAI's model training and costs you no traffic today, because training crawlers send no visitors. Blocking CCBot opts you out of Common Crawl's training data on the same terms. The common choice is to block training agents and allow search agents, so your content stays citable without feeding model training.

Do GPTBot and CCBot both respect robots.txt?

GPTBot: OpenAI documents GPTBot's IP ranges and states it honors robots.txt disallow rules. CCBot: Common Crawl is a nonprofit with a long public record of honoring robots.txt and crawl-delay directives.

Can I tell GPTBot and CCBot apart in my server logs?

Yes — they send different user-agent strings, so grep for each token separately: grep -i "GPTBot" and grep -i "CCBot". Be aware that a user-agent header is self-declared and anything can send either string. For a check that cannot be spoofed, match the request IP against the published ranges: GPTBot at https://openai.com/gptbot.json.

Other comparisons

Part of our directory of every known AI crawler. Last verified: 2026-08-14.