GeoPromptTracker

AI crawler profile

What is ClaudeBot? How to allow or block Anthropic's crawler

User-agent string to match in robots.txt and server logs: ClaudeBot

Operator

Anthropic

Purpose

AI model training

robots.txt

Respects robots.txt

ClaudeBot is Anthropic's primary web crawler, collecting publicly available content that may be used to train the Claude family of models. It is the Anthropic equivalent of OpenAI's GPTBot, and the user agent to target if you want to opt out of Claude training specifically.

Does ClaudeBot respect robots.txt?

Anthropic documents ClaudeBot and states it respects robots.txt directives and anti-circumvention signals.

Verify it's really ClaudeBot

Anthropic publishes no IP-range file for ClaudeBot, so it can only be identified by its user-agent string — which anything can send. Treat traffic claiming this agent as unverified, and prefer a reverse-DNS check where the operator documents one before acting on it.

Block ClaudeBot with robots.txt

robots.txt — block
User-agent: ClaudeBot
Disallow: /

Explicitly allow ClaudeBot

robots.txt — allow
User-agent: ClaudeBot
Allow: /

Block ClaudeBot at the server or CDN

robots.txt is the right first step for ClaudeBot, since Anthropic honors it. Use these only if you want the block enforced rather than requested — for example to stop agents spoofing the user agent. Matching on the user-agent string still trusts a self-declared header — and no IP-range file exists for this agent, so treat it as best-effort.

nginx
if ($http_user_agent ~* "ClaudeBot") {
    return 403;
}
apache — .htaccess
RewriteEngine On
RewriteCond %{HTTP_USER_AGENT} ClaudeBot [NC]
RewriteRule .* - [F,L]
cloudflare — WAF expression
(http.user_agent contains "ClaudeBot")

Find ClaudeBot in your server logs

shell
grep -i "ClaudeBot" /var/log/nginx/access.log | wc -l

Should you block ClaudeBot?

Blocking ClaudeBot opts your content out of Anthropic's model training — a legitimate choice for original content you don't want reproduced by AI. The trade-off: models trained without your content are less likely to know your brand or recommend it unprompted. There is no direct traffic loss today, since training crawlers don't send visitors. See our guide on whether to block AI bots for the full decision framework.

Check and monitor ClaudeBot on your site

Related reading

Frequently asked questions

What is ClaudeBot?

ClaudeBot is Anthropic's web crawler for collecting AI training data. Crawls content to train Anthropic's Claude models.

Does ClaudeBot respect robots.txt?

Anthropic documents ClaudeBot and states it respects robots.txt directives and anti-circumvention signals.

How do I block ClaudeBot?

Add "User-agent: ClaudeBot" followed by "Disallow: /" to your robots.txt file. The change takes effect the next time the bot fetches your robots.txt.

Does blocking ClaudeBot hurt my Google rankings?

No. ClaudeBot is separate from Googlebot, which handles Google Search indexing. Blocking ClaudeBot has no effect on your traditional search rankings.

How can I tell if ClaudeBot is crawling my site?

Search your server access logs for the string "ClaudeBot" — for example: grep -i "ClaudeBot" /var/log/nginx/access.log | wc -l. Our free AI Bot Log Analyzer does this in your browser: paste a log file and it counts hits per AI crawler, including ClaudeBot, with per-path breakdowns.

How do I use ClaudeBot?

You don't — ClaudeBot isn't a tool you run. It's Anthropic's own crawler, operated by Anthropic, that visits your site from their infrastructure. The only control you have over it is whether you allow or block it, via robots.txt or a server rule. If you're looking to crawl other sites yourself, you'd write your own crawler or use a crawling library; sending "ClaudeBot" as your user agent would be impersonating Anthropic.

Does ClaudeBot respect crawl-delay?

Almost certainly not. Crawl-delay was never part of the original robots.txt specification — Google has publicly said it ignores the directive, and the AI crawlers that model their parsers on Google's do the same. Anthropic publishes no position on crawl-delay for ClaudeBot, so treat it as unsupported. If ClaudeBot is hitting your site harder than you want, rate-limit it at the server or CDN instead: a Cloudflare rate-limiting rule or an nginx limit_req zone matched on the user agent will actually be enforced, whereas a crawl-delay line is only a request that this agent likely never reads.

Why is ClaudeBot ignoring my robots.txt?

Anthropic states ClaudeBot honors robots.txt, so if you are still seeing hits the cause is usually one of four things rather than the bot misbehaving. First, robots.txt is cached — Anthropic may be working from a copy fetched up to 24 hours before your change. Second, the rule may not match: robots.txt user-agent matching is on a prefix of the token, and a typo or a trailing character breaks it silently. Third, the block may sit under a different user-agent group than the one this bot reads, since a bot obeys only the most specific group that matches it, not the wildcard group as well. Fourth, the traffic may be something else sending ClaudeBot as its user agent, which anything can do.

Does ClaudeBot execute JavaScript?

Treat it as no unless Anthropic says otherwise. Rendering JavaScript costs an order of magnitude more than fetching HTML, and most AI crawlers — unlike Googlebot, which runs a full headless Chrome — read the raw HTML response and stop there. The practical consequence: anything your page loads client-side after the initial response is likely invisible to ClaudeBot. If your main content is client-rendered, server-render it or pre-render it so the text exists in the first response.

What are ClaudeBot's IP ranges?

Anthropic publishes no IP-range file for ClaudeBot, which means there is no way to verify a request really came from them. Any traffic claiming this user agent should be treated as unverified. If you need certainty, block on the user agent and accept that you may be blocking impostors rather than the real crawler — and note that the absence of published ranges is itself a signal about how seriously the operator treats crawler transparency.

How often does ClaudeBot crawl my site?

There is no published schedule, and no operator commits to one. Training crawlers like ClaudeBot typically sweep in bursts rather than at a steady rate — quiet for weeks, then hundreds of requests over a day or two as a collection run reaches your domain. Larger and more-linked sites are revisited more often. The only way to know for your own site is to measure it: grep your access log for the user agent, or paste the log into our AI Bot Log Analyzer, which breaks hits down by date and path in your browser.

Where is the official Anthropic documentation for ClaudeBot?

Anthropic publishes it at https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler. That page is the authoritative source for the user-agent string and Anthropic's stated crawling policy.

Commonly confused with

Other AI crawlers

Part of our directory of every known AI crawler, refreshed monthly. Last verified: 2026-09-13.