robots.txt, the meta robots tag, and the X-Robots-Tag HTTP header all tell crawlers what they may do with your content, but only one of them is actually documented as an AI opt-out mechanism. robots.txt blocks a crawler before it fetches a page at all, and it is the only mechanism that GPTBot, ClaudeBot, Google-Extended, and CCBot officially support. The meta robots tag and X-Robots-Tag control indexing and snippet behavior after a page has already been fetched, and they are built for conventional search engines like Google rather than for AI training or retrieval crawlers. If the goal is to stop an AI crawler from requesting your content in the first place, robots.txt is the tool; if the goal is to control how a search engine indexes or displays a page it is still allowed to crawl, the meta tag or the header is the right one instead.
Three different layers of control
These three mechanisms sit at three different points in the request lifecycle, and that difference is what decides whether they can stop an AI crawler at all.
A crawler directive is simply a published rule that tells an automated bot what it is, and isn’t, supposed to access or do. A user-agent is the identifying string a bot sends with its requests — such as GPTBot or Googlebot — and it’s the string that directives match against. An HTTP header is metadata a server sends alongside a response, before or separate from the actual page content.
- robots.txt is a plain-text file published at a site’s root that a well-behaved crawler is expected to fetch and check before requesting anything else on that host. It groups rules under
User-agentblocks paired withDisalloworAllowpaths. Because it’s checked first, a crawler that honors it never requests the disallowed content’s body at all. - The meta robots tag is an HTML
<meta>element placed in a page’s<head>, carrying directives such asnoindexornofollowin itscontentattribute. A crawler has to have already fetched and parsed the HTML to see it — it’s an instruction found inside the page, not a gate placed in front of it. - The X-Robots-Tag uses the same directive vocabulary as the meta tag, but sends it as an HTTP response header rather than HTML markup, configured at the server or application layer (for example in
.htaccessor an nginx config block) instead of in the page source. That makes it useful for non-HTML files — PDFs, images, videos — where there’s no<head>to put a meta tag in.
The comparison table
Lined up side by side against official documentation, the three mechanisms differ on where they live, what they can touch, and — critically for this topic — whether any AI crawler has actually said it respects them.
| Mechanism | Where it’s set | Scope | Can it block AI crawlers specifically | Caching / propagation behavior |
|---|---|---|---|---|
| robots.txt | Plain-text file at the site root (https://example.com/robots.txt); one file per protocol, host, and port combination |
Site-wide, or narrowed to a path/directory with Disallow; a separate file is required per subdomain |
Yes — the only mechanism GPTBot, ClaudeBot, CCBot, and Google-Extended document as an opt-out, and it’s honored voluntarily, not enforced | Google caches a robots.txt file for up to 24 hours by default (longer if a refresh fails, adjustable via Cache-Control: max-age); other AI crawlers’ cache windows aren’t publicly documented, so a change may take an unknown amount of time to take effect |
| Meta robots tag | Inside the page’s HTML <head> (Google also honors it in the body) as <meta name="robots" content="..."> |
A single HTML document only; cannot be applied to non-HTML files such as PDFs or images | Not documented by any major AI crawler — OpenAI’s and Anthropic’s own bot documentation describes robots.txt only, with no stated support for parsing meta robots directives | Only takes effect once a crawler fetches and parses the page; if the page is disallowed in robots.txt, the crawler never requests it and never sees the tag at all |
| X-Robots-Tag | HTTP response header, set via server configuration or application code — not in the page body | Any file type (HTML, PDF, image, video) and can be applied in bulk via a server rule (e.g., all .pdf files) or per URL |
Same caveat as the meta tag — no AI crawler’s documentation references X-Robots-Tag as a supported opt-out signal | Same fetch-dependent behavior as the meta tag: the crawler must request the resource and read the header, so an upstream robots.txt disallow still hides it entirely |
How this plays out for GPTBot, ClaudeBot, Google-Extended, and CCBot
Every major AI crawler that documents an opt-out mechanism documents robots.txt, and none of them documents support for the meta robots tag or the X-Robots-Tag header.
GPTBot (OpenAI) is disallowed with User-agent: GPTBot. OpenAI’s own bot documentation lists four separate tokens with different jobs, and disallowing one doesn’t affect the others: GPTBot crawls content for training OpenAI’s models; ChatGPT-User fetches pages when a user or a Custom GPT triggers a live action inside ChatGPT, and OpenAI notes this bot is “not used for crawling the web in an automatic fashion,” so robots.txt rules “may not apply” to it in the same way; OAI-SearchBot crawls pages to surface them in ChatGPT’s search features, and disallowing it keeps a site out of those search answers specifically; OAI-AdsBot checks the safety of pages submitted as ads and doesn’t feed training data at all.
ClaudeBot (Anthropic) is disallowed with User-agent: ClaudeBot. Anthropic’s support documentation lists three tokens: ClaudeBot gathers web content for model training and development; Claude-User retrieves a page when a person asks Claude a direct question that requires fetching that URL; Claude-SearchBot indexes content to improve the quality of search-style results. Anthropic publishes its crawler IP ranges at claude.com/crawling/bots.json so site owners can verify a request is genuinely from one of its bots, and recommends applying the rule to every subdomain, since a robots.txt file doesn’t automatically cover subdomains.
Google-Extended is unusual: it has no HTTP user-agent string of its own. Crawling still happens under Googlebot’s normal identity, but adding User-agent: Google-Extended with Disallow: / to robots.txt tells Google that content it crawls shouldn’t be used to train future Gemini models or the Vertex AI Gemini API, or for grounding. Google states explicitly that “Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search,” which makes it the one token that lets a site separate ordinary search crawling from AI-training use of the same crawl.
CCBot (Common Crawl) is disallowed with User-agent: CCBot. Common Crawl isn’t an AI company itself, but its open dataset is a common ingredient in third-party model training pipelines, so blocking CCBot has a broader downstream effect than blocking any single company’s bot. Common Crawl publishes its crawler IP ranges at index.commoncrawl.org/ccbot.json and warns that other crawlers sometimes falsely identify themselves using the CCBot name.
Worked example: blocking multiple AI crawlers in one robots.txt file
A single robots.txt file can carry a separate rule block for each crawler, and each block only applies to the exact user-agent token it names.
User-agent: GPTBot
Disallow: /
User-agent: ChatGPT-User
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: *
Allow: /
This keeps conventional search crawlers like Googlebot and Bingbot allowed under the catch-all User-agent: * block while opting the five named AI crawlers out of the entire site. Disallow: / blocks every path; swapping it for something narrower, like Disallow: /blog/, would block only that section instead of the whole domain. Each User-agent line starts a new, independent record — rules don’t combine across blocks, and a bot only applies the rules under the exact token that matches its own identity.
The robots.txt-plus-noindex trap
Combining a robots.txt Disallow with a noindex meta tag or X-Robots-Tag on the same page cancels out the second directive, because a crawler that respects Disallow never requests the page to read it.
Google’s own documentation states this directly: if a page is disallowed from crawling, “any information about indexing or serving rules will not be found and will therefore be ignored.” The crawler simply stops at the robots.txt check and never sees the HTML or the response headers for that URL. For AI crawlers this isn’t really a trap so much as a non-issue in practice, since none of them document reading those directives anyway — but it matters if a site is relying on a mix of mechanisms across different bots and assumes a blocked page’s noindex tag is doing extra work. It isn’t; the Disallow alone is what’s taking effect.
Which one should you actually use
Pick based on what you’re trying to prevent, not on which mechanism feels more thorough.
- To stop an AI crawler from fetching your content at all — for training, retrieval, or AI search indexing — use robots.txt with that crawler’s documented token. It’s the only one any of these bots acknowledge.
- To stop a conventional search engine from indexing a specific page that you still want crawled (for internal link discovery, for example), use the meta robots tag with
noindex, and don’t block the page in robots.txt first. - To apply the same indexing control across non-HTML files, or across many URLs at once from a single server rule instead of editing every HTML file, use X-Robots-Tag.
- Remember that none of the three is enforced. They’re voluntary conventions with no authentication behind them. OpenAI’s own documentation concedes as much for
ChatGPT-User, noting that robots.txt rules “may not apply” to that particular bot since it isn’t automated crawling in the traditional sense.
FAQ
Does disallowing GPTBot in robots.txt also keep my pages out of ChatGPT’s search answers?
No. GPTBot and OAI-SearchBot are separate, independently documented tokens — disallowing one has no effect on the other. A site can block GPTBot to opt out of model training while still allowing OAI-SearchBot so pages can appear in ChatGPT’s search-style answers, or do the reverse.
Does a robots.txt Disallow remove data an AI company already collected?
Nothing in OpenAI’s, Anthropic’s, or Common Crawl’s published documentation says a new robots.txt rule retroactively deletes previously crawled or already-trained-on data. These directives govern future crawling behavior going forward; they aren’t a takedown or deletion mechanism, and compliance with them is voluntary on the crawler operator’s part in the first place.
Blocking AI crawlers is one piece of a bigger technical SEO picture: structured data and a content audit both still matter for the crawlers you do want, writing for actual search intent is what gets a page crawled favorably in the first place, and showing up in AI Overviews and chat-based search is the flip side of the same decision about which bots you let in.
Image: “NOIRLab HQ Server Racks” by NOIRLab/NSF/AURA/T. Slovinský, licensed under CC BY 4.0, via Wikimedia Commons.



