A security team blocks every bot with “AI” in its name and closes the ticket. Weeks later, marketing discovers that product pages no longer appear in conversational search, while ordinary search remains unchanged. The policy succeeded at reducing access, but it failed to distinguish model training from search discovery and user-requested retrieval.
An AI crawler robots.txt policy needs more than a copied list of user agents. Providers increasingly separate search indexing, potential model training, and real-time user access into different bots or product tokens. Blocking one may preserve another, and a crawl block may not remove a URL that the system learned about elsewhere. The correct policy maps each access path to a business decision before any directive is deployed.
Robots.txt Controls Crawling, Not Every Form of Discovery
A robots.txt file tells compliant crawlers which URLs they may request. It is a crawl-management mechanism, not a universal privacy or removal system.
Google’s robots.txt guidance warns that a disallowed page can still be indexed as a URL when other pages link to it. Because the crawler cannot read a blocked page, it also cannot see a noindex meta tag inside that page. OpenAI similarly notes that a blocked page may still be surfaced as a title and link when its URL is obtained from another provider or another page; its publisher guidance recommends noindex when the goal is to prevent that result.
Private or sensitive content should therefore rely on authentication and access control, not a voluntary crawler file. Robots rules communicate preferences to identified bots. They do not make public content confidential.
Separate four objectives before editing the file:
- Ordinary search crawling and indexing.
- Search or answer retrieval by an AI service.
- User-initiated fetching during a live request.
- Collection for potential model training.
One global Disallow: / cannot express those distinctions.
Provider Bot Names Map to Different Product Uses
The current bot structure varies by provider. Use official documentation and review it regularly because names and affected products can change.
| Provider token or control | Documented primary use | What blocking can affect | Important boundary |
|---|---|---|---|
| Googlebot | Google Search crawling | Search indexing and eligibility | Robots blocking is not a reliable removal method |
| Google-Extended | Training future Gemini models and grounding in selected Gemini or Vertex experiences | Non-Search AI use covered by the token | Google states it does not affect Google Search |
| OAI-SearchBot | Discovery and inclusion in ChatGPT search summaries and snippets | ChatGPT search visibility and citation | Placement is never guaranteed |
| GPTBot | Collection that may contribute to model training | Potential future training use | OpenAI documents it separately from search discovery |
| ChatGPT user access | Retrieval triggered by a user’s request | Ability to access a page during the task | Operational behavior differs from search indexing |
| Claude-SearchBot | Search indexing and search-result quality | Claude search visibility and accuracy | Separate from model-training collection |
| ClaudeBot | Collection that may contribute to model development | Potential future training use | Separate from user-requested access |
| Claude-User | Retrieval at a user’s direction | Live access to site content in a Claude task | Blocking may reduce user-directed visibility |
| Bing NOCACHE or NOARCHIVE | Page-level control for Bing Chat and related use | How content is included or linked in answers | Implemented as meta controls, not a separate AI crawler |
The table is a policy map, not a promise about every product surface. Always check the current provider page before deployment.

Googlebot and Google-Extended Solve Different Problems
Googlebot is the crawler used for Google Search. Blocking it can prevent Google from reading page content and can damage ordinary Search visibility, including eligibility for generative features that rely on the Search index.
Google-Extended is a product token in robots.txt. According to Google’s crawler documentation, it controls whether content crawled by Google may be used for training future Gemini models and for grounding in Gemini Apps and Grounding with Google Search on Vertex AI. It does not affect Google Search and is not used as a Search ranking signal.
Google-Extended has no separate HTTP user-agent string. Existing Google crawlers fetch the content, while the token communicates a usage preference. A log-monitoring rule that expects a Google-Extended request header will therefore miss the documented mechanism.
Google’s AI Overviews and AI Mode belong to Google Search. Site participation in those features is controlled through Search eligibility, preview directives, and Google’s generative AI Search setting, not through Google-Extended alone.
OpenAI Separates Search Discovery From Potential Training
OpenAI advises publishers to allow OAI-SearchBot when they want pages included in summaries and snippets in ChatGPT search. Its publisher and developer FAQ treats that bot separately from GPTBot, which publishers can disallow for pages they want excluded from potential training.
This separation supports a common policy: allow search discovery while declining training collection. The exact robots file should still be reviewed by engineering and legal teams because path rules, wildcards, subdomains, and CDN behavior can change the outcome.
OpenAI also describes user-driven access separately. A live ChatGPT request may cause a page fetch that is operationally different from building a search index. Decide whether your public documentation, support pages, or tools should remain accessible for those user-directed tasks.
Allowing a crawler creates eligibility, not guaranteed visibility. OpenAI states that ChatGPT search placement depends on multiple factors intended to surface relevant and reliable information.
Anthropic Uses Separate Training, Search, and User Bots
Anthropic documents three robots: ClaudeBot, Claude-SearchBot, and Claude-User. Its crawler guidance maps them to model development, search quality, and user-requested access respectively.
Blocking ClaudeBot signals that future site material should be excluded from training datasets. Blocking Claude-SearchBot can reduce visibility and accuracy in Claude search results. Blocking Claude-User prevents retrieval when a person asks Claude to access the site.
The structure makes the policy choice explicit. A company can reject training collection while preserving search and user-requested access, or apply different rules to public marketing content and licensed archives.
Anthropic states that its bots honor robots.txt and supports the non-standard Crawl-delay extension. Do not assume every provider interprets non-standard directives the same way.
Bing Uses Page Controls Alongside Bingbot
Microsoft’s approach includes existing search crawling plus page-level controls. Bing documented NOCACHE and NOARCHIVE behaviors for Bing Chat in 2023.
Microsoft said pages without either control could be included in answers and potentially used in training. NOCACHE could allow a URL, title, and snippet to appear while limiting broader use. NOARCHIVE could prevent inclusion and linking in Bing Chat. When both were present, Microsoft said it would treat the page as NOCACHE.
These behaviors are not interchangeable with blocking Bingbot. A crawler block can affect ordinary Bing discovery, while page meta directives communicate a more specific serving preference. Because the documentation predates later Copilot and AI Performance products, verify current behavior before implementing a new policy.
Build Policies by Content Class, Not by Entire Domain
Public documentation, product pages, subscriber articles, user profiles, licensed data, and internal portals should not inherit one undifferentiated AI policy.
Create a content inventory with five fields: content class, access status, desired search visibility, desired answer visibility, and training preference. Then map each class to provider-specific controls.
A typical policy might allow search and user-requested access to public documentation, disallow training collection for licensed reports, prevent all automated access to account pages, and leave ordinary search enabled for product pages. The exact choice depends on contracts, privacy obligations, infrastructure, and growth goals.
Keep sensitive systems behind authentication. Do not publish them and expect robots.txt to supply security.
Test the Effective File and the Resulting Visibility
The policy is not complete when the text file is committed. Confirm that robots.txt is publicly reachable at each relevant host and subdomain, returns the expected status, and contains no CMS or CDN override.
Test representative allowed and blocked paths. Inspect server logs for documented HTTP user agents where available, but remember that product tokens such as Google-Extended may not produce a separate request identity.
Then measure outcomes. Search consoles can reveal changes in crawling, indexing, generative impressions, or citations. Answer-level checks can reveal whether pages and brands still appear for relevant prompts. Topify can support a fixed prompt sample across AI systems, allowing teams to compare visibility and cited sources before and after policy changes.
The comparison should be directional. Different platforms refresh on different schedules, and a crawler-policy change may take time to affect discovery.

Maintain an AI Access Registry Instead of a Static List
Bot names, product boundaries, and control semantics evolve. A robots file copied from a one-year-old checklist can be technically valid and strategically wrong.
Maintain a registry containing the provider, token, documented purpose, allowed paths, blocked paths, owner, approval date, documentation URL, test method, and next review date. Review it quarterly and whenever a provider announces a new search, shopping, agent, or training product.
Version the policy and retain the previous file. A rollback is easier when the team knows which rule changed and what metric should recover.
Finally, coordinate SEO, security, legal, infrastructure, and content owners. Robots policy is no longer only a crawl-budget task. It controls whether public evidence can participate in search, generated answers, user-directed tasks, and future models.
Conclusion
An AI crawler robots.txt policy works only when it separates search, answer retrieval, user-directed access, and potential training. Google, OpenAI, Anthropic, and Microsoft expose different tokens and page controls for these purposes. A single block can remove useful discovery while leaving the original governance concern unresolved.
Start with content classes and desired outcomes. Map each provider’s current documentation to those decisions, implement the narrowest rule, and test both technical access and visibility. Keep private material behind authentication, keep a versioned registry, and review the policy as products change. The goal is not to allow or block “AI” as one category. It is to decide which systems may use which public content for which purpose.
FAQ
Does robots.txt keep a page private?
No. Robots.txt communicates crawl preferences to compliant bots. Use authentication or other access controls for private content.
Can I allow ChatGPT search but block OpenAI training?
OpenAI documents OAI-SearchBot for search discovery and GPTBot for potential training separately, allowing publishers to express different preferences.
Does blocking Google-Extended block AI Overviews?
No. Google states that Google-Extended does not affect Google Search. AI Overviews and AI Mode depend on Search eligibility and Google’s generative Search controls.
Why can a blocked URL still appear as a link?
Systems may discover the URL through other pages or providers even when they cannot crawl its content. Use the provider’s supported indexing or removal control when link removal is the actual goal.

Leave a Reply