AI search & GEO

Block GPTBot and Allow OAI-SearchBot: A Robots.txt Guide

Separate OpenAI training crawls from ChatGPT search access, with explicit rules, verification steps and limits of robots.txt.

Published Sep 28, 20263 min readBy RankCrab Team

You can declare different policies for OpenAI's training crawler and search crawler. If your intention is to opt out of training use while allowing discovery for ChatGPT search, do not treat every OpenAI user agent as interchangeable.

Choose the policy before writing rules

Official OpenAI documentation distinguishes these agents. Access does not guarantee that a page will be cited.

AgentDocumented purposePolicy decision
GPTBotCrawling content that may be used for model trainingWhether to permit that training crawl
OAI-SearchBotSurfacing websites in ChatGPT searchWhether to permit search crawling
ChatGPT-UserSome actions initiated by usersSeparate from the automatic search opt-out control; robots rules may not apply

A user-agent name in a request can be spoofed. Declaring a policy and authenticating a request are different tasks.

An example for this specific intention

The following example allows automatic search crawling and disallows GPTBot. It is a template, not a change already made to your website.

User-agent: GPTBot
Disallow: /

User-agent: OAI-SearchBot
Allow: /

Add the groups to your existing file rather than overwriting unrelated Google/Bing rules or sitemap declarations. If your existing OAI-SearchBot group intentionally excludes a path, preserve that decision. A specific group must contain the rules intended for that crawler; do not assume a broad wildcard restriction is inherited exactly as you expect.

Use the robots.txt generator to prepare a file, then compare the entire proposed file with the current one. The generator does not deploy rules or verify crawler traffic.

A four-case review

Intended policyGPTBot groupOAI-SearchBot groupReview question
Search allowed, training disallowedDisallow rootAllow intended public pathsDoes the search-specific group preserve excluded paths?
Both allowedAllow intended pathsAllow intended pathsAre host access rules also permitting requests?
Both disallowedDisallow rootDisallow rootIs the effect on discovery intentional?
Only selected public areas searchableDisallow rootExplicit path policyHave allowed and excluded examples both been tested?

This table describes declared robots policies. It is not a claim that robots.txt is an access-control system, that every fetch follows it, or that a policy removes previously collected material.

Verify publication in layers

  1. Open /robots.txt on the exact hostname serving the content. Preserve a copy and timestamp.
  2. Confirm that the intended groups are in the published response, not only the CMS editor.
  3. Check a public content URL for unintended login challenges, noindex directives or server errors.
  4. Ask the host to inspect a matching request if access is uncertain. Use current published IP information from the official documentation when verifying crawler identity.
  5. Allow for processing, then record new observations without interpreting a lack of requests as proof of a block.

You can use the log analyzer on supplied logs to find claimed crawler requests and response patterns. It cannot authenticate crawler identity from a name alone or show requests missing from your log sample.

Keep the search systems separate

These settings do not control Google AI Overviews. Google documents its own controls in AI features and your website. Nor does adding an llms.txt file override robots rules or create indexing eligibility.

For adjacent concepts, read robots syntax and the limits of llms.txt. Record the business policy, change date and verification evidence so the next developer does not undo a deliberate choice.

Make your next SEO change a useful one.

The free tools are ready now. Our paid product is still in development; join the launch list for an update when it is available.