Block GPTBot and Allow OAI-SearchBot: A Robots.txt Guide
Separate OpenAI training crawls from ChatGPT search access, with explicit rules, verification steps and limits of robots.txt.
You can declare different policies for OpenAI's training crawler and search crawler. If your intention is to opt out of training use while allowing discovery for ChatGPT search, do not treat every OpenAI user agent as interchangeable.
Choose the policy before writing rules
Official OpenAI documentation distinguishes these agents. Access does not guarantee that a page will be cited.
| Agent | Documented purpose | Policy decision |
|---|---|---|
| GPTBot | Crawling content that may be used for model training | Whether to permit that training crawl |
| OAI-SearchBot | Surfacing websites in ChatGPT search | Whether to permit search crawling |
| ChatGPT-User | Some actions initiated by users | Separate from the automatic search opt-out control; robots rules may not apply |
A user-agent name in a request can be spoofed. Declaring a policy and authenticating a request are different tasks.
An example for this specific intention
The following example allows automatic search crawling and disallows GPTBot. It is a template, not a change already made to your website.
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
Add the groups to your existing file rather than overwriting unrelated Google/Bing rules or sitemap declarations. If your existing OAI-SearchBot group intentionally excludes a path, preserve that decision. A specific group must contain the rules intended for that crawler; do not assume a broad wildcard restriction is inherited exactly as you expect.
Use the robots.txt generator to prepare a file, then compare the entire proposed file with the current one. The generator does not deploy rules or verify crawler traffic.
A four-case review
| Intended policy | GPTBot group | OAI-SearchBot group | Review question |
|---|---|---|---|
| Search allowed, training disallowed | Disallow root | Allow intended public paths | Does the search-specific group preserve excluded paths? |
| Both allowed | Allow intended paths | Allow intended paths | Are host access rules also permitting requests? |
| Both disallowed | Disallow root | Disallow root | Is the effect on discovery intentional? |
| Only selected public areas searchable | Disallow root | Explicit path policy | Have allowed and excluded examples both been tested? |
This table describes declared robots policies. It is not a claim that robots.txt is an access-control system, that every fetch follows it, or that a policy removes previously collected material.
Verify publication in layers
- Open
/robots.txton the exact hostname serving the content. Preserve a copy and timestamp. - Confirm that the intended groups are in the published response, not only the CMS editor.
- Check a public content URL for unintended login challenges, noindex directives or server errors.
- Ask the host to inspect a matching request if access is uncertain. Use current published IP information from the official documentation when verifying crawler identity.
- Allow for processing, then record new observations without interpreting a lack of requests as proof of a block.
You can use the log analyzer on supplied logs to find claimed crawler requests and response patterns. It cannot authenticate crawler identity from a name alone or show requests missing from your log sample.
Keep the search systems separate
These settings do not control Google AI Overviews. Google documents its own controls in AI features and your website. Nor does adding an llms.txt file override robots rules or create indexing eligibility.
For adjacent concepts, read robots syntax and the limits of llms.txt. Record the business policy, change date and verification evidence so the next developer does not undo a deliberate choice.