Crawl access — infrastructure note

gptbot robots.txt: find the Disallow and allow GPTBot safely

A broad Disallow rule in robots.txt can quietly drop marketing and documentation pages from OpenAI crawlers. Because crawlers parse user-agent blocks independently and resolve path conflicts using longest-match specificity, quick edits often fail to open the routes you intended. This guide walks through robots.txt evaluation under RFC 9309, testing rules directly with curl, avoiding edge-cache traps, and verifying crawl access on your canonical host.

Updated 14 September 2026 · 14 min read · Technical

Editorial Transparency: Evaluated against RFC 9309 standards, OpenAI crawler documentation, and reproducible server responses. Our Editorial Standards →

Test your domain's live GPTBot robots.txt rules instantly

Check whether GPTBot, OAI-SearchBot, ClaudeBot, or PerplexityBot are permitted or blocked on your public pages with our free online inspector:

Test your robots.txt with AI Crawler Checker

How robots.txt is actually evaluated

Robots.txt evaluation operates strictly per host, per path, and per user-agent group. A crawler fetches https://yourdomain.com/robots.txt directly from the origin root, parses the document into distinct token blocks, selects the single most specific matching group for its identifier, and evaluates rules within that group against the requested path. Directives across different user-agent groups do not inherit or merge, and paths support only the $ and * wildcards specified in RFC 9309.

OpenAI operates three distinct crawlers with separate user-agent tokens:

Allowing GPTBot while leaving a Disallow: / for OAI-SearchBot means ChatGPT Search will still be blocked from retrieving your pages. Review and set crawler access policies per token rather than assuming a single rule covers all OpenAI services.

The longest-match rule

Under the historical 1994 robots.txt draft, the first matching directive took precedence. RFC 9309 (the IETF standard published in September 2022) formalizes the approach used by modern search engines and AI crawlers: the longest matching path wins. When multiple rules in a user-agent block match a target URL, the directive with the most specific (longest) path string decides access. If an Allow and a Disallow match identical path lengths, Allow takes precedence.

This specificity rule often trips up configuration changes. Consider this example:

User-agent: GPTBot
Allow: /
Disallow: /app

When a crawler requests /app/dashboard, both rules match: / (length 1) and /app (length 4). Because /app is longer, Disallow: /app wins and blocks the route. Now consider the inverse configuration:

User-agent: GPTBot
Allow: /docs/
Disallow: /

A request for /docs/api matches Allow: /docs/ (length 6) over Disallow: / (length 1), so the crawler gets access. A request for /pricing matches only Disallow: / and gets blocked. Adding an unanchored Allow: / does not automatically open the entire site if longer Disallow paths exist elsewhere in that group.

Wildcards also adjust the matched character count. The asterisk * matches arbitrary characters, while $ anchors the match to the end of the URL string:

User-agent: GPTBot
Disallow: /*?utm_*
Allow: /products/$

In this snippet, Allow: /products/$ matches only the exact path /products/, leaving deeper subpaths like /products/widget subject to any broader rules. Test each critical URL pattern explicitly to avoid unintended wildcard matching.

Path verification check: Pick your key public routes (such as /, /pricing, /products/, /docs/, and /blog/). For each route, find the longest matching Allow and Disallow prefixes in the active group. The longer path wins, and equal lengths resolve to Allow.

User-agent precedence: why GPTBot sees a different file than you do

Parsers select user-agent groups by exact or most specific token match rather than file order. The resolution flow follows four steps:

  1. Split the file into blocks separated by one or more User-agent: headers.
  2. Identify all groups where a User-agent line contains a case-insensitive substring match for the crawler token. GPTBot matches GPTBot directly, while the wildcard * acts as the fallback.
  3. When multiple blocks match, the group with the longest (most specific) user-agent string is selected. Directives from other blocks are ignored entirely—rules from * do not cascade into custom user-agent blocks.
  4. If no matching block exists, the crawler defaults to unrestricted access.

This isolation can lead to subtle configuration bugs:

# Group 1 — broad staging block
User-agent: *
Disallow: /

# Group 2 — targeted override
User-agent: GPTBot
Allow: /

For GPTBot, Group 2 (length 6) takes precedence over Group 1 (length 1), so only Group 2 rules apply. Public paths are accessible as intended.

However, the following setup behaves differently:

User-agent: *
Allow: /

User-agent: GPTBot
Disallow: /private/
Disallow: /app/

GPTBot evaluates only the GPTBot block. It does not inherit Allow: / from *. Unlisted public paths like /pricing remain accessible only because the default behavior for unmentioned paths is to allow access, not because of directives in the general block. Keeping each user-agent group fully self-contained prevents unexpected fallout when reorganizing rules.

User-agent matching is case-insensitive (gptbot, GptBot, and GPTBot all match), but exact spelling is required. Tokens such as GPT-Bot or ChatGPTBot will fail to match OpenAI crawlers.

When you want identical rules across multiple OpenAI bots, list them consecutively in a single block:

User-agent: GPTBot
User-agent: OAI-SearchBot
Allow: /
Disallow: /app/

This applies the same policy to both bots without duplicating rule blocks.

Allow vs Disallow: debugging when your fix does nothing

Adding an Allow directive has no effect if a competing Disallow rule has a longer matching path for the target URL. Three common path conflicts occur in production:

1. Allow path is shorter than Disallow

User-agent: GPTBot
Disallow: /products/
Allow: /

A request for /products/widget matches Disallow: /products/ (length 10) over Allow: / (length 1), leaving the page blocked. The fix is to explicitly allow the specific subtree:

User-agent: GPTBot
Disallow: /internal/
Allow: /products/
Allow: /pricing
Allow: /docs/

2. Equal path lengths

User-agent: GPTBot
Disallow: /api/
Allow: /api/

Both directives have length 5. RFC 9309 resolves ties in favor of Allow. However, because older or non-standard parsers may handle identical lengths inconsistently, clean up redundant disallows rather than relying on tie-breakers.

3. Subordinate path disallows

User-agent: GPTBot
Allow: /blog/
Disallow: /blog/drafts/
Disallow: /

Here, /blog/hello-world matches Allow: /blog/ (length 6) over Disallow: / (length 1) and is allowed. But /blog/drafts/preview matches Disallow: /blog/drafts/ (length 13) and is blocked. Subordinate rules work as intended when paths are clearly segmented.

To audit any blocked URL, determine the active user-agent group, collect all matching Allow and Disallow prefixes, and compare their character lengths. If neither matches, access is allowed by default; if both match, the longer directive governs.

You can also verify rule evaluation locally using Python's standard urllib.robotparser or Google's open-source C++ parser:

python3 -c "
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.parse(open('robots.txt').read().splitlines())
for path in ['/', '/pricing', '/app/dashboard', '/products/widget']:
    print(path, '->', 'ALLOWED' if rp.can_fetch('GPTBot', path) else 'BLOCKED')
"

Removing an obsolete Disallow: / is cleaner and less error-prone than stacking multiple Allow exceptions to punch holes through it.

Diagnosing crawler blocks with curl

Desktop browsers transmit standard browser user-agent headers and only evaluate the generic fallback block. To inspect the exact directives returned to OpenAI crawlers, send the explicit crawler token via curl -A:

# Fetch active robots.txt on the canonical host as GPTBot
curl -sSL -A "GPTBot" -H "Cache-Control: no-cache" https://yourdomain.com/robots.txt

# Probe public landing and pricing endpoints
curl -sSI -A "GPTBot" https://yourdomain.com/pricing | grep -E "HTTP|x-robots-tag"

Check the matched group in the response. If User-agent: GPTBot or OAI-SearchBot is defined, those directives govern access. Otherwise, fallback rules under User-agent: * apply.

Inspect crawler access across your public routes and generate verified robots.txt fixes with CiteAura's automated site audit.

Start free trial

The canonical redirect trap

Robots.txt files apply per host and protocol. https://example.com/robots.txt and https://www.example.com/robots.txt are evaluated independently. Crawlers inspect only the file hosted on your canonical domain (the host declared in <link rel="canonical"> and sitemaps). Modifying the apex domain while your web server redirects visitors to www leaves crawlers reading the unedited www configuration.

# Confirm host redirect destinations and live headers
curl -sSI https://example.com/robots.txt | grep -E "HTTP|Location"
curl -sSI -A "GPTBot" https://www.example.com/robots.txt | head -n 15

Production configuration patterns

Standard configuration for a web application marketing site:

# Public product pages — open to OpenAI crawlers
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ChatGPT-User
Allow: /
Disallow: /app/
Disallow: /admin/
Disallow: /account/
Disallow: /api/

# General search crawlers
User-agent: *
Allow: /
Disallow: /app/
Disallow: /admin/
Sitemap: https://www.example.com/sitemap.xml

For environments requiring strict default-deny policies, specify explicit permitted routes:

User-agent: GPTBot
Allow: /$
Allow: /pricing$
Allow: /products/
Allow: /docs/
Disallow: /

Because the specific Allow paths are longer than Disallow: /, they take precedence under longest-match evaluation.

Edge caching and log verification

When robots.txt changes fail to take effect live, stale edge cache layers (Cloudflare) or reverse proxy handler precedence (Caddy) are common causes. For complete Caddyfile server configuration blocks and journalctl log queries, refer to Jobs & Troubleshooting in Documentation.

Verification and model expectations

CiteAura audits fetch the live robots.txt from your canonical domain, resolve redirect chains, and evaluate RFC 9309 precedence rules. If a public marketing route is blocked, the audit generates an actionable engineering ticket with reproducible test commands.

Permitting crawler access in robots.txt allows compliant bots to retrieve your pages. It does not guarantee that a model will index the content, retrieve it during prompt evaluation, cite your domain, or rank your brand favorably in answers. Retrieval depends on index recency, prompt relevance, and runtime search intent.

Sources

FAQ: GPTBot and robots.txt

How do I detect GPTBot requests in server logs?

Search web server access logs (Nginx/Apache/Caddy) for the GPTBot/1.0 user-agent string. To confirm legitimate requests from OpenAI, perform reverse DNS lookups on the client IP address (which resolve to *.openai.com).

How do I block GPTBot without affecting traditional search engines?

Declare a dedicated User-agent: GPTBot block containing Disallow: / in robots.txt. RFC 9309 rules ensure that specific user-agent blocks do not affect Googlebot or other general search crawlers.

Should GPTBot be allowed to crawl every route?

No. Keep internal application routes, admin panels, user accounts, shopping carts, and private APIs closed. Allow only public marketing, documentation, and product pages.

How do I confirm a GPTBot robots.txt fix with curl?

Run curl -A GPTBot -H "Cache-Control: no-cache" https://www.example.com/robots.txt on your canonical host to inspect the active directives. Then test individual public URLs to confirm they return HTTP 200 responses.

Will an allowed crawler guarantee a ChatGPT mention?

No. An allow rule enables crawler retrieval; it does not guarantee prompt inclusion, citation generation, or answer ranking.

Why does my fix not appear even after deploy?

This is usually caused by CDN edge caching or proxy routing precedence. Purge the cache for /robots.txt and ensure static file handling takes precedence over reverse proxies.

About the Author: CiteAura Research Team

Search Systems Engineers & Evaluation Analysts

The CiteAura Research Team is the engineering working group behind the open-source CiteAura GEO engine, empirical crawler audits, and verifiable citation benchmarks across generative AI search systems.

Explore Team & Research → Editorial Standards →