Skip to content
Online marketing in the agent era
ROA·Marketing
Menu
SEOAI

OpenAI: Robots.txt May Not Apply to ChatGPT's Fetch Bot

OpenAI says robots.txt may not apply to ChatGPT-User, its user-triggered fetch bot. TollBit data shows 15% of AI fetchers hit disallowed URLs. What to do.

OpenAI ChatGPT fetch bot crawler ignores robots.txt disallow rules in 2026

Key Takeaways

  • In mid-August 2026, Search Engine Journal reported that OpenAI’s crawler documentation explicitly states robots.txt “may not…
  • OpenAI publishes four distinct agents, and they answer to robots.txt differently. According to OpenAI’s documentation:
  • Across the European sites in TollBit’s report, about 15% of identified AI page-fetchers reached URLs the sites had marked…

OpenAI: Robots.txt May Not Apply to ChatGPT’s Fetch Bot

OpenAI’s official crawler documentation now states that robots.txt rules “may not apply” to ChatGPT-User — the bot that fetches a web page when a person asks ChatGPT to read it — and new TollBit data shows about 15% of AI page-fetchers in Europe reached URLs sites had explicitly disallowed in the first half of 2026.

The short version

OpenAI has documented that robots.txt may not apply to ChatGPT-User, its page-fetching agent, because each request is triggered by a user rather than an automatic crawl. TollBit’s State of the Bots report for the first half of 2026 found that ChatGPT-User is disallowed by more sites than any other AI bot, yet it also reached disallowed pages on more sites than any other bot. The practical takeaway for publishers: robots.txt is a request for these user-initiated agents — not enforcement — while OAI-SearchBot, a separate agent, is the one that actually decides whether your site appears in ChatGPT search results.

Key facts

  • OpenAI documents that robots.txt rules “may not apply” to ChatGPT-User because its fetches are user-initiated.
  • About 15% of identified AI page-fetchers in Europe reached disallowed URLs in H1 2026, per TollBit.
  • ChatGPT-User, Bytespider, and Youbot each hit disallowed pages on nearly half the European sites that listed them; ChatGPT-User reached the most.
  • OAI-SearchBot — not ChatGPT-User — determines whether a site shows up in ChatGPT search answers.
  • From September 15, 2026, Cloudflare will block Training and Agent crawlers by default on new domains’ pages with ads.

What happened

In mid-August 2026, Search Engine Journal reported that OpenAI’s crawler documentation explicitly states robots.txt “may not apply” to ChatGPT-User, the agent that fetches a page when a ChatGPT user asks a question. The documentation’s exact language: “Because these actions are initiated by a user, robots.txt rules may not apply.”

The clarification landed alongside new data. TollBit’s State of the Bots report for the first half of 2026 — also covered by PPC Land — found that ChatGPT-User is disallowed by more sites than any other AI bot of its kind, and it reached disallowed pages on more sites than any other bot too. In other words, publishers are blocking it more than any other crawler, and it is bypassing those blocks more than any other crawler.

Which OpenAI bots does robots.txt actually govern?

OpenAI publishes four distinct agents, and they answer to robots.txt differently. According to OpenAI’s documentation:

  • OAI-SearchBot surfaces sites in ChatGPT search results. Disallowing it means your site won’t be shown in ChatGPT search answers, though it can still appear as navigational links. OpenAI notes it can take about 24 hours for a robots.txt change to take effect in search.
  • GPTBot crawls content for training generative AI foundation models. Disallowing it indicates your content should not be used for training.
  • OAI-AdsBot validates the safety of landing pages submitted as ChatGPT ads; OpenAI says this data is not used to train models.
  • ChatGPT-User fetches a page when a user asks ChatGPT or a Custom GPT to read it — and this is the agent where robots.txt “may not apply.”

That split is the crux: the agent that controls your ChatGPT visibility (OAI-SearchBot) is fully robots.txt-compliant, while the agent that fetches pages on demand (ChatGPT-User) carries a documented carve-out.

What does the TollBit data show about bypasses?

Across the European sites in TollBit’s report, about 15% of identified AI page-fetchers reached URLs the sites had marked disallowed. Three agents account for most of it: ChatGPT-User, Bytespider (ByteDance), and Youbot each accessed disallowed pages on nearly half of the European sites that had explicitly listed them — with ChatGPT-User reaching the most sites.

The blocking landscape is uneven. Only 9% of European websites disallow Claude-User, versus 26% in North America; Perplexity-User sits at 13% versus 26%. TollBit treats any request to a disallowed URL as a bypass regardless of what the operator claims — a stance that matters because OpenAI and Perplexity both argue the user-initiated loophole, while Anthropic states all three of its bots respect the file.

What this means (our take)

The headline is not “OpenAI ignores robots.txt.” It’s that robots.txt was never built for user-initiated fetching, and OpenAI has now written that into its documentation. The file was designed for automatic crawlers; when a human asks ChatGPT to “read this page,” the request behaves more like a browser visit than a crawl.

For operators, the real risk is a bad trade. A publisher that blanket-blocks every OpenAI agent has given up ChatGPT search visibility (via OAI-SearchBot) while keeping a fetch control (ChatGPT-User) that carries an explicit carve-out. The smarter posture is granular: allow OAI-SearchBot if you want to appear in ChatGPT answers, decide on GPTBot separately for training, and treat ChatGPT-User as a request you can log but not reliably enforce. Server logs and CDN records — not the robots.txt file — show what actually arrived. The wider shift toward enforcement at the network layer is the same dynamic behind pay-per-crawl reshaping SEO visibility and Google’s tightening of crawl budget for new sites.

The trend is toward control at the edge. Cloudflare is moving crawler decisions to the network layer: from September 15, 2026, new Cloudflare domains will have Training and Agent crawlers blocked by default on pages with ads, while Search crawlers stay allowed. Whether the user-initiated loophole survives is the open question — every major assistant now fetches pages this way.

What to do now

  1. Audit your robots.txt against all four OpenAI agents separately. Don’t use one blanket rule — decide OAI-SearchBot (visibility), GPTBot (training), OAI-AdsBot (ads), and ChatGPT-User (fetching) independently.
  2. Keep OAI-SearchBot allowed if you want ChatGPT search visibility. Blocking it removes you from ChatGPT search answers; the fetching agent doesn’t control that outcome.
  3. Treat ChatGPT-User as a request, not enforcement. Verify what actually arrives in server logs or your CDN, and escalate to a WAF rule if you truly need to block it.
  4. Lean on network-layer controls. If you use Cloudflare, watch for the September 15 default-block change and configure crawler settings at the edge rather than relying on the file alone.
  5. Track the bypass trend. Follow quarterly bot reports like TollBit’s State of the Bots, and revisit your SEO and AI search strategy as the fetch-agent landscape shifts.

FAQ

Does ChatGPT’s fetch bot respect robots.txt?

OpenAI’s documentation says robots.txt rules “may not apply” to ChatGPT-User because its page fetches are initiated by a user rather than an automatic crawl. TollBit data confirms ChatGPT-User reaches disallowed pages on more sites than any other tracked bot.

Which OpenAI bot controls ChatGPT search visibility?

OAI-SearchBot, not ChatGPT-User, determines whether your site appears in ChatGPT search results. Disallowing OAI-SearchBot removes your site from ChatGPT search answers, though it can still show up as a navigational link.

What is the difference between GPTBot and ChatGPT-User?

GPTBot crawls content that may be used to train OpenAI’s generative AI foundation models, and disallowing it opts you out of training. ChatGPT-User fetches a page on demand when a user asks ChatGPT to read it, and OpenAI says robots.txt may not apply to it.

How often do AI bots bypass robots.txt?

About 15% of identified AI page-fetchers on European sites reached disallowed URLs in the first half of 2026, according to TollBit. ChatGPT-User, Bytespider, and Youbot each hit disallowed pages on nearly half the sites that listed them.

How can I block ChatGPT-User if robots.txt doesn’t work?

Use network-layer controls such as a WAF rule or your CDN’s crawler management. Cloudflare is moving crawler controls to the network layer and will block Training and Agent crawlers by default on new domains’ pages with ads starting September 15, 2026.

Sources

R

ROA Marketing Team

ROA Marketing publishes deep, practical playbooks on PPC, SEO, and AI-driven marketing. We test everything we write about on live campaigns.

More articles →
🤖
New Course

Connect Any AI Agent to Google Ads

Build an AI agent that manages campaigns autonomously. MCC setup, OAuth, MCP server — full source code included.

$5 on Gumroad →
📘
Bestseller

Google Ads Expert — Master PPC

12 modules, real CPC benchmarks, bidding decision trees, search term audit protocol. 42,000 words.

$5 on Gumroad →
AI Transparency Disclosure

This content was created with AI assistance and reviewed by human editors before publication, in accordance with the EU AI Act (Article 50). Learn more about our AI practices →