Published August 9, 2026

Should I block AI crawlers like GPTBot from my website?

For most small businesses that want to be recommended by AI: no — and more importantly, "AI crawler" isn't one thing. The bot that collects training data and the bot that builds the index behind ChatGPT's answers are separate crawlers with separate switches. Block the wrong one and you disappear from the answers while gaining nothing you actually wanted.

What this means

When people say "should I block AI bots," they're usually collapsing two very different decisions into one. The first is a rights question: do you want your content used to train someone's model? The second is a marketing question: do you want your business to appear when a customer asks an AI assistant for a recommendation?

The major providers treat these separately, and they say so in their own documentation. OpenAI runs distinct user agents for distinct jobs: GPTBot crawls content that may be used to train its foundation models, while OAI-SearchBot is what surfaces websites inside ChatGPT's search features. OpenAI states the settings are independent — a site owner can allow one and disallow the other — and warns that sites opted out of OAI-SearchBot "will not be shown in ChatGPT search answers" (OpenAI, 2026).

That's the whole trap in one sentence. Blocking GPTBot is a defensible choice about training data. Blocking OAI-SearchBot is a choice to be absent from the answer.

Who this applies to

Two groups should care, for opposite reasons. Businesses that live on discovery — medical and dental practices, law firms, contractors, home services, most local professional services — generally want the search-side crawlers allowed, because being findable is the entire point. Publishers, agencies, and anyone whose content is the product have a real argument for restricting training use, and that argument doesn't require giving up search visibility.

There's a third group that never made a decision at all: businesses blocking AI crawlers by accident. A CDN default, a security plugin, or a robots.txt someone copied off a forum in 2024 can quietly shut the door. Cloudflare, for instance, now offers managed robots.txt that maintains AI crawler directives on a customer's behalf (Cloudflare, 2025). That's useful if you chose it and invisible if you didn't.

How we'd diagnose it

Start with the file itself. Open your-domain.com/robots.txt in a browser — it's public, on every site, and takes ten seconds. Read the Disallow lines and note which user agents they sit under. GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and Google-Extended are the names worth finding.

Then check the layer above it. Your robots.txt can say one thing while your CDN or host enforces another at the network level, which is where a lot of accidental blocking actually lives. Server logs settle it: if a crawler is being turned away, the requests show up as blocked rather than never arriving. Worth knowing that robots.txt is thinly deployed in general — Cloudflare found only about 37% of the top 10,000 domains even have the file (Cloudflare, 2025), so plenty of sites have no stated position at all.

Available options

Benefits, limitations, and tradeoffs

The honest limitation is that robots.txt is a request, not a wall. It's a signal that well-behaved crawlers honor, and it does nothing about crawlers that ignore it. OpenAI's own documentation notes that ChatGPT-User — the agent that fetches a page because a user asked ChatGPT to visit it — is user-initiated, so robots.txt rules may not apply to it (OpenAI, 2026).

Changes also aren't instant. OpenAI notes it can take roughly 24 hours from a robots.txt update for its systems to adjust on the search side (OpenAI, 2026), and Google's crawling cadence ranges from days to months depending on the page. And blocking training data does not retroactively remove anything from models already trained.

What we know

Google's position is unusually direct, and it cuts against a lot of what gets sold as "AI SEO." Google Search Central states there are no additional requirements to appear in AI Overviews or AI Mode, and no special optimizations necessary — the same SEO fundamentals apply. It also says explicitly that you don't need to create new machine readable files, AI text files, or markup, and that no special schema.org structured data is required (Google Search Central, 2025).

On the control side, Google explains that because AI is built into Search, robots.txt directives for Googlebot are the mechanism for managing how a site is crawled for Search — meaning blocking Googlebot removes you from Search including its AI features. Google-Extended is offered separately for limiting AI training and grounding in some of Google's other systems (Google Search Central, 2025). Same architecture as OpenAI's: one switch for visibility, a different switch for training.

As for what site owners are actually doing, Cloudflare's Radar data from mid-2025 showed GPTBot as the most frequently disallowed AI user agent in robots.txt files across the top 10,000 domains (Cloudflare, 2025). Whether all of those owners understood they were making a training decision and not a visibility decision is a fair question.

Next steps

Read your robots.txt today. If you find AI crawler rules you don't remember writing, find out who wrote them and why before you change anything — sometimes there's a real reason. If you find nothing, that's a decision too, just an unmade one. Decide deliberately: search crawlers allowed if you want to be recommended, training crawlers according to how you feel about your content being used that way.

Orlando considerations

Most Central Florida service businesses we look at aren't publishers — their website exists to get found and get calls, not to be the product. For that profile, blocking search-side AI crawlers is close to pure downside. The more common local problem is the accidental block: a site built by a previous agency, on a managed host with AI protections toggled on by default, where nobody involved realized the setting had a marketing consequence. In a market as dense as Orlando's, where an AI assistant has plenty of comparable businesses to name instead, that's an expensive thing to leave unchecked.

Frequently asked questions

Does blocking GPTBot remove my business from ChatGPT?

Not by itself. OpenAI documents GPTBot and OAI-SearchBot as independent settings — GPTBot governs whether your content may be used to train foundation models, while OAI-SearchBot governs whether your site can surface in ChatGPT's search answers. Blocking OAI-SearchBot is the one that keeps you out of those answers.

Can I block AI training but stay visible in Google AI Overviews?

Partly. Google states that robots.txt directives for Googlebot are the control for how sites are crawled for Search, and that AI is built into Search — so blocking Googlebot removes you from Search including its AI features. Google-Extended is a separate control for limiting AI training and grounding in some of Google's other systems.

How do I check whether I'm already blocking AI crawlers?

Open your-domain.com/robots.txt in a browser and read it. Look for Disallow rules under user-agents like GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, or Google-Extended. Also check your CDN or hosting dashboard, since some providers apply AI crawler blocking at the network level regardless of what your robots.txt says.

Do I need an llms.txt file or special AI markup?

Not for Google. Google's documentation states plainly that you don't need to create new machine readable files, AI text files, or markup to appear in its AI features, and that there is no special schema.org structured data required. Other engines have not committed to those formats either, so treat them as optional experiments rather than prerequisites.

References

Cloudflare. (2025, July 1). Control content use for AI training with Cloudflare's managed robots.txt and blocking for monetized content. Cloudflare Blog. https://blog.cloudflare.com/control-content-use-for-ai-training/

Google Search Central. (2025, December 10). AI features and your website. Google for Developers. https://developers.google.com/search/docs/appearance/ai-features

OpenAI. (2026). Overview of OpenAI crawlers. OpenAI Developers. https://developers.openai.com/api/docs/bots

This article is for general informational purposes and isn't a guarantee of placement or performance in any AI system. ChatGPT, Perplexity, Google AI, and similar tools are operated by third parties outside AnswerFoundry's control, and their behavior and documentation change without notice. Crawler names and directives cited here were current as of publication. Results vary by business, market, and competition.

Last updated: August 9, 2026

If you're not sure what your site is currently telling AI crawlers — or whether a setting someone else made is costing you answers — that's one of the first things an AI visibility audit checks.

Find out what AI says about your business.

The audit takes two weeks. The blind spot lasts as long as you let it.

Get Your Free Answer Snapshot