Playbook

Should you allow, block, or charge AI crawlers?

Blocking AI crawlers in your robots.txt protects your content from being trained on and retrieved — and can also pull you out of the AI answers where customers now look. For a business that wants the referral, crawlable is usually the point, and the real risk is blocking the bots by accident.

Updated July 26, 2026Reading time 5 min

The question underneath the question

Somewhere you’ve heard you can block AI bots. Maybe a plugin offered to do it, maybe someone told you to protect your content. So here’s the blunt version: if you block or paywall the AI crawlers, do you disappear from the answers your customers are reading?

For a business whose whole goal is to be found, the honest answer is that you’re mostly working against yourself. But it depends on which bot, doing which job — so it’s worth knowing them by name before you touch anything.

The crawlers to know by name

A handful of user-agents do most of the work. You allow or disallow them in your robots.txt file, and sometimes at your CDN or firewall.

One subtlety: some of these operators run more than one agent — a bulk crawler for building models and a separate, user-triggered fetch for live answers. Blocking one doesn’t always block the other. If you’re unsure which agent does which for a given engine, check that engine’s own published crawler docs rather than guessing at a name.

Allow, block, or charge — the actual tradeoff

There are three doors, and they’re a real decision, not a formality.

Allow is the default. Your pages are crawlable, which makes you eligible to be read, retrieved, and cited. Eligible is not the same as chosen — but you can’t be chosen if you can’t be reached.

Block means a Disallow in robots.txt. It protects your content from being trained on and, for the bots that also fetch live, from being retrieved. That protection is the same move that can take you out of the answers.

Charge is the newer option. Reporting in mid-2026 framed the choice less as allow-or-block and more as which agents you’ll let cite you — with some publishers beginning to charge or paywall AI bots instead of shutting them out entirely. Treat that as an emerging option for large content operations, not a settled standard for a local business.

The catch for a local business

The risk almost nobody chooses on purpose: the dangerous block is the accidental one. An overzealous robots.txt, a security plugin with a “block bad bots” toggle, a CDN bot-fight rule — any of these can wall off GPTBot or PerplexityBot without you ever deciding to. You end up invisible to AI and never made the call. Before you weigh allow-versus-block as a strategy, confirm you haven’t already blocked yourself as an accident.

If your goal is the referral — the customer who asks an engine who to hire and gets sent to you — then being crawlable is the point. Blocking these agents to protect content you actually want quoted is self-defeating. The deliberate question for most local and professional-services businesses isn’t whether to block; it’s whether you’re accidentally blocking already.

The honest exception

This isn’t one-size-fits-all. A content-heavy publisher whose product is its writing faces a genuinely different calculus: its articles are the asset, and letting engines ingest them for free can undercut the business. For that kind of operation, blocking or charging can be a rational defense.

A business that sells a service and wants to be recommended is in the opposite position. Your website isn’t the product; it’s the thing that gets you named. Giving it away to the engines is how you get the referral. Don’t borrow a publisher’s crawler strategy when you’re not in a publisher’s fight.

What to do about it

Know the user-agents. Decide allow, block, or charge on purpose. And check that a plugin or CDN rule hasn’t already made the decision for you.

Before you spend a minute on what AI says about you, make sure it can see you at all — that’s the first thing a Nameworthy report looks at. When you’re ready to fix the whole picture, start with making your business show up in AI, or get the free report.

Common questions

If I block AI crawlers, will I disappear from AI answers?
For a business that wants to be recommended, blocking usually works against you. Several of these crawlers also feed live retrieval, so a robots.txt block can remove you from the answers customers are reading — not just from model training. You won't vanish everywhere at once, since some engines answer partly from memory rather than a live crawl, but you'd be shutting a door you almost certainly want open.
What are the main AI crawler user-agents I should know?
GPTBot (OpenAI), ClaudeBot (Anthropic), PerplexityBot (Perplexity), and Google's NotebookLM crawler, which Google renamed Google-GeminiNotebook in July 2026. Google-Extended is a related control: it's a robots.txt token that governs whether Google uses your content to improve Gemini and Vertex AI, rather than a separate crawler that fetches pages. You allow or disallow each of these in your robots.txt file, and sometimes at your CDN or firewall.

The report is free

Make sure you're not accidentally invisible to AI.

We check what six engines can see — and say — about you, in 48 hours.

Get your AI visibility report