The question underneath the question
Somewhere you’ve heard you can block AI bots. Maybe a plugin offered to do it, maybe someone told you to protect your content. So here’s the blunt version: if you block or paywall the AI crawlers, do you disappear from the answers your customers are reading?
For a business whose whole goal is to be found, the honest answer is that you’re mostly working against yourself. But it depends on which bot, doing which job — so it’s worth knowing them by name before you touch anything.
The crawlers to know by name
A handful of user-agents do most of the work. You allow or disallow them in your robots.txt file, and sometimes at your CDN or firewall.
- GPTBot — OpenAI’s crawler. Controls whether your pages are gathered for its models and search.
- ClaudeBot — Anthropic’s crawler, the one that gathers content on Claude’s behalf.
- PerplexityBot — Perplexity’s crawler. Perplexity is the most citation-driven engine, so being reachable here directly affects whether it can quote you.
- Google-Extended — not a separate crawler, but a robots.txt token that controls whether Google uses your content to improve Gemini and Vertex AI. It governs use, not a distinct fetch.
- Google’s NotebookLM crawler, which Google renamed Google-GeminiNotebook in July 2026. If you had a rule targeting the old name, that change is worth checking.
One subtlety: some of these operators run more than one agent — a bulk crawler for building models and a separate, user-triggered fetch for live answers. Blocking one doesn’t always block the other. If you’re unsure which agent does which for a given engine, check that engine’s own published crawler docs rather than guessing at a name.
Allow, block, or charge — the actual tradeoff
There are three doors, and they’re a real decision, not a formality.
Allow is the default. Your pages are crawlable, which makes you eligible to be read, retrieved, and cited. Eligible is not the same as chosen — but you can’t be chosen if you can’t be reached.
Block means a Disallow in robots.txt. It protects your content from being trained on and, for the bots that also fetch live, from being retrieved. That protection is the same move that can take you out of the answers.
Charge is the newer option. Reporting in mid-2026 framed the choice less as allow-or-block and more as which agents you’ll let cite you — with some publishers beginning to charge or paywall AI bots instead of shutting them out entirely. Treat that as an emerging option for large content operations, not a settled standard for a local business.
The catch for a local business
If your goal is the referral — the customer who asks an engine who to hire and gets sent to you — then being crawlable is the point. Blocking these agents to protect content you actually want quoted is self-defeating. The deliberate question for most local and professional-services businesses isn’t whether to block; it’s whether you’re accidentally blocking already.
The honest exception
This isn’t one-size-fits-all. A content-heavy publisher whose product is its writing faces a genuinely different calculus: its articles are the asset, and letting engines ingest them for free can undercut the business. For that kind of operation, blocking or charging can be a rational defense.
A business that sells a service and wants to be recommended is in the opposite position. Your website isn’t the product; it’s the thing that gets you named. Giving it away to the engines is how you get the referral. Don’t borrow a publisher’s crawler strategy when you’re not in a publisher’s fight.
What to do about it
Know the user-agents. Decide allow, block, or charge on purpose. And check that a plugin or CDN rule hasn’t already made the decision for you.
Before you spend a minute on what AI says about you, make sure it can see you at all — that’s the first thing a Nameworthy report looks at. When you’re ready to fix the whole picture, start with making your business show up in AI, or get the free report.