← All insights
September 11, 2026·Web Maintenance·4 min read

Should You Block AI Crawlers From Your Site? A Straight Answer

Cloudflare’s AI Crawl Control gave website owners a set of options they did not previously have: see which AI bots are hitting your site, block specific ones, and return an HTTP 402 “Payment Required” response with a custom message pointing crawlers toward a licensing conversation. Pay Per Crawl, still in beta, extends that toward automatic usage-based payment.

The tooling is real and it works. Whether you should use it depends almost entirely on what your website is for.

Blocking AI crawlers protects content that people pay you for. It does nothing useful for content whose entire job is to help strangers find your business.

Primo Collab

What the controls actually let you do

The important nuance is that this is not a simple on/off switch, and blanket blocking is not the default posture Cloudflare is pushing. The tools are built for negotiation rather than denial:

  • Block individual crawlers selectively rather than all at once.
  • Return a 402 response with a configurable message, for example directing a crawler to a partnerships email address or a licensing page with pricing.
  • See which bots are actually crawling you, which most site owners have never had visibility into.

Cloudflare has reported that customers are already sending over a billion 402 responses on an average day, which tells you the market is leaning toward “let’s talk about terms” rather than “stay out.”

Not all bots are the same, and this is where people get it wrong

Lumping every AI-related crawler together is the most common mistake, and it is the one that causes real damage:

Crawler typeWhat it doesBlocking effect
Training crawlersCollect content to train future modelsYou are excluded from future model knowledge. No traffic impact today
Retrieval crawlersFetch pages live to answer a user’s current questionYou disappear from AI answers that could have named you
Search crawlersIndex for traditional search resultsYou disappear from search entirely. Almost never what you want

Blocking a training crawler and blocking a retrieval crawler have completely different consequences. The first is a content licensing decision. The second is a visibility decision, and for most small businesses it is a self-inflicted wound.

When blocking makes sense

  • Your content is the product. Paid research, courses, templates, proprietary data. If someone can get the substance for free from a summary, blocking protects revenue directly.
  • You have licensing leverage. Large archives, unique datasets, or original reporting that an AI company would plausibly pay for. The 402 response exists for exactly this conversation.
  • You are under contractual restrictions on how your content may be reused.
  • Crawler volume is genuinely costing you. If aggressive crawling is driving real hosting costs, that is a straightforward operational reason regardless of strategy.

When blocking hurts you

  • Your site exists to be found. Service businesses, local businesses, and most ecommerce. Your content is marketing, and blocking marketing from being seen is not protection.
  • You are trying to get cited in AI answers. You cannot be quoted by a system you have blocked from reading you.
  • You have not checked what you would be blocking. Broad rules written quickly tend to catch search crawlers too, which is a much bigger problem than the one being solved.

A sensible default for a small business site

  1. Turn on visibility first. Look at which bots are actually crawling you before changing anything. Most owners are surprised in both directions.
  2. Leave retrieval and search crawlers alone. These are how you get found.
  3. Decide about training crawlers on principle. There is no wrong answer, but make it a decision rather than a default. If your content is marketing, letting it train models is broadly harmless and occasionally helpful.
  4. Use 402 selectively if you genuinely have content worth licensing, and put a real contact route in the message.
  5. Re-check after a month. Look at whether anything you care about changed in your traffic or AI impressions.

The mistake to avoid

Blocking everything as a precaution because it feels safer. For a business whose website is a marketing asset, that is not caution, it is opting out of discovery to protect content that was never going to be sold. Decide what your content is actually for, then set the rules to match. If the answer is “it exists so people find us,” blocking the systems people now use to find things is working against yourself.

Tell us what you're building.

We keep it small and hands-on. You're talking directly with the people building your site, not a sales rep.

See what clients say