All articles Bot detection

How to Block LLM Training Crawlers Like GPTBot, ClaudeBot, and CCBot

AI companies now crawl the open web at industrial scale to build training corpora, and your original content is fair game unless you say otherwise. The problem is that a robots.txt line is a polite request, not a lock — you need a way to tell honest crawlers apart from impostors and enforce the boundary.

Who the major LLM crawlers are

Each large AI vendor operates a declared crawler with a documented user-agent. The ones worth knowing:

  • GPTBot — OpenAI’s training crawler.
  • ClaudeBot — Anthropic’s crawler for model training.
  • CCBot — Common Crawl, whose datasets feed dozens of downstream models.
  • Google-Extended — a robots.txt token that governs Gemini training without affecting Google Search.
  • PerplexityBot and Applebot-Extended — answer-engine and Apple Intelligence crawlers.

Each publishes verifiable IP ranges or supports reverse-DNS confirmation. That verification step is the whole game: a user-agent string is trivially forged, so treat it as a claim to be checked, not proof.

robots.txt: necessary but not sufficient

Start with an explicit policy. A robots.txt block like the following signals intent to compliant crawlers:

User-agent: GPTBot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Google-Extended
Disallow: /

This works for crawlers that choose to obey. It does nothing for a scraper that spoofs GPTBot from a residential proxy to look legitimate, or one that simply ignores the file. Publishers routinely find their paywalled articles in training sets despite a clean robots.txt, because the disallow was never enforced — only requested.

Verifying a crawler actually is who it claims

Before you block or allow based on user-agent, confirm the identity:

  1. Reverse DNS the source IP and check it resolves to the vendor’s domain, then forward-resolve back to the same IP.
  2. Match against published IP ranges where the vendor lists them.
  3. Watch request behavior — a real training crawler fetches broadly and slowly; a content thief targets your highest-value pages and paginates aggressively.

Prynt handles this server-side. Every request carries a stable visitorId plus Smart Signals that flag datacenter origin, proxy use, and automation traits, so a fake GPTBot coming from a rotating proxy pool is exposed even though its header looks perfect. You can read more about how declared-bot identity is confirmed in our bot detection overview.

Allow the crawlers you want, block the rest

Blocking is rarely all-or-nothing. Many sites want Googlebot and Bingbot for search, want to appear in answer engines, but do not want their catalog vacuumed into a training set. A policy engine lets you split those decisions:

  • Allow verified search indexers unconditionally.
  • Allow or rate-limit verified answer-engine crawlers based on business goals.
  • Block verified training crawlers on premium content paths.
  • Challenge or block anything claiming to be a known bot that fails verification.

The key is deciding per-crawler and per-path, then enforcing it consistently rather than hoping a text file is honored.

Enforcing at the edge, not just declaring intent

Enforcement means acting on the verification result at request time. When Prynt classifies a request, you receive a response you can branch on: serve normally, throttle, return a 403, or route to a challenge. Because classification is server-side, a scraper cannot inspect client code to learn how it was caught, and the stable visitorId ties repeat offenders together across IP rotations — the same actor cycling through a hundred proxies still resolves to one identity.

Pair this with logging. Knowing which crawlers hit which URLs, and how often, turns an abstract policy into a measurable one: you can see the day a new AI crawler shows up in your logs and decide before it has taken your whole archive.

Where the risk concentrates

Not every page needs the same protection. Concentrate enforcement where the value and the exposure are highest: original reporting and analysis, proprietary research, structured data others would pay for, and anything behind a paywall. Marketing pages and boilerplate matter far less if they end up in a training set, so spending your strictest controls there wastes effort. Map your content by how costly it would be to see reproduced, then tighten verification and blocking on the top tier while leaving low-value pages open. This keeps compliant crawlers happy on the material you do not mind sharing and reserves friction for the archive that actually earns you money.

Getting started

You do not need to rebuild your stack to draw the line. Start with an honest robots.txt so compliant crawlers self-select out, then add verification and enforcement for everyone else. Prynt is free to start, so you can classify live traffic and see how much of it is declared AI crawlers versus spoofed scrapers before committing to a plan. Explore detection on real requests in the playground, or compare tiers on our pricing page when you are ready to enforce policy in production.

Try it free

Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.

Keep reading