Blocking every bot is a mistake — it takes you out of search results and AI answers along with the scrapers. The real objective is discrimination: welcome the crawlers that send you traffic and revenue, and stop the ones vacuuming your content. Those two groups often hit the exact same URLs, which is why user-agent filtering fails.
Which bots you actually want
Not every automated visitor is an adversary. The ones worth keeping:
- Search indexers — Googlebot, Bingbot, and regional equivalents that put you in results.
- Answer-engine crawlers — bots that surface your content in AI-generated answers, which increasingly drive referral traffic.
- Monitoring and preview bots — uptime checkers, and link-preview fetchers for social platforms and chat apps.
- Partner and feed crawlers — shopping feeds, syndication, and legitimate aggregators you have a business reason to serve.
Block these and you lose visibility, referrals, and revenue. The trick is letting them in without letting impostors ride the same user-agent through the door.
Why user-agent allowlisting fails
The naive approach — allow any request whose user-agent says “Googlebot” — is exactly backwards, because that string is the easiest thing in the world to forge. Scrapers routinely impersonate trusted crawlers precisely because so many sites allowlist on the header alone. If your rule is “trust the user-agent,” you have effectively published an open invitation to anyone who reads it.
Verifying declared identity
Real allowlisting verifies the origin, not the claim:
- Reverse DNS the source IP and confirm it resolves to the operator’s domain.
- Forward-resolve that hostname back to the original IP to defeat spoofed PTR records.
- Cross-check published IP ranges for operators that maintain them.
- Watch behavior — verified crawlers fetch broadly and respect pacing; impostors target high-value pages and race through pagination.
Prynt does this server-side. It confirms whether a request claiming to be a known crawler actually originates from that operator, and pairs the result with Smart Signals that flag datacenter origin, proxy use, and automation traits. A fake Googlebot arriving from a residential proxy is exposed even though its header is perfect, and the stable visitorId ties an impersonator’s rotating IPs back to one actor. Our scraping protection page covers this verification flow in depth.
Building the policy
With verified identity in hand, express your intent as a simple decision table:
- Verified good bot → allow, unmetered.
- Unverified request claiming to be a good bot → block or challenge; a real crawler always verifies.
- Verified but unwanted crawler (for example a training crawler on premium content) → block or rate-limit per your policy.
- Unverified automation with no bot claim → treat as a likely scraper; throttle or block on protected paths.
This gives every visitor the right experience: search engines index freely, answer engines surface your content, and scrapers hit a wall.
Keeping it maintainable
Bot ecosystems change monthly as new AI crawlers appear. Rather than hand-maintaining IP lists, lean on a service that tracks operators and updates verification logic for you, and log every bot decision so you can audit what was allowed and blocked. When a new crawler shows up in your logs, you want to make a deliberate choice about it, not discover months later that it took your archive.
Handling the long tail of smaller bots
The major search and answer engines are easy to reason about, but most sites also see a long tail of smaller crawlers: SEO tools, brand-monitoring services, academic researchers, and niche aggregators. Some are harmless or even useful; others are thin cover for scraping. Do not try to enumerate them all by hand. Instead, default this tail to the same verification-and-behavior test you apply to everyone: if it verifies as a known, wanted operator, allow it; if it makes no verifiable claim and shows automation traits on protected paths, treat it as scraping. This keeps your allowlist small and principled rather than an ever-growing list of exceptions you can never finish maintaining.
Review your allowlist on a schedule
An allowlist is not set-and-forget. New crawlers launch, operators change IP ranges, and a bot that was benign last year may behave differently today. Put a recurring review on the calendar, driven by your bot logs, so decisions stay deliberate rather than accreting by accident.
Getting started
Allowlisting good bots is about verification, not trust. Prynt is free to start, so you can verify declared crawlers on live traffic, see which “Googlebots” are real, and build an allowlist that holds up. Try it against your own traffic in the playground.
Try it free
Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.