All articles Network & IP

Verifying Declared Bots With Reverse DNS Before You Trust the User-Agent

“User-Agent: Googlebot” is a string, and strings are free. Before you allowlist any crawler, you have to prove it is who it claims to be — otherwise your allowlist is an open door with a helpful label. Reverse DNS verification is how that proof works.

The impersonation problem

Most sites want to treat search engines and answer engines generously — index freely, no rate limits, full access. The obvious way to recognize them is the user-agent header. The obvious problem is that any scraper can send the identical header. Attackers impersonate trusted crawlers precisely because allowlisting on user-agent is so common, so the very rule meant to help good bots becomes the vector that lets bad ones in.

How forward-confirmed reverse DNS works

The reliable check has two directions, and you need both:

  1. Reverse lookup (PTR) — take the request’s source IP and resolve it to a hostname. A real Googlebot IP resolves to a googlebot.com or google.com host.
  2. Forward lookup (A/AAAA) — resolve that hostname back to an IP address and confirm it matches the original source IP.

The second step is what makes it trustworthy. An attacker can point a PTR record on their own IP at crawl-instance.googlebot.com, but they cannot make Google’s DNS forward-resolve that hostname to the attacker’s address. Only the operator controls the forward record, so the round trip proves genuine ownership.

Layering IP ranges and behavior

Reverse DNS is strong for operators that maintain proper records, but you can reinforce it:

  • Published IP ranges — several crawler operators publish their address blocks; matching against them adds a second independent check.
  • Behavioral consistency — a verified crawler fetches broadly and respects pacing, while an impostor that somehow passes DNS still targets high-value pages and paginates aggressively.
  • Automation traits — headless or instrumented-browser fingerprints that no legitimate first-party crawler would exhibit.

Prynt runs these checks server-side and combines them into a single verdict. It confirms declared-crawler identity, and pairs that with Smart Signals for datacenter origin, proxy use, and automation — so a request claiming to be a search bot from a residential proxy is flagged even if its header is flawless. Our network overview explains how origin intelligence feeds these decisions.

The other half: bots that never declare

Reverse DNS only helps when a client makes a claim you can verify. The stealthier threat is the scraper that sends a plain browser user-agent and never pretends to be a bot at all. Reverse DNS says nothing about it, because there is nothing to confirm. That is why verification alone is incomplete — you also need signals that catch undeclared automation:

  • Headless and automation-framework fingerprints.
  • Residential-proxy and datacenter-origin indicators.
  • A stable device identity that ties a rotating scraper back to one actor across IP changes.

Combining declared-bot verification with undeclared-automation detection covers both halves of the problem.

Putting it into practice

A workable policy looks like this:

  1. If a request claims to be a known crawler, verify it with forward-confirmed reverse DNS and IP-range checks.
  2. Allow verified good bots; block or challenge failed claims — a real crawler always verifies.
  3. For requests making no bot claim, evaluate automation and network signals and apply your scraping controls.

Let a service maintain the operator lists and verification logic so you are not hand-editing IP ranges every time a new AI crawler launches.

Caching verification results

Forward-confirmed reverse DNS involves live lookups, and doing them on every request adds latency. Cache the result per identity for a sensible window so a verified crawler is not re-resolved on each fetch, while still re-checking often enough that a compromised or reassigned IP does not keep its trusted status forever. A managed service handles this caching for you, along with the operator lists, so verification stays fast without going stale.

When operators lack proper records

Not every legitimate crawler maintains clean reverse DNS, and some newer AI operators are still maturing their infrastructure. For those, fall back to published IP ranges and behavioral consistency, and treat a missing PTR record as a reason to rate-limit rather than fully trust, not necessarily to block. The point is to grade confidence, not to demand perfection from every operator, so a genuinely good but imperfectly configured crawler is not punished as harshly as a spoofer.

Getting started

Never trust a user-agent you have not verified. Prynt is free to start, so you can run declared-bot verification and undeclared-automation detection on live traffic and see how many “Googlebots” fail the round trip. Try it against your own requests in the playground.

Try it free

Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.

Keep reading