Your API is your most valuable surface and your least defended one. The same endpoints that power a partner’s integration are the ones a competitor scrapes to clone your dataset, and both requests look like plain JSON over HTTPS.
Volume alone won’t tell them apart. A legitimate customer syncing millions of records and a scraper harvesting your catalog can hit identical request counts. The difference lives in who and what is calling — the client environment, the network origin, and the access pattern.
Where per-key rate limiting breaks down
Rate limiting per API key assumes the abuser has one key and one IP. Serious scrapers have neither. They register multiple free accounts to mint a pool of keys, then rotate requests across residential and datacenter proxies so no single key or IP crosses your threshold.
The result: your limits protect you from accidents and amateurs while professional scrapers glide underneath, distributing load precisely to stay invisible. Meanwhile, over-aggressive limits are the first thing that breaks a paying customer’s legitimate bulk job.
Fingerprint the caller, not the request
Prynt shifts the question from “how many requests” to “who is making them.” For browser-originated and SDK-instrumented traffic, a stable visitorId identifies the actual device behind rotating keys and IPs, so twenty keys driven by one machine collapse into a single actor. Server-side Smart Signals add network and automation context:
- Datacenter IP origin, where scrapers overwhelmingly run.
- Residential proxy detection for the pools that dodge IP reputation.
- Automation frameworks and headless clients driving requests.
- Key-to-device fan-out: many API keys tied to one device or subnet.
- Endpoint enumeration: sequential or exhaustive access patterns no real integration produces.
Cross-referenced against Prynt’s reputation network, a device seen abusing one customer’s API arrives pre-flagged at yours.
Separate abuse from your best customers
The goal is surgical: throttle the scraper, never the integration. A workable model:
- Allowlist known integrations by key and server IP so partner traffic bypasses scoring entirely.
- Score unknown callers on network origin, device linkage, and access pattern.
- Step up, don’t slam: challenge medium-risk callers with proof-of-work or verification before hard-blocking.
- Block high-confidence abuse: distributed keys from one device over datacenter IPs enumerating endpoints.
This preserves the developer experience your real users depend on while making mass scraping expensive and slow.
Wire it into your gateway
Evaluate signals at the API gateway or an auth-request layer before the request reaches business logic. Prynt returns a result server-side in milliseconds, so you decide allow, challenge, or deny without adding meaningful latency. Read our bot detection overview for how the request-level signals fit together, then attach the caller’s device and reputation context to each key.
Log every decision. Scraping campaigns evolve, and the record of which devices, subnets, and key clusters you throttled is what lets you tighten thresholds without guessing. Instrument the response headers you return, too — a well-behaved integration that hits a limit should get a clear 429 with a Retry-After so it backs off gracefully, while an abusive client gets nothing useful to optimize against. That distinction keeps your developer experience honest even as your defenses get more aggressive.
Measure the leak you closed
Track requests denied by reason, bandwidth reclaimed, and — critically — legitimate-integration error rates, which should stay flat. Teams often discover a handful of devices behind hundreds of keys accounted for the bulk of “API growth” that was really data exfiltration.
Handle the gray-area cases
Not every high-volume caller is a scraper, and not every scraper is malicious. Search engine crawlers, uptime monitors, and your own customers’ analytics jobs all generate machine traffic you want to serve. The mistake is a blanket “block all automation” rule that severs these along with the abuse. Maintain an allowlist of verified good bots and known integration server IPs, and confirm crawler identity against published address ranges rather than trusting a user-agent string, which anyone can forge.
For everything else, prefer graduated friction over an immediate wall. A medium-risk caller can be met with a proof-of-work challenge or a short verification that a real integration passes once and caches, while a scraper’s distributed, disposable clients pay the cost on every rotation. That asymmetry — cheap for legitimate callers, expensive for abusers — is what lets you tighten controls without a wave of support tickets from partners.
You will not stop every scraper, but you can make your API a poor target: each rotation now costs a fresh device and a clean network, not just another free key.
Try the signals against your own traffic in the Prynt playground, or start free on our pricing page and put a real cost on scraping your API.
Try it free
Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.