Repricing tools let a competitor track your catalog and undercut you within minutes of a price change. If your margins keep eroding on exactly the SKUs you moved, automated price scraping is a likely cause — and it usually hides inside traffic that looks human.
Why price scraping is hard to see
Modern scrapers do not look like the crude bots of a decade ago. They run headless Chrome to execute your JavaScript, rotate through residential proxy pools so every request comes from a fresh consumer IP, and pace themselves to blend into normal shopper volume. Rate limiting by IP fails because no single IP makes enough requests to trip a threshold. User-agent filtering fails because they send a normal browser string.
What gives them away is the pattern across the whole session, not any single request: methodical pagination through category pages, requests for product JSON endpoints no human clicks directly, no mouse movement or scroll behavior, and identical timing rhythms.
Signals that separate a scraper from a shopper
Effective detection layers several independent signals so no single evasion defeats it:
- Automation traits — headless browser fingerprints, missing or spoofed canvas and WebGL, driver artifacts from Puppeteer or Playwright.
- Network origin — datacenter ASNs, known proxy ranges, and residential-proxy indicators that reveal an IP is being resold for automation.
- Behavioral absence — no pointer entropy, instantaneous navigation, and access to endpoints only a script would target.
- Identity persistence — a stable device identity that survives IP and cookie rotation, so the same scraper is recognized across thousands of “different” visitors.
Prynt combines these server-side and returns a single verdict per request, along with reason codes explaining why. That stable visitorId is the piece IP-based tools miss: when a scraper cycles through a residential proxy network, the identity stays constant even as the source address changes on every hit. See how the reputation layer shares abuse signals across sites in our network overview.
Protecting the endpoints that matter
Scrapers go where the data is cleanest, which is usually your structured endpoints rather than rendered HTML. Prioritize:
- Product and pricing APIs — the JSON your own frontend calls. These are the richest, easiest-to-parse target.
- Category and search pages — the enumeration surface a scraper walks to discover every SKU.
- Availability and inventory checks — real-time stock data that competitors want as badly as price.
Put device-level verification in front of these paths so an unverified automated client gets throttled or blocked before it can enumerate your catalog, while a logged-in shopper on a real phone sails through.
Responding without punishing real customers
The goal is not to nuke every bot — it is to keep false positives near zero so you never block a paying customer. A graduated response works best:
- Allow verified search and shopping crawlers you want in feeds.
- Rate-limit borderline automated traffic so scraping becomes too slow to be useful.
- Serve stale or rounded data to suspected scrapers, poisoning the value of what they collect.
- Block outright high-confidence scrapers hitting pricing APIs from proxy networks.
Because Prynt runs server-side and returns a confidence score with reason codes, you can tune where each band starts and act on the evidence rather than a blunt allow/deny.
Measure before and after
Instrument your pricing pages so you know the baseline: what fraction of requests are automated, from which ASNs, targeting which SKUs. Retailers are routinely surprised to find that a double-digit percentage of “traffic” to their catalog is competitor bots. Once you can see it, you can decide how aggressively to respond, and you can prove the impact when scraping volume drops and margin on repriced SKUs recovers.
A note on legal versus technical defense
Many retailers ask whether they can simply sue a scraper. In practice, public price data occupies a legal gray area, cases are slow and jurisdiction-dependent, and a competitor can resume from a new entity before a ruling lands. Terms-of-service violations help, but they are hard to enforce against an anonymous proxy operation. Technical detection is the defense you actually control: it works in real time, applies equally to every offender regardless of who they are, and does not require you to identify a defendant. Treat legal options as a backstop for egregious, identifiable cases, and rely on detection and enforcement for the daily reality of automated repricing.
Getting started
You can profile your own scraping problem before spending a dollar. Prynt is free to start, so drop it in front of your catalog, watch the automated share of traffic in real time, and decide where to draw the line. When you are ready to enforce blocking across production endpoints, compare plans on our pricing page.
Try it free
Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.