All articles Bot detection

Detecting nodriver and Other CDP-Light Automation

Every generation of browser automation is shaped by the detections that caught the one before. Selenium leaked navigator.webdriver and chromedriver’s injected variables. Puppeteer and Playwright in headless mode leaked a dozen inconsistencies between the headless shell and a real browser. undetected-chromedriver patched chromedriver to hide its fingerprints. nodriver, from the same author, takes the next step: it drops WebDriver and chromedriver entirely.

That makes it a good case study in what still works when the obvious artifacts are gone.

What nodriver changes

nodriver is an asynchronous Python library that launches an ordinary installed Chrome and controls it directly over the Chrome DevTools Protocol (CDP). There is no WebDriver session and no driver binary in between.

Compared with the tools that came before, that has consequences:

ArtifactSelenium / chromedriverPuppeteer / Playwright (headless)nodriver
navigator.webdriver === trueYes, unless patchedYes, unless patchedNot set by default
chromedriver-injected globalsYesNoNo
Headless rendering differencesDepends on modeOftenUsually runs headed
Real Chrome binary and profileVariesOften bundled ChromiumYes, the installed Chrome
Real Chrome TLS handshakeYesYesYes

The result is a browser whose JavaScript environment looks like a normal Chrome, because it is one. It also tends to be careful about which CDP domains it enables, since some of them create side effects a page can observe. Details change between releases, which is the point of the project: assume each known artifact will eventually be removed.

Artifacts that mostly disappear

Detections built on these signals degrade sharply against nodriver and similar tools:

  • WebDriver flags. No WebDriver, no flag.
  • Driver globals. No chromedriver, nothing injected.
  • Headless tells. Missing plugins, odd screen sizes, software rendering in headless mode, and user agents containing “HeadlessChrome”. A headed real Chrome has none of them.
  • TLS mismatches. A Python HTTP client claiming to be Chrome has a TLS handshake that does not match Chrome. nodriver traffic comes from Chrome itself, so the JA4 fingerprint is Chrome’s. The TLS_AUTOMATION reason code is excellent against HTTP-level bots and simply silent here.
  • Some CDP side effects. Techniques that detect an attached debugger through observable side effects, covered in the CDP artifact detection post, catch some configurations and miss others, depending on what the script enables.

None of this means artifacts are useless. Most automation is not this careful, and cheap checks still remove the bulk of it. The Selenium, Puppeteer and Playwright guide and the undetected-chromedriver post cover those. But a detection strategy that relies on artifacts alone has a ceiling, and nodriver sits above it.

Signals that survive

When the browser is real, look at what the automation does with it, where it runs, and how much of it there is.

Input behavior

A CDP script dispatches mouse and keyboard events through the protocol. The events are delivered by the browser itself, so a page sees them as trusted, but their shape is still the script’s. Common tells:

  • Pointer paths that are perfectly straight, or that jump directly to the target with no approach.
  • Clicks in the exact geometric center of elements, every time.
  • Keystrokes with uniform intervals, or entire fields filled in a single event.
  • No scrolling, hovering or idle movement before a form submit.
  • Focus moving between fields faster than a person reads labels.

Scripts can add jitter and curves, and good ones do. Synthetic variation still tends to be statistically different from human variation over a whole session, which is why behavioral analysis looks at distributions rather than individual events. Prynt’s behavioral Smart Signal and the AUTOMATION_BEHAVIOR reason code come from this layer, and form-level behavior feeds formBot and FORM_BOT.

Timing

Automation is impatient in characteristic ways: it submits as soon as the DOM is ready, waits for exactly the same duration on every run, or completes a multi-step flow faster than the pages can be read. Timing between page load, first interaction and submission is a cheap feature and hard to fake consistently across thousands of sessions.

Environment

nodriver runs on a machine somebody pays for, and at scale that machine is rarely a personal laptop on a home connection.

  • Virtual machines. Cloud and local VMs leave hardware traces in graphics and system properties. virtualMachine and VIRTUAL_MACHINE.
  • Datacenter and proxy networks. Traffic from cloud IPs, or routed through residential and mobile proxies to hide them. DATACENTER, PROXY, RESIDENTIAL_PROXY, VPN.
  • Inconsistency. A browser timezone that does not match the IP’s location, or a locale that does not match either, is weak alone and useful in combination.

Scale

The decisive signal is often economic. One person with nodriver is a minor problem. A farm of nodriver instances creating accounts, scraping or testing credentials produces volume, and volume concentrates on devices.

  • Velocity per visitorId: too many identifications in too short a window. VELOCITY.
  • Accounts per device: a browser profile reused across many signups shows up in accountsOnDevice, with MULTI_ACCOUNT and ACCOUNT_SHARING.
  • Profile churn: operators who create a fresh profile per account produce many new visitorIds that share an IP range, a VM fingerprint and identical behavior.

Putting it together

No single signal catches nodriver reliably; the combination does. A server-side check reads all of them from one event:

const event = await prynt.getEvent(requestId);   // @prynt/node, secret key
const s = event.smartSignals ?? {};

const automationLike = s.behavioral?.result || s.formBot?.result;
const hostedEnv      = s.virtualMachine?.result || s.datacenter?.result || s.residentialProxy?.result;
const scaled         = event.accountsOnDevice.count >= 2;

if (event.decision === 'block' || (automationLike && (hostedEnv || scaled))) {
  return deny();
}
if (automationLike || hostedEnv) {
  return requireChallenge();
}

For the middle ground, a proof-of-work challenge is a better step-up than a CAPTCHA. prynt.challenge() returns { passed, passToken }, and your server validates the single-use token. It costs a real user a moment of computation and costs a farm real CPU across every instance, which changes the economics without asking anyone to click on traffic lights.

Expect the arms race

nodriver will keep improving, and so will its successors. Detections that depend on one implementation detail will keep breaking. Detections that depend on the shape of automated work, how it moves, where it runs and how much of it there is, hold up because changing them costs the operator money.

The bot detection page summarizes the full signal set, and you can watch how your own browser scores on the playground. If you report confirmed bots back through POST /v1/outcomes, those labels mark the reported device, account or IP, so the next visit from any of them is flagged from the start.

Try it free

Prynt is device intelligence with a free tier — visitor IDs, bot & fraud Smart Signals, and behavioral biometrics, powered by a cross-site network. Start free.

Keep reading