pydoll is a free, open source data extraction & web scraping project written in HTML and released under MIT. It has 7,088 GitHub stars, 406 forks and 23 open issues, and was last pushed 30 hours ago. On this registry it ranks #26 of 45 tracked projects in Data Extraction & Web Scraping, with 5 head-to-head comparisons available. It gained 2 stars over the last 3 tracked days.

What is pydoll?

Pydoll is a stealth-first Python library that automates Chromium-based browsers directly over the DevTools Protocol, without a WebDriver, for developers who scrape, test, or automate sites that block ordinary bots.

What it is

Pydoll is a browser automation library for Python that lives in the web scraping and end-to-end testing ecosystem. It drives Chrome directly over the Chrome DevTools Protocol (CDP) using a WebSocket connection, so there is no driver binary in the loop. It sits alongside tools such as Playwright and Puppeteer as a Chromium automation option, but it is written for Python from the ground up on top of asyncio, with type checking through mypy.

The concrete problem it solves is the wall that appears once automation leaves a developer machine. A scraper that runs fine locally tends to hit captchas and Cloudflare challenges when it runs for real, because standard automation leaks signals: the WebDriver binary, the navigator.webdriver flag, and inconsistent browser fingerprints. Pydoll removes the driver binary entirely and also tries to make the browser present a coherent identity, so ordinary bot protection has less to detect. Its own documentation is explicit that results on behavioral challenges depend on the browser and the IP reputation behind it, so it is a mitigation rather than a guarantee.

Key capabilities

  • Fingerprint injection through tab.apply_fingerprint(), aligning User-Agent, Client Hints, navigator, WebGL, canvas, screen, fonts, timezone, and locale. The overrides survive toString and prototype introspection and propagate into Web Workers, so lie-detection checks such as CreepJS's do not flag them.
  • Humanized interactions: mouse movement along Bezier curves, realistic typing, and scroll physics.
  • Zero WebDrivers — a direct CDP connection over WebSocket, with no driver binary and no navigator.webdriver flag.
  • An async, fully typed API built on asyncio and checked with mypy, giving IDE autocompletion and static error checking.
  • Network control for intercepting requests to block ads and trackers, monitoring traffic for API discovery, and making authenticated HTTP requests that inherit the browser session.
  • Full support for shadow roots, including closed ones, and for cross-origin iframes, queried and driven through the same API.
  • Structured extraction: define a Pydantic model, call tab.extract(), and receive typed, validated data instead of querying element by element.

Who uses it and how

  • Scraping teams whose pipelines break on captchas and Cloudflare challenges in production but work locally, and who need the browser to look consistent rather than merely headless.
  • QA and test engineers writing end-to-end tests against Chromium, using the async API inside existing Python test suites.
  • Crawler and anti-detection work where the browser must pass behavioral challenges such as Cloudflare Turnstile or reCAPTCHA v3, in combination with a reputable IP.
  • Projects that need to reach inside shadow DOM, including closed shadow roots, and cross-origin iframes without switching tooling.
  • Data extraction pipelines that want validated output from a Pydantic model rather than manual parsing of individual elements.

Getting started

The README points to the Documentation and Getting Started pages at https://pydoll.tech/, and the library requires Python 3.10 or newer. No Docker image or hosted option is described; it is installed and used as a Python library.

How it compares

Among the similar tools named in its own topics, Pydoll is grouped with Playwright and Puppeteer as Chromium automation, but it differs in mechanism: it uses a direct CDP connection over WebSocket rather than a WebDriver, and it ships fingerprint injection and humanized input as first-class features. Developers already committed to a JavaScript test runner have little reason to move, while Python teams that have been fighting detection have a more targeted option.

When to use it — and when not to

Choose Pydoll when the target is Chromium and the pain is detection, fingerprints, or driver version mismatches. Be aware that it is maintained by a single person who states plainly that releases and issue replies may be slower than usual, and that Firefox support is a stated goal tied to reaching 10k stars rather than a shipped feature. It is also the wrong pick for anyone who cannot operate their own browser environment and reputable IP, since passing behavioral challenges depends on both.

project readme (upstream, from github) — read inline

Pydoll Logo

The stealth-first browser automation library for Python.
No WebDriver, no navigator.webdriver flag, humanized clicks and typing.

Tests Ruff CI MyPy CI Python >= 3.10 Ask DeepWiki

Documentation · Getting Started · Features · Support

You have probably watched a scraper work on your machine, then hit a wall of captchas and Cloudflare challenges the moment it ran for real. That wall is what Pydoll is built around. It drives Chrome directly over the DevTools Protocol, so there is no WebDriver binary and no navigator.webdriver flag to give you away, and it clicks, types, and scrolls like a real person. That is often enough to get past the bot protection that stops ordinary automation, all behind an async, fully typed API.

Why Pydoll?

  • Fingerprint injection: Make the browser report a fully consistent identity with tab.apply_fingerprint(): User-Agent, Client Hints, navigator, WebGL, canvas, screen, fonts, timezone and locale, all aligned. The injected overrides survive toString and prototype introspection and propagate into Web Workers, so lie-detection checks like CreepJS's don't flag them.
  • Humanized interactions: Mouse movement along Bezier curves, realistic typing, and scroll physics. Often enough to pass behavioral challenges like Cloudflare Turnstile or reCAPTCHA v3, depending on your browser and IP reputation.
  • Zero WebDrivers: A direct CDP connection over WebSocket. No driver binary, no navigator.webdriver flag, no version-matching headaches.
  • Async and typed: Built on asyncio, type-checked with mypy. Full IDE autocompletion and static error checking.
  • Network control: Intercept requests to block ads/trackers, monitor traffic for API discovery, and make authenticated HTTP requests that inherit the browser session.
  • Shadow DOM and iframes: Full support for shadow roots (including closed) and cross-origin iframes. Discover, query, and interact with elements inside them using the same API.
  • Structured extraction: Define a Pydantic model, call tab.extract(), and get typed, validated data back. No manual element-by-element querying.

[!NOTE] A word from the maintainer. Pydoll is currently maintained by a single person, and I'm a bit stretched at the moment, so new releases and replies to issues may take a little longer than usual. To be clear: the project is not dead, and it is not going anywhere. Development continues; it's just moving at a calmer pace for now.

A goal to aim for: once the project reaches 10k stars, I plan to ship Firefox support, a big step that opens up a whole new range of possibilities for the library. Momentum like that is exactly the kind of incentive that makes a feature this large worth taking on, so if you'd like to see it happen, that's the push it needs.

Top Sponsors

SerpApi
Web Search API for your AI apps. Available in Markdown and JSON for any integration.
IPcook
Residential proxies for stealth browser automation: 55M+ IPs in 185+ locations, rotating & sticky sessions, city-level targeting, HTTP & SOCKS5, 99.99% uptime, sub-0.5s responses, pay-as-you-go traffic that never expires. Use WELCOME20 for 20% off.
The Web Scraping Club
The #1 newsletter dedicated to web scraping. Read their full, independent review of Pydoll.
NodeMaven
The most efficient proxy provider for web scraping and automation: ZIP targeting, 99.9% uptime, filtered high-quality IPs, no KYC. Use PYDOLL35 for 35% off Mobile & Residential, or PYDOLL40 for 40% off ISP (Static) proxies.
NiuProxy
Rotating residential proxies with a special deal for Pydoll users: 10TB at $0.35/GB or 1TB at $0.50/GB. Use PAY2 for 10% off your recharge.

Sponsors


PYDOLL 15% off

1GB free via our link

AI-native testing cloud

Proxies for automation
➕ Your logo here
Become a sponsor

Learn more about our sponsors · Become a sponsor

Installation

pip install pydoll-python

No WebDriver binaries or external dependencies required.

Getting Started

1. Stealthy Automation

The impera

readme truncated — read the full docs on github

Frequently asked questions

Is pydoll free to use?

pydoll is open source under the MIT licence. There is no licence fee and no seat count — you can self-host it or, where the project offers one, pay a vendor for a managed version instead.

What does pydoll do?

Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.

What is pydoll written in?

pydoll is primarily written in HTML. Its source is publicly available at https://github.com/autoscrape-labs/pydoll, and it has 7,088 GitHub stars.