Open source web-crawler projects
Every project in the registry tagged web-crawler, ranked by real GitHub adoption.
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架
Open-source web scraping API. Turn any website into clean markdown or structured JSON. Anti-detect browser, proxy auto-selection, self-hosted. One command: make
Foundational low latency web data collecting in Rust
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
CLI tool for saving a faithful copy of a complete web page in a single HTML file (based on SingleFile)
Run a high-fidelity browser-based web archiving crawler in a single Docker container
Undetected web-scraping & seamless HTML parsing in Python!
Related tags
Frequently asked questions
How many open source web-crawler projects are there?
This registry tracks 7 projects tagged web-crawler, with 24,764 GitHub stars between them. The most-adopted is crawlab at 12,272 stars.
Are these web-crawler projects free to use?
Yes — 7 of the 7 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which web-crawler project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these web-crawler projects still maintained?
6 of the 7 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.