Open source crawler projects

Every project in the registry tagged crawler, ranked by real GitHub adoption.

projects 25 combined stars ★ 639K refresh nightly
01 firecrawl ★ 182K

last push5 hours ago languageTypeScript licenseAGPL-3.0
02 Scrapling ★ 82K

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDce

last push3 days ago languagePython licenseBSD-3-Clause
03 scrapy ★ 64K

Scrapy, a fast high-level web crawling & scraping framework for Python.

last push10 hours ago languagePython licenseBSD-3-Clause
04 EasySpider ★ 45K

A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。

last push13 hours ago languageJavaScript licenseAGPL-3.0
05 lux ★ 32K

👾 Fast and simple video download library and CLI tool written in Go

last push6 months ago languageGo licenseMIT
06 Scrapegraph-ai ★ 31K

Python scraper based on AI

last push10 days ago languagePython licenseMIT
07 crawlee ★ 26K

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or G

last push8 hours ago languageTypeScript licenseApache-2.0
08 colly ★ 26K

Elegant Scraper and Crawler Framework for Golang

last push38 hours ago languageGo licenseApache-2.0
09 proxy_pool ★ 24K

Python ProxyPool for web spider

last push3 months ago languagePython licenseMIT
10 Douyin_TikTok_Download_API ★ 20K

🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and pla

last push3 days ago languagePython licenseApache-2.0
11 katana ★ 18K

A next-generation crawling and spidering framework.

last push4 days ago languageGo licenseMIT
12 Maxun ★ 17K

No-code web scraping, crawling, and extraction platform

last push13 hours ago languageTypeScript licenseAGPL-3.0
13 crawlab ★ 12K

Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架

last push7 months ago languageGo licenseBSD-3-Clause
14 webmagic ★ 12K

A scalable web crawler framework for Java.

last push9 months ago languageJava licenseApache-2.0
15 crawlee-python ★ 9.5K

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, P

last push35 hours ago languagePython licenseApache-2.0
16 JMComic-Crawler-Python ★ 7.3K

Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀

last push4 days ago languagePython licenseMIT
17 pydoll ★ 7.1K

Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.

last push31 hours ago languageHTML licenseMIT
18 trafilatura ★ 6.8K

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

last push6 days ago languagePython licenseApache-2.0
19 node-crawler ★ 6.8K

Web Crawler/Spider for NodeJS + server-side jQuery ;-)

last push3 months ago languageTypeScript licenseMIT
20 DotnetSpider ★ 4.1K

DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework

last push6 months ago languageC# licenseMIT
21 puppeteer-sharp ★ 3.9K

Headless Chrome .NET API

last push10 hours ago languageC# licenseMIT
22 SCrawler ★ 2.2K

🏳️‍🌈 Media downloader from any sites, including Twitter, Reddit, Instagram, BlueSky, TikTok, Threads, Facebook, OnlyFans, YouTube, Pinterest, PornHub, XHamste

last push1 months ago languageVisual Basic .NET licenseGPL-3.0
23 crw ★ 1.0K

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra

last push5 hours ago languageRust licenseAGPL-3.0
24 seonaut ★ 786

Open source SEO audit tool.

last push4 months ago languageGo licenseMIT
25 reader ★ 560

Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

last push29 days ago languageTypeScript licenseApache-2.0

Related tags

← all tags

Frequently asked questions

How many open source crawler projects are there?

This registry tracks 25 projects tagged crawler, with 639,419 GitHub stars between them. The most-adopted is firecrawl at 181,627 stars.

Are these crawler projects free to use?

Yes — 25 of the 25 carry an explicit open-source licence across 5 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which crawler project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these crawler projects still maintained?

18 of the 25 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.