category

Open Source Data Extraction Web Scraping Tools

45 open source tools in this category.

projects 45 combined stars ★ 726K refresh nightly

All Data Extraction Web Scraping tools

Crawl4AI ★ 84K

LLM-ready web crawler built for AI data pipelines

rank001 licenseApache-2.0 written inPython typeData Extraction & Web Scraping
Scrapling ★ 82K

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDce

rank002 licenseBSD-3-Clause written inPython typeData Extraction & Web Scraping
scrapy ★ 64K

Scrapy, a fast high-level web crawling & scraping framework for Python.

rank003 licenseBSD-3-Clause written inPython typeData Extraction & Web Scraping
EasySpider ★ 45K

A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。

rank004 licenseAGPL-3.0 written inJavaScript typeData Extraction & Web Scraping
changedetection.io ★ 34K

Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock

rank005 licenseApache-2.0 written inPython typeData Extraction & Web Scraping
lux ★ 32K

👾 Fast and simple video download library and CLI tool written in Go

rank006 licenseMIT written inGo typeData Extraction & Web Scraping
CloakBrowser ★ 32K

Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.

rank007 licenseMIT written inPython typeData Extraction & Web Scraping
Scrapegraph-ai ★ 31K

Python scraper based on AI

rank008 licenseMIT written inPython typeData Extraction & Web Scraping
crawlee ★ 26K

Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or G

rank009 licenseApache-2.0 written inTypeScript typeData Extraction & Web Scraping
colly ★ 26K

Elegant Scraper and Crawler Framework for Golang

rank010 licenseApache-2.0 written inGo typeData Extraction & Web Scraping
stagehand ★ 24K

The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.

rank011 licenseMIT written inTypeScript typeData Extraction & Web Scraping
proxy_pool ★ 24K

Python ProxyPool for web spider

rank012 licenseMIT written inPython typeData Extraction & Web Scraping
Douyin_TikTok_Download_API ★ 20K

🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and pla

rank013 licenseApache-2.0 written inPython typeData Extraction & Web Scraping
katana ★ 18K

A next-generation crawling and spidering framework.

rank014 licenseMIT written inGo typeData Extraction & Web Scraping
Maxun ★ 17K

No-code web scraping, crawling, and extraction platform

rank015 licenseAGPL-3.0 written inTypeScript typeData Extraction & Web Scraping
SeleniumBase ★ 13K

APIs for browser automation, testing, and bypassing bot-detection. Includes CDP Mode: A stealthy configuration for chromium that passes every bot detection test

rank016 licenseMIT written inPython typeData Extraction & Web Scraping
crawlab ★ 12K

Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架

rank017 licenseBSD-3-Clause written inGo typeData Extraction & Web Scraping
webmagic ★ 12K

A scalable web crawler framework for Java.

rank018 licenseApache-2.0 written inJava typeData Extraction & Web Scraping
jsoup ★ 11K

jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.

rank019 licenseMIT written inJava typeData Extraction & Web Scraping
camofox-browser ★ 11K

Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.

rank020 licenseMIT written inJavaScript typeData Extraction & Web Scraping
pinchtab ★ 10K

High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.

rank021 licenseMIT written inGo typeData Extraction & Web Scraping
crawlee-python ★ 9.5K

Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, P

rank022 licenseApache-2.0 written inPython typeData Extraction & Web Scraping
firecrawl-mcp-server ★ 7.5K

🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.

rank023 licenseMIT written inTypeScript typeData Extraction & Web Scraping
JMComic-Crawler-Python ★ 7.3K

Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀

rank024 licenseMIT written inPython typeData Extraction & Web Scraping
rod ★ 7.1K

A Chrome DevTools Protocol driver for web automation and scraping.

rank025 licenseMIT written inGo typeData Extraction & Web Scraping
pydoll ★ 7.1K

Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.

rank026 licenseMIT written inHTML typeData Extraction & Web Scraping
trafilatura ★ 6.8K

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

rank027 licenseApache-2.0 written inPython typeData Extraction & Web Scraping
node-crawler ★ 6.8K

Web Crawler/Spider for NodeJS + server-side jQuery ;-)

rank028 licenseMIT written inTypeScript typeData Extraction & Web Scraping
curl_cffi ★ 6.5K

Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.

rank029 licenseMIT written inPython typeData Extraction & Web Scraping
ferret ★ 6.0K

Declarative data automation language and Go runtime for structured extraction workflows.

rank030 licenseApache-2.0 written inGo typeData Extraction & Web Scraping
skills ★ 5.9K

Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, indepen

rank031 licenseMIT written inPython typeData Extraction & Web Scraping
google-maps-scraper ★ 5.9K

scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,emai

rank032 licenseMIT written inGo typeData Extraction & Web Scraping
DotnetSpider ★ 4.1K

DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework

rank033 licenseMIT written inC# typeData Extraction & Web Scraping
puppeteer-sharp ★ 3.9K

Headless Chrome .NET API

rank034 licenseMIT written inC# typeData Extraction & Web Scraping
brightdata-mcp ★ 2.6K

A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

rank035 licenseMIT written inJavaScript typeData Extraction & Web Scraping
recipe-scrapers ★ 2.2K

Python package for scraping recipes data

rank036 licenseMIT written inPython typeData Extraction & Web Scraping
article-extractor ★ 1.9K

To extract article from given URL

rank037 licenseMIT written inTypeScript typeData Extraction & Web Scraping
OpenSERP ★ 1.4K

Self-hosted SERP API for Google, Bing, Yandex, and more

rank038 licenseMIT written inGo typeData Extraction & Web Scraping
parsera ★ 1.4K

Lightweight library for scraping web-sites with LLMs

rank039 licenseGPL-2.0 written inPython typeData Extraction & Web Scraping
socid-extractor ★ 1.1K

⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

rank040 licenseMIT written inPython typeData Extraction & Web Scraping
crw ★ 1.0K

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra

rank041 licenseAGPL-3.0 written inRust typeData Extraction & Web Scraping
pdf_oxide ★ 1.0K

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry

rank042 licenseApache-2.0 written inRust typeData Extraction & Web Scraping
scrapecraft ★ 698

🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.

rank043 licenseMIT written inPython typeData Extraction & Web Scraping
Stealth-Requests ★ 563

Undetected web-scraping & seamless HTML parsing in Python!

rank044 licenseMIT written inPython typeData Extraction & Web Scraping
reader ★ 560

Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

rank045 licenseApache-2.0 written inTypeScript typeData Extraction & Web Scraping

Frequently asked questions

How many open source Data Extraction Web Scraping tools are there?

This registry tracks 45 open source Data Extraction Web Scraping projects, with 726,351 combined GitHub stars. The list is ranked by stars and refreshed nightly.

What is the most popular open source Data Extraction Web Scraping project?

Crawl4AI leads this category with 83,747 GitHub stars, followed by Scrapling.

Are these Data Extraction Web Scraping tools free?

Yes — every project listed here is open source. Some also offer paid hosted versions alongside the free self-hosted option; the licence for each project is shown on its card.