category
Open Source Data Extraction Web Scraping Tools
45 open source tools in this category.
All Data Extraction Web Scraping tools
LLM-ready web crawler built for AI data pipelines
🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDce
Scrapy, a fast high-level web crawling & scraping framework for Python.
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock
👾 Fast and simple video download library and CLI tool written in Go
Stealth Chromium that passes every bot detection test. Drop-in Playwright replacement with source-level fingerprint patches. 30/30 tests passed.
Python scraper based on AI
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or G
Elegant Scraper and Crawler Framework for Golang
The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.
Python ProxyPool for web spider
🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and pla
A next-generation crawling and spidering framework.
No-code web scraping, crawling, and extraction platform
APIs for browser automation, testing, and bypassing bot-detection. Includes CDP Mode: A stealthy configuration for chromium that passes every bot detection test
Distributed web crawler admin platform for spiders management regardless of languages and frameworks. 分布式爬虫管理平台,支持任何语言和框架
A scalable web crawler framework for Java.
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
Stealth headless browser for AI agents — bypass Cloudflare, bot detection, and anti-scraping. Drop-in Puppeteer/Playwright replacement.
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, P
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
Python API for JMComic | 提供Python API访问禁漫天堂,同时支持网页端和移动端 | 禁漫天堂GitHub Actions下载器🚀
A Chrome DevTools Protocol driver for web automation and scraping.
Pydoll is a library for automating chromium-based browsers without a WebDriver, offering realistic interactions.
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Web Crawler/Spider for NodeJS + server-side jQuery ;-)
Python binding for curl-impersonate fork via cffi. A http client that can impersonate browser tls/ja3/http2 fingerprints.
Declarative data automation language and Go runtime for structured extraction workflows.
Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, indepen
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,emai
DotnetSpider, a .NET standard web crawling library. It is lightweight, efficient and fast high-level web crawling & scraping framework
Headless Chrome .NET API
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Python package for scraping recipes data
To extract article from given URL
Self-hosted SERP API for Google, Bing, Yandex, and more
Lightweight library for scraping web-sites with LLMs
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra
The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry
🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.
Undetected web-scraping & seamless HTML parsing in Python!
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
Frequently asked questions
How many open source Data Extraction Web Scraping tools are there?
This registry tracks 45 open source Data Extraction Web Scraping projects, with 726,351 combined GitHub stars. The list is ranked by stars and refreshed nightly.
What is the most popular open source Data Extraction Web Scraping project?
Crawl4AI leads this category with 83,747 GitHub stars, followed by Scrapling.
Are these Data Extraction Web Scraping tools free?
Yes — every project listed here is open source. Some also offer paid hosted versions alongside the free self-hosted option; the licence for each project is shown on its card.