Open source data-extraction projects

Every project in the registry tagged data-extraction, ranked by real GitHub adoption.

projects 18 combined stars ★ 363K refresh nightly
01 firecrawl ★ 182K

last push3 hours ago languageTypeScript licenseAGPL-3.0
02 Scrapling ★ 82K

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl! Don't be shy, join here: https://discord.gg/EMgGbDce

last push3 days ago languagePython licenseBSD-3-Clause
03 Scrapegraph-ai ★ 31K

Python scraper based on AI

last push10 days ago languagePython licenseMIT
04 stagehand ★ 24K

The SDK to extract data and interact with any site on the web. Get started with Claude Code, Codex, Eve, Mastra, and more.

last push4 hours ago languageTypeScript licenseMIT
05 Maxun ★ 17K

No-code web scraping, crawling, and extraction platform

last push11 hours ago languageTypeScript licenseAGPL-3.0
06 ferret ★ 6.0K

Declarative data automation language and Go runtime for structured extraction workflows.

last push26 hours ago languageGo licenseApache-2.0
07 skills ★ 5.9K

Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, indepen

last push24 days ago languagePython licenseMIT
08 brightdata-mcp ★ 2.6K

A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

last push8 hours ago languageJavaScript licenseMIT
09 recipe-scrapers ★ 2.2K

Python package for scraping recipes data

last push8 days ago languagePython licenseMIT
10 contextgem ★ 2.0K

ContextGem: Effortless LLM extraction from documents

last push1 months ago languagePython licenseApache-2.0
11 article-extractor ★ 1.9K

To extract article from given URL

last push29 days ago languageTypeScript licenseMIT
12 parsera ★ 1.4K

Lightweight library for scraping web-sites with LLMs

last push9 months ago languagePython licenseGPL-2.0
13 socid-extractor ★ 1.1K

⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

last push5 days ago languagePython licenseMIT
14 crw ★ 1.0K

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra

last push3 hours ago languageRust licenseAGPL-3.0
15 pdf_oxide ★ 1.0K

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry

last push11 hours ago languageRust licenseApache-2.0
16 scrapecraft ★ 698

🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.

last push9 months ago languagePython licenseMIT
17 Stealth-Requests ★ 563

Undetected web-scraping & seamless HTML parsing in Python!

last push6 months ago languagePython licenseMIT
18 reader ★ 560

Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

last push29 days ago languageTypeScript licenseApache-2.0

Related tags

← all tags

Frequently asked questions

How many open source data-extraction projects are there?

This registry tracks 18 projects tagged data-extraction, with 363,380 GitHub stars between them. The most-adopted is firecrawl at 181,627 stars.

Are these data-extraction projects free to use?

Yes — 18 of the 18 carry an explicit open-source licence across 5 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which data-extraction project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these data-extraction projects still maintained?

15 of the 18 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.