Open source scraping projects
Every project in the registry tagged scraping, ranked by real GitHub adoption.
Scrapy, a fast high-level web crawling & scraping framework for Python.
Python scraper based on AI
Elegant Scraper and Crawler Framework for Golang
A scalable web crawler framework for Java.
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Lightweight library for scraping web-sites with LLMs
⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites
Related tags
Frequently asked questions
How many open source scraping projects are there?
This registry tracks 8 projects tagged scraping, with 319,360 GitHub stars between them. The most-adopted is firecrawl at 181,627 stars.
Are these scraping projects free to use?
Yes — 8 of the 8 carry an explicit open-source licence across 5 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which scraping project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these scraping projects still maintained?
6 of the 8 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.