Open source web-scraping projects
Every project in the registry tagged web-scraping, ranked by real GitHub adoption.
Scrapy, a fast high-level web crawling & scraping framework for Python.
Best and simplest tool for website change detection, web page monitoring, and website change alerts. Perfect for tracking content changes, price drops, restock
jsoup: the Java HTML parser, built for HTML editing, cleaning, scraping, and XSS safety.
High-performance browser automation bridge and multi-instance orchestrator with advanced stealth injection and real-time dashboard.
scrape data from Google Maps. Extracts data such as the name, address, phone number, website URL, rating, reviews number, latitude and longitude, reviews,emai
🤖 AI-powered web scraping editor with visual workflow builder. Build, test & deploy web scrapers using natural language. Powered by ScrapeGraphAI & LangGraph.
Undetected web-scraping & seamless HTML parsing in Python!
Related tags
Frequently asked questions
How many open source web-scraping projects are there?
This registry tracks 7 projects tagged web-scraping, with 127,508 GitHub stars between them. The most-adopted is scrapy at 64,388 stars.
Are these web-scraping projects free to use?
Yes — 7 of the 7 carry an explicit open-source licence across 3 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which web-scraping project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these web-scraping projects still maintained?
5 of the 7 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.