Open source webscraping projects
Every project in the registry tagged webscraping, ranked by real GitHub adoption.
Create agents that monitor and act on your behalf. Your agents are standing by!
Lightweight library for scraping web-sites with LLMs
Undetected web-scraping & seamless HTML parsing in Python!
Frequently asked questions
How many open source webscraping projects are there?
This registry tracks 3 projects tagged webscraping, with 51,886 GitHub stars between them. The most-adopted is huginn at 49,967 stars.
Are these webscraping projects free to use?
Yes — 3 of the 3 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which webscraping project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these webscraping projects still maintained?
1 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.