Open source data-collection projects
Every project in the registry tagged data-collection, ranked by real GitHub adoption.
A visual no-code/code-free web crawler/spider易采集:一个可视化浏览器自动化测试/数据采集/网页爬虫软件,可以无代码图形化的设计和执行爬虫任务。别名:ServiceWrapper面向Web应用的智能化服务封装系统。
Open-source data integration for modern teams
🚀 Self-hosted TikTok & Douyin scraper and no-watermark video downloader — async REST API, MCP server, CLI and web console for posts, profiles, comments and pla
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
Sync and transform data from any source to any destination
Open-source data integration for modern stacks
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Related tags
Frequently asked questions
How many open source data-collection projects are there?
This registry tracks 7 projects tagged data-collection, with 108,532 GitHub stars between them. The most-adopted is EasySpider at 44,558 stars.
Are these data-collection projects free to use?
Yes — 6 of the 7 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which data-collection project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these data-collection projects still maintained?
7 of the 7 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.