Open source content-extraction projects

Every project in the registry tagged content-extraction, ranked by real GitHub adoption.

projects 3 combined stars ★ 10K refresh nightly
01 firecrawl-mcp-server ★ 7.5K

🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.

last push13 hours ago languageTypeScript licenseMIT
02 article-extractor ★ 1.9K

To extract article from given URL

last push29 days ago languageTypeScript licenseMIT
03 x-reader ★ 962

Universal content reader MCP Server for 10+ platforms

last push19 days ago languagePython licenseMIT

← all tags

Frequently asked questions

How many open source content-extraction projects are there?

This registry tracks 3 projects tagged content-extraction, with 10,353 GitHub stars between them. The most-adopted is firecrawl-mcp-server at 7,478 stars.

Are these content-extraction projects free to use?

Yes — 3 of the 3 carry an explicit open-source licence across 1 distinct licence, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which content-extraction project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these content-extraction projects still maintained?

3 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.