Open source html-to-markdown projects
Every project in the registry tagged html-to-markdown, ranked by real GitHub adoption.
Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML
Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra
Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.
Related tags
Frequently asked questions
How many open source html-to-markdown projects are there?
This registry tracks 4 projects tagged html-to-markdown, with 190,064 GitHub stars between them. The most-adopted is firecrawl at 181,627 stars.
Are these html-to-markdown projects free to use?
Yes — 4 of the 4 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which html-to-markdown project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these html-to-markdown projects still maintained?
4 of the 4 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.