Open source html-to-markdown projects

Every project in the registry tagged html-to-markdown, ranked by real GitHub adoption.

projects 4 combined stars ★ 190K refresh nightly
01 firecrawl ★ 182K

last push3 hours ago languageTypeScript licenseAGPL-3.0
02 trafilatura ★ 6.8K

Python & Command-line tool to gather text and metadata on the Web: Crawling, scraping, extraction, output as CSV, JSON, HTML, MD, TXT, XML

last push6 days ago languagePython licenseApache-2.0
03 crw ★ 1.0K

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scra

last push3 hours ago languageRust licenseAGPL-3.0
04 reader ★ 560

Open source web infrastructure for AI. Scrape, crawl, and automate the web, clean markdown, browser sessions, ready for your agents.

last push29 days ago languageTypeScript licenseApache-2.0

Related tags

← all tags

Frequently asked questions

How many open source html-to-markdown projects are there?

This registry tracks 4 projects tagged html-to-markdown, with 190,064 GitHub stars between them. The most-adopted is firecrawl at 181,627 stars.

Are these html-to-markdown projects free to use?

Yes — 4 of the 4 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which html-to-markdown project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these html-to-markdown projects still maintained?

4 of the 4 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.