Open source pdf-extraction projects

Every project in the registry tagged pdf-extraction, ranked by real GitHub adoption.

projects 3 combined stars ★ 17K refresh nightly
01 xberg ★ 9.3K

Polyglot document intelligence with a Rust core: extract text, metadata, images, tables, and structured data from 106 formats across 140 file extensions, plus c

last push7 hours ago languageRust licenseMIT
02 unstract ★ 7.2K

LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows

last push59 minutes ago languagePython licenseAGPL-3.0
03 signaturepdf ★ 831

Free open-source web software for signing PDF (alone or with others) and also organize pages, edit metadata and compress pdf

last push10 days ago languageJavaScript licenseAGPL-3.0

← all tags

Frequently asked questions

How many open source pdf-extraction projects are there?

This registry tracks 3 projects tagged pdf-extraction, with 17,396 GitHub stars between them. The most-adopted is xberg at 9,321 stars.

Are these pdf-extraction projects free to use?

Yes — 3 of the 3 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which pdf-extraction project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these pdf-extraction projects still maintained?

3 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.