Open source pdf-parser projects

Every project in the registry tagged pdf-parser, ranked by real GitHub adoption.

projects 3 combined stars ★ 92K refresh nightly
01 PaddleOCR ★ 90K

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports

last pushyesterday languagePython licenseApache-2.0
02 pdfx ★ 1.1K

A free-floating 2D Canvas for processing multiple PDF files simultaneously

last push1 months ago languageTypeScript licenseMIT
03 pdf_oxide ★ 1.0K

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry

last push11 hours ago languageRust licenseApache-2.0

← all tags

Frequently asked questions

How many open source pdf-parser projects are there?

This registry tracks 3 projects tagged pdf-parser, with 91,825 GitHub stars between them. The most-adopted is PaddleOCR at 89,729 stars.

Are these pdf-parser projects free to use?

Yes — 3 of the 3 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which pdf-parser project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these pdf-parser projects still maintained?

3 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.