Open source ocr projects
Every project in the registry tagged ocr, ranked by real GitHub adoption.
Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports
Effortlessly organize your digital documents
Powerful screen capture and file sharing tool for Windows
PyMuPDF is a high performance Python library for data extraction, analysis, conversion & manipulation of PDF (and other) documents.
Assist in organizing your piles of documents, resulting from scanners, e-mails and other sources with miminal effort.
Papermerge OSS
Capture web pages, text, and images. Find them fast.
Related tags
Frequently asked questions
How many open source ocr projects are there?
This registry tracks 7 projects tagged ocr, with 188,253 GitHub stars between them. The most-adopted is PaddleOCR at 89,729 stars.
Are these ocr projects free to use?
Yes — 7 of the 7 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which ocr project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these ocr projects still maintained?
7 of the 7 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.