Open source data-quality projects

Every project in the registry tagged data-quality, ranked by real GitHub adoption.

projects 3 combined stars ★ 15K refresh nightly
01 great_expectations ★ 12K

Always know what to expect from your data.

last push16 hours ago languagePython licenseApache-2.0
02 odd-platform ★ 1.4K

First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

last push4 days ago languageJava licenseApache-2.0
03 duckle ★ 1.3K

Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,

last push7 hours ago languageRust licenseApache-2.0

Related tags

← all tags

Frequently asked questions

How many open source data-quality projects are there?

This registry tracks 3 projects tagged data-quality, with 14,533 GitHub stars between them. The most-adopted is great_expectations at 11,798 stars.

Are these data-quality projects free to use?

Yes — 3 of the 3 carry an explicit open-source licence across 1 distinct licence, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which data-quality project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these data-quality projects still maintained?

3 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.