Open source data-transformation projects
Every project in the registry tagged data-transformation, ranked by real GitHub adoption.
Collect, transform, and route logs and metrics in one tool
Build data pipelines with SQL and Python, ingest data from different sources, add quality checks, and build end-to-end flows.
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,
Frequently asked questions
How many open source data-transformation projects are there?
This registry tracks 3 projects tagged data-transformation, with 25,597 GitHub stars between them. The most-adopted is Vector at 22,581 stars.
Are these data-transformation projects free to use?
Yes — 3 of the 3 carry an explicit open-source licence across 2 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which data-transformation project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these data-transformation projects still maintained?
3 of the 3 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.