Open source data-pipeline projects
Every project in the registry tagged data-pipeline, ranked by real GitHub adoption.
Open-source data integration for modern teams
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
Flink CDC is a streaming data integration tool
Where data access meets operational intelligence
YAML-defined workflow orchestration, no database required
Data observability for modern data teams
A lightweight stream processing library for Go
CLI task management & automation tool
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestio
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,
zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Dat
SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for ever
Related tags
Frequently asked questions
How many open source data-pipeline projects are there?
This registry tracks 16 projects tagged data-pipeline, with 85,698 GitHub stars between them. The most-adopted is Airbyte at 22,082 stars.
Are these data-pipeline projects free to use?
Yes — 15 of the 16 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.
Which data-pipeline project should I choose?
The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.
Are these data-pipeline projects still maintained?
13 of the 16 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.