Open source data-pipeline projects

Every project in the registry tagged data-pipeline, ranked by real GitHub adoption.

projects 16 combined stars ★ 86K refresh nightly
01 Airbyte ★ 22K

Open-source data integration for modern teams

last push5 hours ago languagePython license
02 shardingsphere ★ 21K

Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.

last push10 hours ago languageJava licenseApache-2.0
03 debezium ★ 13K

Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.

last push7 hours ago languageJava licenseApache-2.0
04 flink-cdc ★ 6.5K

Flink CDC is a streaming data integration tool

last push17 hours ago languageJava licenseApache-2.0
05 whodb ★ 5.0K

Where data access meets operational intelligence

last push6 hours ago languageGo licenseApache-2.0
06 Dagu ★ 4.0K

YAML-defined workflow orchestration, no database required

last push18 hours ago languageGo licenseGPL-3.0
07 Elementary Data ★ 2.4K

Data observability for modern data teams

last push17 hours ago languageHTML licenseApache-2.0
08 go-streams ★ 2.2K

A lightweight stream processing library for Go

last push8 months ago languageGo licenseMIT
09 doit ★ 2.1K

CLI task management & automation tool

last push7 months ago languagePython licenseMIT
10 multiwoven ★ 1.7K

🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.

last push25 hours ago languageRuby licenseAGPL-3.0
11 olake ★ 1.5K

OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestio

last push12 hours ago languageGo licenseApache-2.0
12 duckle ★ 1.3K

Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,

last push12 hours ago languageRust licenseApache-2.0
13 zerocode ★ 1.0K

zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Dat

last push5 days ago languageJava licenseApache-2.0
14 seatunnel-web ★ 880

SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).

last push8 months ago languageJava licenseApache-2.0
15 conduit ★ 611

Conduit streams data between data stores. Kafka Connect replacement. No JVM required.

last push16 hours ago languageGo licenseApache-2.0
16 genblaze ★ 554

Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for ever

last push19 hours ago languagePython licenseMIT

Related tags

← all tags

Frequently asked questions

How many open source data-pipeline projects are there?

This registry tracks 16 projects tagged data-pipeline, with 85,698 GitHub stars between them. The most-adopted is Airbyte at 22,082 stars.

Are these data-pipeline projects free to use?

Yes — 15 of the 16 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which data-pipeline project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these data-pipeline projects still maintained?

13 of the 16 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.