Open source data-engineering projects

Every project in the registry tagged data-engineering, ranked by real GitHub adoption.

projects 19 combined stars ★ 251K refresh nightly
01 Apache Superset ★ 75K

Self-serve data exploration and dashboard builder

last push5 hours ago languagePython licenseApache-2.0
02 Kestra ★ 28K

Declarative workflow orchestration for data, AI, and infra

last push5 hours ago languageJava licenseApache-2.0
03 Prefect ★ 24K

Python-native workflow orchestration for data and ML teams

last push7 hours ago languagePython licenseApache-2.0
04 Airbyte ★ 22K

Open-source data integration for modern teams

last push5 hours ago languagePython license
05 taipy ★ 19K

Turns Data and AI algorithms into production-ready web applications in no time.

last push1 months ago languagePython licenseApache-2.0
06 dagster ★ 16K

An orchestration platform for the development, production, and observation of data assets.

last push6 hours ago languagePython licenseApache-2.0
07 great_expectations ★ 12K

Always know what to expect from your data.

last push6 hours ago languagePython licenseApache-2.0
08 CocoIndex ★ 12K

Ultra-fast data transformation for AI with lineage

last push34 hours ago languageRust licenseApache-2.0
09 Mage ★ 8.8K

Magical data pipeline tool for seamless transformations

last push6 days ago languagePython licenseApache-2.0
10 GrowthBook ★ 8.4K

Open-source A/B testing and feature flagging platform

last push5 hours ago languageTypeScript license
11 Evidence ★ 6.9K

Transform SQL into beautiful data stories

last push28 hours ago languageTypeScript licenseMIT
12 CloudQuery ★ 6.5K

Sync and transform data from any source to any destination

last push18 hours ago languageGo licenseMPL-2.0
13 aws-sdk-pandas ★ 4.1K

pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager,

last push2 days ago languagePython licenseApache-2.0
14 devlake ★ 3.1K

Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering

last push22 hours ago languageGo licenseApache-2.0
15 multiwoven ★ 1.7K

🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.

last push24 hours ago languageRuby licenseAGPL-3.0
16 duckle ★ 1.3K

Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,

last push12 hours ago languageRust licenseApache-2.0
17 Bacalhau ★ 871

Revolutionizing data processing with Compute Over Data

last push5 days ago languageGo licenseApache-2.0
18 conduit ★ 611

Conduit streams data between data stores. Kafka Connect replacement. No JVM required.

last push16 hours ago languageGo licenseApache-2.0
19 Orbital ★ 360

Connect APIs, databases, and streams without glue code

last push3 months ago languageTypeScript license

Related tags

← all tags

Frequently asked questions

How many open source data-engineering projects are there?

This registry tracks 19 projects tagged data-engineering, with 250,611 GitHub stars between them. The most-adopted is Apache Superset at 74,816 stars.

Are these data-engineering projects free to use?

Yes — 16 of the 19 carry an explicit open-source licence across 4 distinct licences, so there is no licence fee. Where a project also sells a hosted or enterprise version, the self-hosted path remains free.

Which data-engineering project should I choose?

The list above is ranked by GitHub stars, but stars measure attention rather than fit. Check three things on each card: the licence (permissive versus copyleft), the language it is written in, and the last-push date — a high-star project that has not been pushed in a year is a liability.

Are these data-engineering projects still maintained?

19 of the 19 were pushed in the last 90 days, and every card shows its exact last-push date so you can see the rest. Sort your shortlist by that date before committing to a migration.