category
Open Source Data Engineering & Integration Tools
Open source data engineering & integration tools
All Data Engineering & Integration tools
Open-source data integration for modern teams
Unify data models and metrics across your entire stack
Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.
Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.
Ultra-fast data transformation for AI with lineage
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Zero-ETL, infinite possibilities. Live query APIs, code & more with SQL. No DB required.
Broadcast, Presence, and Postgres Changes via WebSockets
Sync and transform data from any source to any destination
Flink CDC is a streaming data integration tool
Open-source BI for modern data teams
🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
Where data access meets operational intelligence
pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager,
A data integration framework
A system for agentic LLM-powered data processing and ETL
Scalable and efficient data transformation framework - backwards compatible with dbt.
Fast, Simple and a cost effective tool to replicate data from Postgres to Data Warehouses, Queues and Storage
Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering
Convert wearable data into actionable health insights with open algorithms
API for Current cases and more stuff about COVID-19 and Influenza
Data observability for modern data teams
A lightweight stream processing library for Go
CLI task management & automation tool
CherryUSB is a tiny and beautiful, high performance and portable USB host and device stack for embedded system with USB IP
🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.
Database Reporting Tool and Tasks (.Net)
A Python stream processing engine modeled after Yahoo! Pipes
The open document intelligence platform for builders and hackers - DMS for the agentic world
Hop Orchestration Platform
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestio
Actively maintained fork of Alibaba DataX — a fast, versatile ETL tool for RDBMS/NoSQL data transfer
Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality,
The Data Change Processing platform
zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Dat
SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).
Conduit streams data between data stores. Kafka Connect replacement. No JVM required.
Genblaze is an open source Python SDK for orchestrating generative AI media pipelines across video, audio, and image providers with built in provenance for ever
Seamless data import for your applications
Frequently asked questions
How many open source Data Engineering & Integration tools are there?
This registry tracks 39 open source Data Engineering & Integration projects, with 195,504 combined GitHub stars. The list is ranked by stars and refreshed nightly.
What is the most popular open source Data Engineering & Integration project?
Airbyte leads this category with 22,082 GitHub stars, followed by Cube.
Are these Data Engineering & Integration tools free?
Yes — every project listed here is open source. Some also offer paid hosted versions alongside the free self-hosted option; the licence for each project is shown on its card.