OLake Go
OLake Go is a high-performance, open-source data ingestion engine for replicating databases, S3, and Kafka into
Apache Iceberg (or plain Parquet).
Built for scalable, real-time pipelines, OLake Go provides a simple web UI and CLI - used to ingest into vendor-lock-in free Iceberg tables supporting all the query-engines/warehouses.
Read the docs and benchmarks at
olake.io/docs.
Join our active community on
Slack.
[!NOTE] 🎉 OLake Fusion is now live! — Automate your Apache Iceberg Table Maintenance. Check it out here → github.com/datazip-inc/olake-fusion 🎉
OLake Go — Super-fast Sync to Apache Iceberg
OLake Go supports replication from transactional databases such as PostgreSQL, MySQL, MongoDB, Oracle, DB2, and MSSQL, event-streaming systems like Apache Kafka and Object-store like S3, into open data lakehouse formats such as Apache Iceberg or Plain Parquet — delivering blazing-fast performance with minimal infrastructure cost.
🚀 Why OLake Go?
- 🧠 Smart sync: Full + CDC replication with automatic schema discovery & schema evolution
- ⚡ High throughput: 580K RPS (Postgres) & 338K RPS (MySQL)
- ➡️ Exactly once delivery & Arrow writes: Accuracy with speed.
- 💾 Iceberg-native: Supports Glue, Hive, JDBC, REST catalogs
- 🖥️ Self-serve UI: Deploy via Docker Compose and sync in minutes
- 💸 Infra-light: No Spark, no Flink, no Kafka, no Debezium
📊 Benchmarks
Full Load
| Source → Destination | Full Load | Relative Performance (Full Load) | Full Report |
|---|---|---|---|
| Postgres → Iceberg (as of 30th Jan 2026) |
5,80,113 RPS | 12.5× faster than Fivetran | Full Report |
| MySQL → Iceberg (as of 30th May 2026) |
1,39,773 RPS | 1.91× faster than Fivetran | Full Report |
| MongoDB → Iceberg (as of 5th Feb 2026) |
37,879 RPS | - | Full Report |
| Oracle → Iceberg (as of 30th Jan 2026) |
5,26,337 RPS | - | Full Report |
| Kafka → Iceberg (as of 27th Feb 2026) |
2,09,065 MPS (Bounded Incremental) | 1.23x slower than Flink | Full Report |
| MSSQL → Iceberg (as of 09th June 2026) |
3,45,866 MPS | 4.32x faster than Fivetran | Full Report |
CDC
| Source → Destination | CDC | Relative Performance (CDC) | Full Report |
|---|---|---|---|
| Postgres → Iceberg (as of 30th Jan 2026) |
55,555 RPS | 2× faster than Fivetran | Full Report |
| MySQL → Iceberg (as of 30th May 2026) |
59,951 RPS | 1.52× faster than Fivetran | Full Report |
| MongoDB → Iceberg (as of 5th Feb 2026) |
10,692 RPS | - | Full Report |
🔧 Supported Sources and Destinations
Sources (Databases)
| Source | Full Load | CDC | Incremental | Notes | Documentation |
|---|---|---|---|---|---|
| PostgreSQL | ✅ | ✅ pgoutput |
✅ | wal2json deprecated |
Postgres Docs |
| MySQL | ✅ | ✅ | ✅ | Binlog-based CDC | MySQL Docs |
| MongoDB | ✅ | ✅ | ✅ | Oplog-based CDC | MongoDB Docs |
| Oracle | ✅ | WIP | ✅ | JDBC based Full Load & Incremental | Oracle Docs |
| DB2 | ✅ | - | ✅ | JDBC based Full Load & Incremental | DB2 Docs |
| MSSQL | ✅ | ✅ | ✅ | Full Load, CDC & Incremental | MSSQL Docs |
Source (S3)
| Source | Full Load | CDC | Incremental | Notes | Documentation |
|---|---|---|---|---|---|
| S3 | ✅ | - | ✅ | Ingests from Amazon S3 or S3-compatible (MinIO, LocalStack) | S3 Docs |
Source (Kafka)
| Source | Bounded Incremental | Notes | Documentation |
|---|---|---|---|
| Kafka | ✅ | Latest offset bounded incremental sync | Kafka Docs |
Destinations
| Destination | Format | Supported Catalogs |
|---|---|---|
| Iceberg | ✅ | Glue, Hive, JDBC, REST (Nessie, Polaris, Unity, Lakekeeper, AWS S3 tables) |
| Parquet | ✅ | Filesystem |
Writer Docs
Apache Iceberg Docs
- Catalogs
- AWS Glue Catalog
- REST Catalog
- Generic
- Lakekeeper
- Nessie
- S3 Tables
- Unity
- Apache Polaris
- JDBC Catalog
- Hive Catalog
- Azure ADLS Gen2
- Google Cloud Storage (GCS)
- MinIO (local)
- Catalogs
Parquet Writer
- AWS S3 Docs
- [Google Cloud Storage (GCS)](https://olake.io/docs/writers/parquet/config/#using-gcs-co