OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.
github.com/datazip-inc/olakeIndexed at de2b96eUp to date with upstream
OLake Go is a codebase documentation engine for data-lake ingestion: it consumes source-DB change streams or backfills plus target lakehouse configuration, then runs a sync pipeline that produces Iceberg-native artifacts (or Plain Parquet) with exactly-once delivery semantics, surfaced through the OLake Go UI and optional CLI automation. It’s used by teams who want to replicate transactional systems—like PostgreSQL/MySQL/MongoDB/Oracle/DB2/MSSQL—into open lakehouse formats without building or operating Spark/Flink/Debezium-style infrastructure.
Written by gpt-5.4-nano from the index
374 files and 3,058 symbols in 10 modules, led by Go (237 files) and Java (33).
Start reading
Code health
This codebase scores 6.4 out of 10 for code health, which we rate fair. It also scores maintainability 6.9 and static performance 9.7 out of 10. The three are scored separately and never blended into one number. Risk is concentrated, as it usually is: 36 of 374 files are git hotspots, and they average 3.7 — which is where the fixes pay off most.
1 thing worth doing in the week to Sep 29, the last indexed commit, 1 of them now.
Clean up what this week's commits left: 7 serious findings in 6 files
1 commit in the last 7 days added or worsened them, and they are still open. Start with destination/iceberg/olake-iceberg-java-writer/src/main/java/io/olake/iceberg/rpc/OlakeRowsIngester.java, which has 1 critical.
Add a test coverage report
Without one, Repowise cannot tell tested code from untested code, so every test-related action says “unknown”.
Point Claude Code, Cursor, Codex or VS Code at this repository. Your agent gets the index this page is built from: cited answers, callers, history and health for any file.
https://api.repowise.dev/mcp/datazip-inc/olakeOne command, run anywhere. Adds the server to your local Claude Code config.
Docsclaude mcp add --transport http repowise https://api.repowise.dev/mcp/datazip-inc/olake \ --header "Authorization: Bearer rw_live_xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
This repo is public, so anyone with a Repowise account and an API key can connect.
Try asking
Prefer local? pip install repowise && repowise init indexes your own checkout, no account needed.
For maintainers
Both badges are public, cached, and update on their own after every index. Nothing to install.
Links to this page. Adding it also re-indexes the repo every week.
[](https://repowise.dev/repo/datazip-inc/olake)Average health across every file, straight from the latest index.
[](https://repowise.dev/repo/datazip-inc/olake)OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3. This page is a map of the datazip-inc/olake repository, written primarily in Go, rebuilt from the source each time it is indexed. Repowise parses every symbol, computes a dependency graph, scores per-file code health from complexity, duplication, test coverage and churn, mines git history for hotspots and ownership, and lifts the architectural decisions into documentation you can read here or query through MCP.
The codebase has 374 files and 3,058 symbols in 10 modules, led by Go, Java and Shell. Code health is 6.4 out of 10, rated fair.
Use the links above to open each view, or connect this repository to your agent for grounded answers inside Claude, Cursor, Codex or VS Code.