Delta Lake

Delta Lake is an open-source storage layer that brings ACID transactions to Apache Spark and big data workloads.

Reviewed by 7wData

On this page

Profile

Delta Lake is an open-source storage layer that adds ACID transactions, schema enforcement, and time travel to data lakes, enabling reliable batch and streaming data processing on cloud object storage.

Delta Lake is an open-source storage layer that brings ACID transactions to Apache Spark and big data workloads. It was originally developed at Databricks in 2019 and later donated to the Linux Foundation. The project is governed by the Delta Lake community and is hosted on GitHub under the delta-io organization, where it has accumulated over 8,800 stars and 2,100 forks as of June 2026.

Delta Lake provides a transactional metadata layer on top of existing data lakes (e.g., S3, ADLS, HDFS), enabling features like schema enforcement, time travel, and unified batch/streaming. The project's core is written in Scala and Java, with connectors for Spark, Flink, and a standalone kernel. In June 2024, Databricks acquired Tabular, a company founded by the original creators of Apache Iceberg, for over $1 billion (per CNBC), signaling a push to unify the Delta Lake and Iceberg formats.

The acquisition brought Tabular's co-founders—Ryan Blue, Daniel Weeks, and Jason Reid—into Databricks. As of early 2026, Delta Lake remains the most widely deployed open-source table format for data lakes, though it faces increasing competition from Apache Iceberg and Apache Hudi. The project's governance and roadmap are heavily influenced by Databricks, which employs the majority of core committers.

No independent financials for Delta Lake are available, as it is not a standalone entity; it is a project within the Databricks ecosystem. Databricks itself reported over $1.6 billion in revenue in its fiscal year ending January 2025 (per The Information), but Delta Lake's specific adoption metrics are not disclosed. Recent activity on the GitHub repository shows ongoing development for Spark 4.2 compatibility and changelog V2 features.

Track Delta Lake and 240+ vendors.

335k+ subscribers read the daily AI & data note. One email, both newsletters. Unsubscribe anytime.

Who buys this

  • Data engineering teams building ETL pipelines on Apache Spark
  • Enterprises migrating from on-premise data warehouses to cloud data lakes
  • Organizations running machine learning workflows that require consistent training data snapshots
  • Companies using Databricks as their primary data and AI platform
  • Teams operating streaming data pipelines that need exactly-once semantics

Strengths and what to watch

Strengths

  • Open-source project with strong community adoption: 8,800+ GitHub stars and 2,100+ forks as of June 2026.
  • Backed by Databricks, which provides sustained engineering investment and integration with its commercial platform.
  • ACID transactions on cloud object storage (S3, ADLS, GCS) without requiring a separate database cluster.

Watch for

  • Governance is effectively controlled by Databricks, which employs most core committers; community concerns about vendor influence persist.
  • Competition from Apache Iceberg (backed by AWS, Netflix, and others) and Apache Hudi (backed by Uber) is intensifying, with Iceberg gaining rapid adoption in 2025-2026.
  • The Tabular acquisition (2024) created internal tension over format unification; the promised Delta-Iceberg integration has not yet produced a unified standard.

Key Information

Industry
Formats
Founded
2019

Frequently Asked Questions

What is Delta Lake and what does it do?

Delta Lake is an open-source storage layer that adds ACID transactions, schema enforcement, and time travel to data lakes. It enables reliable batch and streaming data processing on cloud object storage like S3, ADLS, and HDFS.

Who created Delta Lake and who maintains it now?

Delta Lake was originally developed at Databricks in 2019 and later donated to the Linux Foundation. The project is governed by the Delta Lake community and hosted on GitHub under the delta-io organization, but Databricks employs most core committers.

How does Delta Lake compare to Apache Iceberg and Apache Hudi?

Delta Lake faces increasing competition from Apache Iceberg, backed by AWS and Netflix, and Apache Hudi, backed by Uber. Iceberg has gained rapid adoption in 2025-2026. Databricks acquired Tabular in 2024 to unify Delta Lake and Iceberg formats.

What are the main benefits of using Delta Lake for data lakes?

Delta Lake provides ACID transactions on cloud object storage without needing a separate database cluster. It offers schema enforcement, time travel for data versioning, and unified batch and streaming processing, making data lakes more reliable and consistent.

Is Delta Lake controlled by Databricks and what does that mean for users?

Yes, Databricks effectively controls Delta Lake's governance and roadmap, employing most core committers. This vendor influence raises community concerns, but it also ensures sustained engineering investment and tight integration with Databricks' commercial platform.

Can Delta Lake handle streaming data pipelines with exactly-once semantics?

Yes, Delta Lake supports streaming data pipelines that require exactly-once semantics. Its ACID transactions and unified batch/streaming processing enable reliable data ingestion and processing, making it suitable for real-time analytics and machine learning workflows.

Sources

  1. github.com — GitHub repository statistics (stars, forks, commits) and recent commit history showing Spark 4.2 compatibility work.
  2. techcrunch.com — Databricks acquisition of Tabular for over $1 billion, bringing Iceberg creators into Databricks.
  3. ir.delta.com — Delta Air Lines Q1 2026 financial results (irrelevant to Delta Lake; included to show the search returned airline data, not Delta Lake).
  4. www.prnewswire.com — Duplicate of Delta Air Lines earnings (irrelevant to Delta Lake).