Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Business Intelligence 2021 • By Yves Mulkers

Apache Iceberg: The Hub of an Emerging Data Service Ecosystem?

Apache Iceberg: The Hub of an Emerging Data Service Ecosystem?
4 min read
Amazon S3, Apache Hadoop, Apache Hive
Curated from datanami.com →

Engineers at Netflix and Apple created Apache Iceberg several years ago to address the performance and usability challenges of using Apache Hive tables in large and demanding data lake environments. Now the data table format is the focus of a burgeoning ecosystem of data services that could automate time-consuming engineering tasks and unleash a new era of big data productivity.

Apache Hive emerged over a decade ago to make Apache Hadoop clusters look and function more like standard relational databases accessible through SQL. While Hadoop usage has waned in the age of cloud data lakes like AWS S3 and Azure Data Lake Storage (ADLS), the Hive legacy continues, both as a query engine for large, batch-oriented analytic jobs but arguably more so as a table format and a metadata catalog used by other query engines, including Apache Spark and Presto, among others.

In this manner, Hive acts as a unifying layer that enables these engines to function on a common set of data stored in Hadoop clusters and S3-compatible data lakes. While Hive marked a significant step forward in big data storage and analytics a decade ago, its technical limitations today are forcing data engineers and analysts to embark upon expensive and time-consuming workarounds to store and analyze massive data sets effectively.

One of the big problems with Hive is that it doesn’t adapt well to changing datasets. It can handle static data just fine, but if a user or an application makes changes to the data, such as an ETL job that updates a Parquet file, then those changes have to be coordinated with other applications or users. If this coordination does not happen, then the data is at risk of becoming corrupt and giving the wrong answer when queried.

This was one of the main drivers behind the Apache Iceberg project. Engineers at Apple and Netflix started the Iceberg project around 2018 to address the limitations in using Hive tables to store and query massive data sets. Ryan Blue, a senior engineer at Netflix and the PMC Chair of the Apache Iceberg project, recently discussed the genesis of Iceberg and the direction it’s headed in a session at the Subsurface 2021 conference, which was sponsored by Dremio and held last month.

“Iceberg exists because Netflix slowly realized we needed a new table format,” Blue said. “Many different services and engines were using Hive tables. But the problem was, we didn’t have that correctness guarantee. We didn’t have atomic transactions. And sometimes changes from one system caused another system to get the wrong data and that sort of issue caused us to just never use these services, not make change to our tables, just to be safe.”

The number one goal of the Iceberg project was to ensure correctness in the data, Blue said.

“Quite simply, tables shouldn’t lie to you when you query them,” he said. “It’s a really simple thing. But we survived for a very, very long period of time where these tables were being updated, or your file system was, say, S3 and didn’t provide a consistent listing, your tables could easily lie to you.”

Iceberg, which is written in Java and also offers a Scala API, effectively solves this dilemma by enforcing transactional consistency in the data, even when it’s accessed by multiple applications. According to Dremio’s description of Iceberg, the Iceberg table format “has similar capabilities and functionality as SQL tables in traditional databases but in a fully open and accessible manner such that multiple engines (Dremio, Spark, etc.) can operate on the same dataset.”

In addition to support for atomic transactions, the second major obstacle the Iceberg project tackled was enabling operations to be performed at a finer-grained level than simply partitioning the data level, Blue said.

“We needed to be able to rewrite data at the file level in order to do more efficient writes,” he said. “We wanted appends that could append to multiple partitions at the time.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • QlikCompany
  • LiveboardsTool
  • Alteryx DesignerTool
  • TelliusCompany
  • MicrostrategyCompany
  • ModeCompany
  • NarratorCompany
  • LightdashCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datanami.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

Artificial Intelligence: What Is Reinforcement Learning
Business Intelligence

Artificial Intelligence: What Is Reinforcement Learning

4 min read • 2018
How Open Data Is Changing Lives in Cities Around the World
Data Management

How Open Data Is Changing Lives in Cities Around the World

2 min read • 2016
The Implications of Emotional AI in the Legal World
Big Data

The Implications of Emotional AI in the Legal World

3 min read • 2022
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.