Teradata CTO pours cold water on lakehouse concept

Despite industry efforts to get both data exploration and business analysis workloads onto a single “lakehouse” system, separate data lakes and warehouses are still required for effective enterprise analytics and BI systems, Teradata’s CTO tells The Register.
Speaking at the vendor’s London conference last week, Stephen Brobst sought to set his vision apart from recent trends espoused by rivals that increasingly see data management, analytics, BI, and machine learning bundled together on one platform.
“You need to have a unified architecture, but they are discrete things. There is a difference between the raw data, which is really data lake, and the data product, which is the enterprise data warehouse,” Brobst says.
Moves from rival players in the broad enterprise data and analytics markets appear to have set a different path. Born from the world of Hadoop and Apache Spark, Databricks has a long history in data lakes – where businesses dump structured and unstructured data for analytics and exploration – and has more recently added SQL support to its hybrid lakehouse system, where it encourages users to support both data exploration and regular business analytics workload.
Databricks, Snowflake, Cloudera, and Google are also betting on the two distinct workloads in the same environment with their approaches to data warehouses and data lakes respectively.
Although Teradata launched its own data lake in August, in part by improving optimization for object stores such as AWS S3, Brobst said there was an important distinction between where businesses put their raw data and the data warehouse, which optimizes query performance and controls governance.
He explains that although some Teradata customers use Databricks for their data lake, he advises against implementations where they persist the data, add key assignments, and do some light integration.
“This is actually not very useful, because you don’t want to have more copies of data than is necessary. If it adds value, OK, fine, but my view is that if you’ve done the hard work of adding the key structures and the homogenization of the data types, just put it in the … enterprise data warehouse,” he says.
Famous for his Hawaiian shirts and dramatic gesticulation during impassioned discussion of data warehousing architecture, Brobst graduated in computer science at UC Berkeley and gained a PhD at MIT. He helped found Teradata and has been one of the driving forces behind the growth of data warehousing in business for four decades.
He says the data lake is a “robust” concept, but distinct from the data warehouse, although they should be interoperable and within the same logical architecture.
“Vendors are all trying to claw it into their direction, but when done right the data lake is a good concept to land the raw data and have a low cost retention of that data in its original form,” Brobst says.
“We can afford to keep that raw data in the data lake and retain it and then the data scientists and the power users, they can use that raw data and explore it and decide which data surgically should be promoted in the data product and which one shouldn’t, so it’s not all or nothing anymore.


