In Search of the Modern Data Stack

The modern data stack is many things to many people. It’s multi-cloud! It’s the data mesh! It’s BI plus AI! To get a better visibility into just what the modern data stack is, how it’s evolving, and why it all matters, we look to Fivetran’s “Multi-Cloud Modern Data Stack: Fireside Chat with Industry Trailblazers” for some insight.
For Ali Ghodsi, the CEO and co-founder of Databricks, things are pretty clear: The modern data stack is the open lakehouse architecture, which combines elements of a data warehouse and a data lakes to provide high-quality data in support of BI and AI workloads. It’s all about making things simpler.
“What’s going to happen over the next five years,” Ghodsi says during the fireside chat, “is “companies like Fivetran and Databricks and many others that are to come [are] going to re-envision how things are done in this amazing new infrastructure that we have. And it’s going to be much, much simpler. You’ll be able to actually do much more with it, and move faster.”
George Fraser, the CEO and co-founder of cloud-based ELT company Fivetran that hosted the chat, doesn’t necessarily share Ghodsi’s view of the lakehouse. For Fraser, many attempts at centralization have failed, which is he has said in the past that data lakes like S3 and ADLS are legacy tech.
“I think the core of the modern data stack is really the modern cloud data warehouse, and I would include Databricks in that category,” Fraser says (adding “you can yell at me later, Ali.”). “These modern analytical stores are so much faster. You can unplug your OLAP cube. You don’t need to do that anymore…There’s all these components that used to exist that just go away.”
Martin Casado, who has overseen a16z investments in both Databricks and Fivetran, isn’t quite so sure that the modern data stack will coalesce around data warehouses like those offered by Snowflake, Databricks, Google Cloud, AWS, and Microsoft Azure.
“You kind of presuppose that the data warehouse is central, and it’s clearly important,” he tells Fraser. “But in our investigations, when we talked a bunch of customers, it’s pretty clear to us that we’re seeing a multiplicity of data stores out there, and new emergent architectures. It may not be the data warehouse. It’s not obvious to me, certainly, that that’s going to be core.”
As the senior director of data analytics services at Google Cloud, Sudhir Hasbe gets his hands dirty in a lot of different products: BigQuery, Dataflow, Dataproc, Composer, Data Fusion, Data Catalog, Dataprep and PubSub. He unabashedly has a Google-centric view of what the modern data stack entails. “I think I would love to have all data on Google Cloud,” Hasbe quips. “It’s not going to happen.”
Actually, Google Cloud is the most progressive of the cloud giants when it comes to supporting a multi-cloud strategy. With its DataPlex offering, Google Cloud is also on the leading edge of adopting data fabric (or data mesh) approaches to federating management of data stored in different locations.
“Data is distributed in an organization across different clouds, and that’s here to stay for a very long time,” Hasbe says. “So the real question is, how can we enable organizations to leverage all data across all of these platforms, and provide capabilities that are going to be seamless?”
The logical view of the modern data stack gets more complicated when one considers two additional questions: Who is going to use it, and how is the data going to be managed? These may be afterthoughts for small teams. But in large enterprises with multiple departments that don’t necessarily see eye to eye on data (and which may actively be competing with each other), it becomes a tough question.
“When you have a single copy of data that can be accessed by different engines, the problem is people will create multiple copies,” Hasbe said.


