Could the data mesh solve your data lake scaling issues?

It’s good enough for Netflix and Exasol customer Zalando, but is data mesh architecture the right approach for your organization and its data democratization journey? Exasol CTO Mathias Golombek investigates: – What data mesh architecture is – Why you might implement a data mesh – The pros and cons of data mesh – The pioneers of data mesh architecture
In my recent blog series I delved into one of 2021’s hottest data topics – data democratization – exploring how it can fit into a business’ overarching data strategy along with some practical advice on n in your own organization.
For today’s follow up, I’m introducing another contemporary data concept – the data mesh. I’ll explore the link between data democratization and data mesh as a means to connect siloed data and create a self-service data infrastructure that makes data highly available and easily discoverable for the people who need it.
To be clear, I’m not advocating data mesh as a silver bullet to all the issues people experience with data lakes. It’s a concept that works for some, not everyone. Ultimately, you’ll need to make up your own mind.
The cloud is one of, if not the, most disruptive driver of radically new data-architecture approaches. But to fully understand what’s driving the need for data mesh, we need to appreciate the mess many organisations find themselves in when they try to scale their data.
Ananth Packkildurai’s article in Data Engineering Weekly contains a great analogy for the sad state of data infrastructure in many organizations. He likens the modern data generation process to the equivalent of writing a dictionary without any definitions, shuffling the words up randomly and then hiring expensive and analysts to try and make sense of it all. While this analogy certainly doesn’t apply to every organization it definitely resonates – and is at the core of why the data mesh principle has gained such a following over the last few years.
To write about data mesh and not acknowledge the ground-breaking work of its creator, ThoughtWorks consultant Zhamak Dehghani, would be unforgivable. Her papers: How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh and Data Mesh Principles and Logical Architecture have become required reading on the topic and I urge you to check them out, if you haven’t already.
To summarize, Dehgani’s data mesh theory argues that data platforms based on traditional data warehouse or data lake models have common failure modes that mean they don’t scale well. Instead of centralized lakes, or warehouses, data mesh advocates the shift to a more de-centralized and distributed architecture that fuels a self-serve data infrastructure and treats data more as a self-contained product.
Dehghani maintains that as your data lakes grow, so too does the complexity of the data management involved. In a traditional lake architecture you’ve typically got producers of data who generate it and send it into to the data lake.


