How to unlock value from data lakes

3 min read
Curated from intelligentcio.com →

With data lakes crucial to organizations due to their ability to store large volumes of highly diverse data from multiple sources, Paul Leahy, Country Manager ANZ, Qlik, asks how a data lake can be effectively used to create value for your business.

There’s a common saying that data is the new water. Like water, data must be filtered before it is used, but unlike water, data is not limited by supply. That’s why companies have for a long time been moving to data lakes, where unfiltered, fast flowing data is stored in massive data dams for future analysis.

Put another way, data lakes are where the unfiltered water is, while analytics is the filtering and refining process, extracting value for the company.

Data lakes are important to organizations due to their ability to store large volumes of highly diverse data from multiple sources. It is this promise of cost-effective rapid storage that drove initial interest in data lakes; organizations wanted to overcome the costs and delays associated with storing data in traditional data warehouses.

So how can you effectively use a data lake to create value for your business? It starts with a cloud-based model.

The rise of the cloud-based data pond

Early on-premise installations of data lakes faced criticism for perceived challenges with security and governance performance issues, as well as the cost of maintaining and managing dedicated data centers.

Yet modern cloud-based data lakes have helped overcome many of those challenges. Big vendors including Amazon, Microsoft and Google offer managed cloud environments for data lakes replacing the capital cost of an on-prem Hadoop environment with an elastic consumption model, where organizations pay for what they use. They also mitigated some of the security and management challenges, allowing businesses to focus on data usage, rather than maintaining the environment.

The consumption model encouraged users to avoid dumping all their data into a central lake and instead load what they need for analytics, leading to smaller, purpose-built cloud data lakes or cloud-based data ponds. The rise of data science and Machine Learning platforms, as well as availability of SQL-based analytic services made accessing and analyzing stored data easier. This has meant data insights are made available faster to business users.

The challenge of making real-time analytics-ready data available to data consumers however remains. Traditional data integration approaches slow the data pipeline, making data outdated even before it is processed and ready for analysis. Then there are challenges around data trust and accessibility to data consumers.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at intelligentcio.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.