7 Misconceptions About Data Lakes

4 min read
Curated from data-informed.com →

As big data becomes an increasingly integral part of business operations, companies are looking for ways to make data ingestion and analysis easier in order to deliver faster, more insightful results. The data lake has emerged as a popular big data tool thanks to its inherent ability to support the accumulation of data in the original format from a potentially infinite number of sources, such as social media, ticketing systems and sensors. Correlations across data sources are extremely valuable because of their potential to generate many more learnings than a single data source. Used wisely, the resulting insights empower professionals to make data-informed business decisions and build smarter automated processes.

Because the data lake is a relatively new concept, there is a significant amount of misinformation regarding what they are and how they work. Through my interactions with professionals across different industries and fields, and after going through dozens of articles on this topic, it became obvious that a few clarifications are in order.

Here are the most common misconceptions about the data lake:

The data lake can be more precisely defined as a paradigm, rather than a technology. It stems from the realization that keeping and processing data of various types and formats in one place, especially unstructured data from an entire organization, enables the emergence of insight that cannot be derived from single data sources. At its core, the data lake supports the big data endeavors of organizations by paving the way for the discovery of brand new and actionable insights.

2) The data lake is interchangeable with the data warehouse.

Concentrating data from different sources for correlation and modeling purposes is not a new idea. It has been around for years and has been substantiated by data warehouses. However, data warehouses and data lakes are very different concepts, due to their distinct types of focus on data structure, data usage, and data size.

It’s relatively easy to create complex reports if all the available data can fit into columns with records that match each other, as is the case with a data warehouse. The complexity arises from data that might not be in a format the user knows beforehand or where there is no structure to speak of. Examples of unstructured data sources are social media sources, web articles, invoices, and sensor data. Compared to the data warehouse, the data lake is uniquely suited to analyze both structured and unstructured data.

3) The data lake delivers insight on its own.

The data lake must be put to use by professionals with an understanding of the realities behind business processes. It also works best when used in conjunction with a set of processing, interrogation, transformation, and visualization tools. This means organizations need access to the right software & hardware toolkit, as well as the right pool of talent, to fully tap into the benefits of the data lake.

As the complexity of the questions/queries increase, the required tools and technologies change. This is why it is important to approach the operation of the data lake with agile processes, from both a technical and a business perspective.

4) It’s very hard to work with several multi-format data sources in the data lake.

One of the major concerns about the data lake is that even if you can store data in multiple formats, most apps will not be able to operate on all of them simultaneously. Most apps support only a handful, mostly structured data types. Less known is the fact that between the actual data and the apps there are middleware layers that can be inserted to provide “data virtualization.” Data virtualization integrates solutions that unify data formats, providing the apps with the required data type.

At the same time, many organizations that build data lakes also create so called data services. These can be virtual files that link to processes instead of actual files.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at data-informed.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.