Data Lakes Prove Key to Modern Data Platforms

Big Data is quickly becoming a business’s best resource. Refined and analyzed data can help companies to improve operations and uncover insights. But as data begins to flow in from multiple sources, it’s important to prevent it from being siloed by pooling it in a central repository.
Enter the data lake: an architecture that acts as a storage space for data so engineers and IT teams can easily access it for future use.
Moreover, a data lake can form the firm foundation of a modern data platform. This was the case for telecom provider TalkTalk, which tapped Microsoft Azure’s Data Lake, Data Lake Analytics and Data Factory cloud-based hybrid data integration tools to increase data maturity.
Tapping Azure, the company quickly developed real-time and batch data pipelines while introducing security best practices and DevOps processes. This change brought data to the forefront of the company’s architectural decisions.
With this architecture in place, TalkTalk is looking to “scale the modern data platform, introducing real-time and event-driven ingestion and processing, as well as migrating on-premises systems to the cloud,” says Ben Dyer, TalkTalk’s head of data technology and architecture.
So, how do companies go about building and using a data lake? A good place to start is to understand the architecture and what it can deliver.
Data lakes store data of any type in its raw form, much as a real lake provides a habitat where all types of creatures can live together.
A data lake is an architecture for storing high-volume, high-velocity, high-variety, as-is data in a centralized repository for Big Data and real-time analytics. And the technology is an attention-getter: The global data lakes market is expected to grow at a rate of 28 percent between 2017 and 2023.
Companies can pull in vast amounts of data — structured, semistructured and unstructured — in real time into a data lake, from anywhere. Data can be ingested from Internet of Things sensors, clickstream activity on a website, log files, social media feeds, videos and online transaction processing (OLTP) systems, for instance. There are no constraints on where the data hails from, but it’s a good idea to use metadata tagging to add some level of organization to what’s ingested, so that relevant data can be surfaced for queries and analysis.
“To ensure that a lake doesn’t become a swamp, it’s very helpful to provide a catalog that makes data visible and accessible to the business, as well as to IT and data management professionals,” says Doug Henschen, vice president and principal analyst at Constellation Research.
Data lakes should not be confused with data warehouses. Where data lakes store raw data, warehouses store current and historical data in an organized fashion.
IT teams and data engineers should think of a data warehouse as a highly structured environment, where racks and containers are clearly labeled and similar items are stacked together for supply chain efficiency.
The difference between a data lake and a data warehouse primarily pertains to analytics.
Data warehouses are best for analyzing structured data quickly and with great accuracy and transparency for managerial or regulatory purposes. Meanwhile, data lakes are primed for experimentation, explains Kelle O’Neal, founder and CEO of management consulting firm First San Francisco Partners.


