How to Successfully Deploy an Enterprise Data Lake

How to Successfully Deploy an Enterprise Data Lake As enterprises try to extract more value from their data, the notion of building a “data lake” has gained traction. Data lakes are repositories of large amounts of structured and unstructured data that can be processed without the restrictions of a traditional data warehouse. By pooling a variety of data types, an enterprise ostensibly can uncover new insights to improve competitiveness. However, taking every scrap of data a company has and storing it all in one place is generally slow and inefficient. For actionable insights, you need to: a) have the right data; b) have the right sources; and c) process them in the right way. Simply installing Hadoop is not a strategy. This eWEEK slide show offers industry-information suggestions from Mark Gibbs, senior product manager at SnapLogic, on how to make a data lake work.
Most enterprises employ a broad set of software-as-a-service (SaaS) applications and devices that yield high volumes of data. Salesforce, Workday and Zendesk, for example, all hold key business insights. Streaming data from internet of things (IoT) sensors provides information about the health of equipment, and social streams reveal how customers perceive your product. But you need to combine this data with traditional, on-premises sources to get a full view of your business.
Knowing what you hope to achieve will help you decide which of these numerous data sources to ingest. That doesn’t mean you need to know in advance the insights you’ll get, but decide if your primary objective is to improve your product, refine your supply chain, be more efficient as a company or something else. Having broad goals in mind will allow you to know which data sources to prioritize.
Avoid “dumping,” so that your data lake doesn’t become a swamp of useless information. Enterprises are often tempted to throw everything into their data lake without a coherent strategy.


