Do Content Managers Need a Data Lake?

The data lake concept arrived 10 years ago as the answer to common complaints about information silos and the volumes and complex varieties of information. It was a way to bring all the data together from multiple business applications and data systems into one centralized place, in whatever form it arrived, without the need for processing or structuring. The data lake was supposed to be the dream of fast-tracking structured data and unstructured data, such as videos and documents, into a one-stop repository shop for business insights, realized.
The data lake could hold vast amounts of raw data until needed for any use across the enterprise, even for use cases not yet identified. Unlike a data warehouse, which both stores and processes data, data lakes were designed to save time and provide flexibility, allowing consumers to just get the data in there and do what they need with the pure data down the line.
Ten years later it’s time to ask: are data lakes delivering on the promise? Data professionals are calling data lakes “dead” or “bad.” This article examines the current state of data lakes for the modern content manager and what to look for when deciding if data lakes are right for you or if a new concept like the data lakehouse is a better answer.
A data lake isn’t a technology product that you acquire; it’s an architecture or approach to storing and organizing data. You can have one data lake or several, depending on the organization. It can be in the cloud, on-premises or a hybrid.
Someone asked recently if they should get a data warehouse or a data lake. These two ideas are complementary, not an either-or choice. The data lake holds data in its pure form, while data warehouses store processed, cleansed and structured data.
Content or unstructured data usually comes into the data lake without a model to structure the content and without metadata to describe it. The data is not consistent, standardized or trustworthy. In order for search engines to be able to search the content and for analytics applications to extract insights from it, the architecture may need to supplement the data lake with services that simplify ingestion of content into the data lake and tag the content with metadata.
Related Article: Don’t Be Afraid of the Dark: Bring Dark Data Into the Light
Data lakes provide a fast data stream in, without requiring a model to profile just what it is and how important it is. Therefore data lakes can encourage a “keep everything” approach to managing content throughout its lifecycle. But this stops being so much of a lake and becomes more of a data dumping ground.


