The difference between a data swamp and a data lake? 5 signs

As companies collect increasing amounts of data and store it, they risk creating data swamps. Sometimes, what started as a data lake turns into a data swamp. Data lakes and data swamps are both data repositories, but data swamps are highly disorganised.
In short, a data lake equips companies to retrieve and use their data effectively. But, data swamps can make both those tasks exceptionally difficult and perhaps impossible. Here are five signs that what you think of as a data lake is actually a data swamp:
Metadata is information that describes other data. When appropriately used within a data lake, it acts as a tagging system that enables people to search for different kinds of data. Metadata can also create a tiered storage structure that stops a data lake from turning into a data swamp. Companies might organise their data with metadata tags denoting the source of the data or how it relates to a company event.
It’s also worthwhile to depend on metadata to help describe time frames or the age of the data. If an organisation made a metadata tag titled “2018 User Feedback Forms,” that metadata describes both the type and age of the information. Some metadata tags are less specific, such as “Twitter.” Even in that case, the people working with the data can use more than one metadata tag for a piece of information, thereby adding context to it.
Data swamps don’t have metatags. Then, the people accessing the data run into a problematic scenario where they may know exactly what kind of information they want to find but have no idea how to go about doing it.
Some company leaders get so excited about the fact that it’s now relatively easy to collect data that they start doing it without a clear goal in mind. A data lake can transform into a data swamp when companies don’t set parameters about the kinds of data they want to gather and why.
When enterprises can’t or won’t set limits on data amounts, they could find that what was once a well-organised data lake is now a data swamp flooded with information they may never need. Corporate silos can exacerbate the common problem of gathering data without rhyme or reason.
Perhaps departments have differing opinions about which kinds of data are most useful to a company at a given time. For example, the marketing department would likely want a different type of information than what’s most prized by the human resources department. Bringing relevance to data and ensuring it goes into a data lake instead of a swamp means getting everyone on the same page about when, why and how to acquire data.
Companies leaders should also adopt future-oriented mindsets data collection. But, when doing that, they must be careful not to fall into the trap of gathering data “just in case.” Making clearly defined goals about data usage helps prevent overeagerness when collecting the information.
Data governance defines how to treat data, who should handle it, where the data goes, how long companies retain the information and more.


