How to Clean Your Data Swamp Without Draining It

Organizations want data lakes. But too often, they end up instead with data swamps that are of little use for transforming data into value. If you’re struggling to turn your data swamp into a clear data lake, keep reading for data lake organization and storage best practices.
In case it’s not clear, the term data swamp is something of a play on words. It’s not a term that professional data scientists frequently use.
Data scientists do, however, talk about data lakes. A data lake is a body of data that you transform, modify or analyze to gain valuable insights.
Ideally, your data lake will be clear, smooth and actionable. But if it’s not, your data lake looks more like a swamp – a murky, burdensome, difficult-to-maintain body of data that is difficult to turn into value.
Data swamps arise when you face challenges like the following:
By definition, data lakes can contain multiple types and structures of data. But that doesn’t mean you can simply throw data into a data lake in a willy-nilly fashion. You need a data governance framework to ensure that your data does not become so complex and unruly that it is difficult to analyze.
To transform your data into value, you need to be able to convert data to formats that your analytics tools support. Data conversion is the process that connects your data lake to your data analytics operation.
This can be challenging to do when conversion tools don’t support the types of data in your data lake. It’s a common problem when your data lake includes data from legacy infrastructure, such as mainframes, and you attempt to analyze it using modern tools.
Getting data into a data lake quickly can also be a problem, especially when dealing with legacy environments and tools. To go back to the mainframe example, your legacy mainframe infrastructure may lack the ability (without the help of third-party tools) to offload data in real time into a data lake that lives in the cloud.
Even the best governed, most transformable data is of little use if crucial information is missing. For example, a data lake composed of machine data that is missing information from certain devices won’t give you a reliably clear idea of what is in your infrastructure.


