Data Hygiene The real cost of dirty data and 5 tips to improve data quality

Growing up, we’re taught to keep our hands clean. As children we’re told ‘wash your hands before dinner’ and ‘wash your hands after playing outside.’ Even as adults, public health programs show us how to wash our hands effectively to prevent disease.
Now, as primary data custodians for our respective organizations, we still need to keep it clean.
This time we’re not talking about your hands, but your data.
The term dirty data is not just a catchy alliteration. It is, in fact, an extremely serious problem.
How serious, you might ask.
According to Experian’s 2019 Global Data Management Research: “We see, year after year, that despite our ambitions, many businesses fail to take full advantage of the opportunity that data can provide to improve customer interactions to increase business performance.”
In the US alone, an IBM survey revealed that bad data costs the economy $3.1 trillion every year. Moreover, more than 30% of business leaders are not confident with the data they’re using to make key business decisions, while 27% of respondents are uncertain.
Dirty data typically refers to data that is poorly structured, has inaccuracies or is incomplete. It impacts various industries differently. However, whichever industry an enterprise is operating in, the negative repercussions are equally damaging to a business’ overall health.
In the financial services industry, dirty data goes beyond financial loss. Inaccurate and incomplete data can lead to regulatory breaches, delayed decisions due to manual checks, and sub-optimal trade strategies just to name a few.
Businesses that use and rely on a CRM for lead nurturing and customer segmentation are likewise negatively affected. Culled statistics show that while 67% of businesses use CRM data for customer targeting, 60% believe that their overall data health is unreliable.
Even the healthcare industry is not spared. According to Healthcare Finance, supplies management is one area that dirty data affects the most. Supply costs account for 20% – 30% of operating expenses, yet this area is often mismanaged due to incorrect or incomplete data. Dirty data also affects inventory. Making sure that medical supplies are present where they are needed which could spell the difference between life and death.
How can dirty data inflict so much damage? Data Doctor Thomas Redman explains, “The reason dirty data costs so much is that decision makers, managers, knowledge workers, data scientists, and others must accommodate it in their everyday work. And doing so is both time-consuming and expensive. The data they need has plenty of errors, and in the face of a critical deadline, many individuals simply make corrections themselves to complete the task at hand. They don’t think to reach out to the data creator, explain their requirements, and help eliminate root causes.”
Dirty data can affect all business type and industries, and it can wreak havoc even in today’s most advanced digital projects.
Initiatives involving artificial intelligence and modern data governance are the prime examples of how dirty data impacts digital transformation projects.
According to a study by market research firm Dimensional Research, 8 out of 10 AI and machine learning projects have stalled due to poor data quality, while 96%, “have run into problems with data quality, data labeling required to train AI, and building model confidence.”
This is bad news as AI, big data, and machine learning now rank highly in the top priorities of many companies’ digital transformation initiatives.
The utility sector is one of the industries most affected by this AI, machine learning, and data governance stall caused by dirty data.


