Data Observability Goes Far Beyond Data Quality Monitoring and Alerts

There’s no question that bad data hurts the bottom line.
Bad customer data costs companies six percent of their total sales, according to a UK Royal Mail survey. The UK Government’s Data Quality Hub estimates organizations spend between 10% and 30% of their revenue tackling data quality. For multi-billion-dollar companies, that can easily be hundreds of millions of dollars per year. Still another estimate by IBM pegs the overall cost of poor data quality for U.S. businesses at $3.1 trillion per year.
The cost of poor quality and unreliable data is set to skyrocket as companies transform into data-driven digital enterprises. Data will be required to underpin even more of their business models, internal processes, and will drive key business decisions. Access to data, therefore, is critical, but ensuring that that data is reliable becomes absolutely mission-critical.
Conventional wisdom says that data quality monitoring tools are the solution. As a result, the market is flush with these solutions. However, with the shift towards distributed, cloud-centric data infrastructures, data quality monitoring tools have rapidly become outdated. They were designed for an earlier generation of application environments, and are unable to scale, they are too labor-intensive to manage, too slow at diagnosing and fixing the root causes of data quality problems, and hopeless at preventing future ones.
It’s important to understand the technical reasons why data quality monitoring tools and their passive, alert-based approach have aged so poorly. And I will argue that rather than choosing a legacy technology, forward-looking organizations should look at a multidimensional data observability platform that is purpose-built for modern architectures to rapidly fix and prevent data quality problems, and automatically maintain high data reliability.
Data Quality: More than Just Data Errors
Let’s start with a refresher of what we mean by data quality. One of the biggest misconceptions about data quality is that it only means clean and error-free data. Acceldata Rohit Choudhary explained, there are actually six critical facets to data quality:
In other words, data can be error-free (accurate) yet have missing or redundant elements (completeness) that prevent you from completing analytical jobs. Or your data may be stored using different units or labels, creating (consistency) problems that wreak havoc on calculations. So can stale and out-of-date data (freshness). Or the schemas and structure of your data can differ wildly from dataset to dataset. This lack of normalization (validity) can make your data nearly impossible to aggregate and query together. You want to be able to identify uniqueness in your data sets because it is the most important dimension for establishing that there is no duplication. Measuring data uniqueness is done through analysis and comparison against other data records in your environment. Any data set with a high uniqueness score ensures you will have no — or only minimal — duplicates, and that builds trust in your data.
Here’s a real-world example of poor data quality wreaking havoc despite technically being error-free. In 1999, a NASA spaceship to Mars — the Mars Climate Orbiter — was lost due to a data consistency issue. While NASA used metric units, its contractor, Lockheed Martin, used imperial English units. As a result, Lockheed’s software calculated the Orbiter’s thrust in pounds of force, while NASA’s software ingested the numbers using the metric equivalent, newtons. This resulted in the NASA probe dipping 100 miles closer to the planet than expected, causing the Orbiter to either crash on Mars or fly out towards the sun (NASA doesn’t know, as communication with the probe was already lost). And it resulted in the premature end of NASA’s $327.6 million mission.
Too many data quality monitoring tools focus on trying to keep data error-free, when there are five other crucial aspects of data quality.


