Skip to content
7wData Data and AI tools, companies, events, podcast
  • Tools
  • Companies
  • Podcast
  • Articles
  • Events
  • Newsletter
  • Research
  • Sponsor

Table of Contents

Apache Hadoop 2018 • By Yves Mulkers

Data is inherently messy. Is that really such a bad thing?

Data is inherently messy. Is that really such a bad thing?
4 min read
Accuracy and precision, Data Quality, human rights
Curated from infoworld.com →

A data quality expert once told me that vendors providing data quality software solutions should always ensure 100 percent quality data, and if they didn’t, they should be liable for any ensuing issues. I disagreed with that harsh assessment then—and still do. The truth is, sometimes 100 percent data quality isn’t necessary and could even hinder an organization’s ultimate business goals.

As much as you would like our data to be perfect and pristine, to conform to your established dimensions of data quality, it isn’t. While there’s been renewed focus in recent years on the importance of data quality for achieving higher-value data and improving machine learning, data quality is not a new problem. Tools to address data quality have existed since at least the early 1990s, and MIT held its first International Conference on Information Quality back in 1996. 

After 20 to 25 years, you might expect that we would have mastered data quality! So why is 100 percent complete, clean, consistent, and accurate data still so difficult to achieve?  

The answer lies in changing your mindset: Data quality is contextual, not universal. It’s time for us to accept and expect that data is messy: incomplete, nonstandard, inconsistent, inaccurate, and out of date—but that’s not necessarily a bad thing. By understanding the contexts that make data messy, you can focus your efforts on addressing data quality issues where they are most critical, and to tolerate the rest where other factors are more important—in other words, put data quality in the right place at the right time. 

Not all data is created equal. We all have names—identifiers by which we are recognized. In seminars I’ve given, I’ve asked the question: “Is ‘John Doe’ good data?” Almost unanimously, the answer is no because it is considered fictitious and often used as test data. Yet “John Doe” is common and valid in health care or police investigations as the name for an unknown male (someone who does or did actually exist), in legal cases, as part of a Twitter handle for more than 100 people last I checked—not to mention there are real people with that name. The name John Doe is complete, consistent, and can be accurate. But you need to understand the context before you can say whether it is good, bad, or simply needs additional processing logic.

Numeric values and dates can be equally challenging. Just think about a rating scale from 1 to 5.  Is 1 the best rating, or is 5? Or a value of 100—is that a perfect grade, a high Fahrenheit temperature, an age, or an invalid credit rating? You need context (supplied via documentation, help, policies, metadata, etc.) to understand the data correctly, and to implement the right data quality checks and rules. You must then determine whether there is a data quality issue at all, and if so, whether it’s one around which you need data quality measurements and processes.

How you incorporate data into your operations and systems is another factor impacting your consideration of data quality. Building custom applications for every organizational function is expensive. Over time, you’ve replaced many of these with software packages and even suites of systems such as enterprise resource planning (ERP) products. Each of these products, as well as your homegrown applications, have systemic requirements. Enforcing a single, consistent organization-wide standard, whether for dates (annual calendar vs. timestamp vs. Julian date), Boolean values (T/F vs. Y/N vs. 1/0), or other codes, would be quixotic at best and otherwise resource- and revenue-consuming. The same is true for third-party data, including the increasing variety of open data available.

The definitions and semantics of data impact consistency of data as well. The definition of “customer,” for example, may vary depending on whether you are in marketing, order fulfillment, or finance.

In the 7wData directory

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Compare the tools & companies behind this topic

Browse the directory →
  • KKafkaCompany
  • LakebaseTool
  • ClickHouseTool
  • ClouderaCompany
  • Astra DBTool
  • Data TransferTool
  • DremioCompany
  • FFlinkCompany

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at infoworld.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.

Want the structural read on any AI or data company?
INS7GHTS

Want a sharper read on this topic?

Ask ins7ghts how the players compare, what people are actually shipping with, and where the trade-offs land.

Tweet LinkedIn Bluesky Threads Email

Related Articles

An Artificial Intelligence Definition for Beginners
Data Management

An Artificial Intelligence Definition for Beginners

4 min read • 2016
Transforming Data Insights: How AI Revolutionizes Data Management
Data Management

Transforming Data Insights: How AI Revolutionizes Data Management

12 min read • Jul 2024
Data revolutionized healthcare before
Data Management

Data revolutionized healthcare before, can it be done again?

3 min read • 2020
7wData

Independent reporting on AI and data: daily newsletter, podcast, deep dives.

Read

  • Ins7ghts newsletter
  • AI Beat newsletter
  • Latest articles
  • Podcast
  • Research guides

Use

  • Tools directory
  • Company directory
  • Research
  • Events
  • ins7ghts

Company

  • About
  • Contact
  • Sponsor a slot
  • Media kit
  • RSS feed

Follow

  • LinkedIn
  • X
  • YouTube
  • Instagram

© 2026 7wData. Independent. Belgium-based.

Privacy Cookies Terms Imprint Cookie settings
INS7GHTS
New · ins7ghts Drops

The AI governance conversation already moved. Most 2026 plans missed it.

Drop #1 · 60 pages · Launch week €99 (then €149) · ends Thu 9 July

Read Drop #1 →
Cookies on 7wData

We use strictly necessary cookies for the site to work, and optional analytics cookies to understand how readers use 7wData. We never share your data with advertisers. See our Cookie Policy.

Get the AI & data signal, daily. 335k+ already do.
Thanks. Check your inbox to confirm.
Get the AI & data signal

One curated email a day. 335k+ data & AI professionals already read it.

No spam. Unsubscribe anytime.

Check your inbox.

We just sent a confirmation. Click the link to start receiving the daily signal.