6 dimensions of data quality boost data performance

4 min read
Curated from techtarget.com →

Artificial intelligence and machine learning can generate quality predictions and analysis, but first require organizations be trained on high quality data, starting with the six dimensions of data quality.

The old adage of computer programming — garbage in, garbage out — is just as applicable to today’s AI systems as it was to traditional software. Data quality means different things in different contexts, but, in general, good quality data is reliable, accurate and trustworthy.

“Data quality also refers to the business’ ability to use data for operational or management decision-making,” said Musaddiq Rehman, principal in the digital, data and analytics practice at Ernst & Young.

In the past, ensuring the quality of data meant a team of human beings would fact-check data records, but as the size and number of data sets increases, this becomes less and less practical and scalable.

Many companies are starting to use automated tools, including AI, to help with the problem.

By the end of this year, 60% of organizations will leverage machine-learning-enabled data quality technology in order to reduce the need for manual tasks, according to Gartner. To maximize these data quality tools, mastering the six dimensions of data quality should help ensure effective data performance.

Data accuracy is all about whether the data in the company systems matches that out in the real world or another verifiable source. “For an accuracy metric to provide valuable insights, there is typically the need for reference data to verify its accuracy,” said Rehman. For example, vendor data could be checked against a third-party data supplier database, or invoice amounts typed into a bookkeeping system could be checked against the paper documents. The biggest bane when it comes to accuracy is human data entry, whether it’s employees or customers themselves doing the typing. “One letter separates the postal abbreviations for Alabama, Alaska and Arkansas,” said Doug Henschen, vice president and principal analyst at Constellation Research. “One digit difference in an address or phone number makes the difference in being able to connect to a customer.” Even with all the recent progress in digitizing back-end systems and improving customer-facing interfaces, systems are still vulnerable to errors, he said. Good user interface design can help a lot here. For example, many customer-facing address forms have a built-in address checker to confirm that an address does, in fact, exist. Similarly, credit card numbers and email addresses can be checked at the time of the manual entry. The latest technology on this front is the customer data platform. “CDPs are primarily designed to resolve identities and tie together information associated with one person to create a single customer record,” Henschen said. But it can also help to ensure accuracy and keep records up to date as customers change jobs, get married and divorced, move, or get new email addresses. Most data quality tools offer functionality to validate addresses and perform other standard accuracy checks. They can also be used to profile data, so if someone enters something unexpected, it will send an alert. IBM, SAP, Attacama, Informatica and other leaders in Gartner’s Magic Quadrant for data quality solutions offer AI-powered data quality rule creation with a self-learning engine. Unfortunately, despite the new technologies coming on board, the accuracy problem is getting worse, not better, according to a survey of nearly 900 data experts released in September by data quality vendor Talend. For example, the percentage of respondents who said their data was up to date fell dramatically since this time last year. Only 28% rated their data “very good” in timeliness — down from 57% in 2021. The percentage of respondents who rated their data “very good” on accuracy fell as well, from 46% in 2021 to 39% this year.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at techtarget.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.