Data Quality Is Also An AI Problem

3 min read
Curated from forbes.com →

Artificial intelligence (AI) continues its rise to prominence within the business world. The number of companies using AI today and the range of problems AI is being applied to are both increasing steadily. However, there is one issue that is plaguing AI just as much as it has plagued analytics of all kinds over the years—data quality.

Organizations put tremendous resources behind ensuring the quality of their data. This is necessary due to the broad range of ways that data quality can be compromised. Users might input data incorrectly, a system setting might lead to an incorrect code being assigned to certain actions, or a typo might end up in a script developed to facilitate data transformation. These are among the many potential sources of poor data quality.

The reality is that the data quality issue will never be “solved” no matter how much an organization budgets or how sincere its intentions may be. This is because the business environment—and the systems supporting it—is always in flux. New quality issues can arise at any time from any number of directions, including immediately after we certify that the data is pristine at a given point in time.

As organizations delve into AI, data quality will be as big of an issue as ever. This is because any AI processes that use traditional data sources will be just as dependent on those sources being of high quality as any other analytic processes. However, AI is also making use of a wide range of new data sources and data types. The methods of the past that are typically borrowed for a new data source fall apart when entering the realm of new data types used by AI.

Data such as images, text and videos have not been used to any significant degree in the past using non-AI methodologies. The data quality issues with these data types are also different than those of the past. Let’s take the example of images and consider a few ways that different data quality issues come into play.

• Data quality for model building. On the input side, images are often “tagged” to facilitate building a model. For example, a picture will be tagged as “containing a cat” or not, “containing a hot dog” or not, etc. Humans do this tagging, and humans can make errors. How do we find those errors and correct them? It can be easy to automate flagging that a price is clearly too low, an invoice is too high or an age can’t be true. With image tagging, it is very difficult to find an error without having a second person look for it and correct it. Detecting tagging errors mathematically is incredibly difficult, if not impossible, today.

• Data quality for model scoring.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at forbes.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.