Data Quality Impacts on Downstream Analytics and Machine Learning

Advanced analytics, including artificial intelligence and machine learning (AI/ML) are hot areas for investment right now, and for good reason. Organizations have access to more data than ever before, and to the extent that business leaders can draw useful insights from that information, they stand to gain significant new efficiencies and achieve a meaningful advantage over their competitors.
Unfortunately, there are some risks associated with diving head-first into advanced analytics and AI/ML without first attending to data integrity. Data integrity encompasses data quality, which ensures that information is accurate, complete, and consistent. But data quality is just one of the essential elements of data integrity. The others are integration, data enrichment, location intelligence, and data governance.
As the global leader in data integrity, Precisely addresses all five of these elements, ensuring that business leaders have access to accurate, complete, and fully contextual information. We help businesses to ensure that their data and the insights and decisions that are driven by that data can be fully trusted.
Machine learning has the potential to delve deeply into an organization’s data assets with the end goal of automating decisions or delivering insights to support better decisions. Property and casualty insurers, for example, are applying sophisticated algorithms to examine claims information as soon as it becomes available. By analyzing both structured and unstructured data and comparing it to a growing body of information about past claims, including the severity and type of damage, the location of the loss, geospatial factors surrounding the claim, and more, insurance companies can flag cases that merit closer attention, or highlight claims that should be expedited. Using Precisely’s powerful location intelligence technology, they can even pre-position adjusters in the areas likely to suffer the greatest losses from a major weather event.
These kinds of use cases are already being applied to achieve major efficiencies. With increased scale, the benefits accruing from those efficiencies become even greater. At the same time, though, the errors and inefficiencies that can potentially arise as a result of poor data quality will scale up as well. Given our example from above, poor data quality could result in too many false positives when detecting potential cases of fraud. That, in turn, can lead to wasted effort in the fraud department, not to mention frustrated customers.
Consider what happens when some portion of an organization’s decision-making process is delegated to a machine learning algorithm. If data quality is not maintained at sufficiently high levels, ML algorithms will “learn poorly” and may develop a skewed interpretation of reality which will subsequently form the basis for automated decisions or recommendations.
One company we know had deployed a web-based order form on which they asked customers to select their industry from a drop-down list. It was a required field, but it had no impact on pricing or any other feature of relevance to the customer.


