Data quality and data preparation in self-service

2 min read
Curated from blogs.sas.com →

The rise of self-service analytics, and the idea of the ‘citizen data scientist’, has also brought a number of issues to the fore in organizations. In particular, two common areas of discussion are the twin pillars of data quality and data preparation.

There is no doubt that good quality, well-prepared data is essential for any analytics process, particularly self-service. Without it, nobody can trust the results, and the process will be largely pointless. One of my colleagues, however, likened self-service with unknown data quality as like driving at 100mph in fog: it’s unlikely to turn out well, but a better driver may be able to handle it more easily and for longer. And a better driver, equipped with good tools such as high quality navigational aids, is likely to be in an even better position. In other words, quality matters not just for the data, but also the analyst and the tools.

But whose responsibility is data quality? And how do data quality and data preparation actually fit together? Many business users would say that both were an IT responsibility. The IT department, after all, is responsible for setting out the rules on data governance, and (probably) only they have the ability to clean the data properly, and make sure that it is the necessary quality. But I think this is an abdication of responsibility by business users.

Business users are the primary owners of their own data. They generate it, they know it well, and they also understand best when something is wrong. It is far easier for them than the IT department to recognize a problem with data quality. Let’s consider a well-known example: coding data in hospitals (the information that tells you the patient’s diagnosis and any procedures or treatments). Who is more likely to recognize that a code is incorrect, the person who put it in and therefore knows what it should say, and will recognize that (say) nobody under Department X should have procedure Y, or the technical specialist who manages the IT system and is nominally responsible?

The answer is clearly the person who provided the input. So does this mean that the data quality and preparation process should also be self-service? I think the answer is yes, to a certain extent.

Once you introduce self-service analytics, business users can really start to see the benefits of good quality data.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at blogs.sas.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.