Data Strategy Drives Better Quality, Cost Savings

4 min read

Technologies and techniques such as robotics, machine learning, natural language processing, and blockchain present a defining opportunity for the finance function. However, companies that have not adequately invested in data will find their digital transformation efforts frustrated.

The Financial Education & Research Foundation is working on a report with EY to help senior-level financial executives better understand data strategy and quality needed for their digital transformation.

As part of that effort we speak with Professor Dave Waterman, an adjunct lecturer at Georgetown University and a data scientist at U.Group.

A transcript of the discussion appears below the podcast player.

FERF: What are some of the ways that low-quality data impacts an organization?

Dave Waterman: The issue with low-quality data is that it’s going to take more effort to make it useful, and so there’s going to be an increased time and increased cost at every step along the way. To load that data, to analyze it, to process it, to use it for making decisions, all of those things are going to be more difficult and more complicated as the data quality degrades.

As a simple example, if your neighbor’s kid comes over to you and asks you to help out with her lemonade stand and asks you to look over her financials, if she comes to you with an Excel spreadsheet, you think, ‘Oh, this is great. She has her business in order. I can just look it over, tell her where she’s doing well, and where she needs to improve.’ But when you start taking a closer look, you see that maybe ten percent of the fields are missing from the spreadsheet or some of the column names don’t make sense. Maybe they’re abbreviated and you can’t figure out what they are. Or the data’s not consistent. Some of the fields have numbers in them, some of them have words in them. All of those things are going to make it more difficult for you to figure out what the data is telling you.

That’s an example of the kinds of data that we see every day. It’s not uncommon for me to see a data set where ten percent of the fields are missing. That doesn’t mean that that data is useless. It just means that I have to do a lot more work to impute what could be there: What does that missing data mean, and how can I work around it or work with it?

All of those things are going to slow down the work that you do. They’re probably going to have to involve a lot more conversations between the involved parties. If that’s a difficult communication to make, then that’s going to slow you down even more. All of those things are going to result in less accuracy in your results, less clarity in your decisions, less detail in your conclusions, and so that’s going to lead to missed opportunities.

FERF: You mentioned that you often see data that’s incomplete or messy. What are some of the reasons why data can be of poor quality?

Waterman: The most common reason would be because it was input by a person and that there was no validation done when that data was inputted. If you allow people to enter data manually and you don’t have any checks before it gets stored, you’re going to see some really crazy stuff.

You can have data validation that lets stuff through, and that can be a problem too. But, really, you need to architect your data structures from the very beginning with a knowledge and understanding of what’s going to be going into those fields so that you can set clear limits so that it makes it much more difficult for individuals to enter in poor data.

And you can still have data quality issues from data that comes automatically. Think about if you have a computer server somewhere that’s running and it’s making records over time. Maybe the network connection goes down, and so you lose a day’s worth of data from that server. Those things happen too. Sometimes it’s unavoidable. But the biggest problem that we see generally comes from when an individual has the ability to enter data on their own.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at financialexecutives.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.