Meet the data quality dimensions

Data is valuable if it is of good quality. Data quality dimensions are the characteristics against which we measure quality.
In this article, we will look at these dimensions, what they track and why each is important.
To ensure that data is trustworthy, it is important to understand the dimensions of data quality. Data quality dimensions will help you assess if your data is good enough to use or if you need to make improvements. A single dimension will not be sufficient to assess the quality of your data. You will need to select the dimensions that will describe it as fit for the purpose it’s intended to be used.
While we recognise that organisations may define different quality dimensions, we recommend these six dimensions, as defined by the Data Management Association UK (DAMA(UK):
We have accuracy when data reflects reality. For example, this can refer to correct names, addresses or represent factual and up to date data. A likely place for errors to occur is right at the start, during data collection. Look closely at each data field to see if the values look plausible. The height of a person recorded as 15 cm is clearly wrong.
Real-world information can change over time. This makes accuracy quite challenging to monitor. A change in the personal circumstances of a claimant may affect the housing benefit that the person is entitled to. You should regularly review data that is likely to change over time.
High data accuracy allows you to produce analytics than can be trusted; it also leads to correct reporting and confident decision-making.
Data is considered complete when all the data required for a particular use is present and available to be used. It’s not about ensuring 100% of your data fields are complete. It’s about determining what data is critical and what is optional.
Consider patient records consisting of personal details and medical history. Missing information on allergies is a serious data quality problem because the consequences can be severe. On the other hand, if there are gaps in email addresses, this may not have an impact on patient care.
Completeness applies not only at the data item level but also at the record level in terms of assessing whether your data set meets your expectations of what’s comprehensive. This measure helps you understand if missing data will have an impact on its use and affect the reliability of insights that you gather from it.
Completeness is not the same as accuracy as a full data set may still have incorrect values. You may have full information about people claiming benefits, but this does not mean that the information is correct.
Uniqueness measures the number of duplicates. Data is unique if it appears only once in a data set. A record can be a duplicate even if it has some fields that are different. For example, two patient records may have different addresses and contact numbers, but if they both refer to the same patient there is duplication. In this instance, it is likely that healthcare providers may miss critical information because it is in the duplicate.
Duplication is a particular risk when combining data sets. You should always check your data for uniqueness.


