The Golden Record: Explained

2 min read

Where ‘big data’ appears to be the skeleton key that will unlock everything and all you want to know about your business, there’s more than meets the eye when it comes to understanding your data. Yes, clean data will unlock incredible value for your enterprise; inaccurate records, on the other hand, are a significant burden on our productivity.

This is why we all seek the “Golden Record”.

The Golden Record is the ultimate prize in the data world. A fundamental concept within Master Data Management (MDM) defined as the single source of truth; one data point that captures all the necessary information we need to know about a member, a resource, or an item in our catalogue – assumed to be 100% accurate.

Its power is undeniable. However, where we have multiple databases, working out how to achieve such perfection is hard to ascertain. As such, we must first understand the benefits of a golden record.

So, let’s step back to your childhood and consider how imperfect information can cause havoc in any system.

To explain, we are going to go back to a very simple example – You’re 13-years old. Imagine sitting in your classroom. Everyone has arrived, and the teacher is about to run through the register.

They have drawn the names from various local databases and put them on paper without properly checking who is meant to be where.

After a few minutes, something isn’t right.

The teacher is repeating themselves. No one is sure why. The process is taking much longer than it should and the kids are becoming restless. We take a closer look at the register, then everything becomes clear.

Seemingly duplicated records – the bane of the Master Data Management world.

When building databases from disparate sources, we often run into the issue of duplication. Whether resulting from incomplete entries, changes that occur over time or some other reason, this is a significant issue for any enterprise that relies on vast volumes of information.

As you may imagine, if we were to expand the rollcall example to include hundreds of thousands of names, the overhead of duplication becomes exponentially worse, with every process draining an increasing volume of resource.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datasciencecentral.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.