Data catalogs fuel increased efficiency, speed-to-insight

As organizations collect more data and develop more data assets, data catalogs can be critical.
Data catalogs are organized repositories where data users and analysts can search for andfind the data they need for their work.
With organizations amassing terabytes and petabytes of data and, depending on the size of the enterprise, building hundreds and perhaps thousands of reports, dashboards, data models and other data assets, finding the table, chart or data set without knowing exactly where to look can be difficult — and maybe almost impossible.
Data catalogs solve that problem, indexing data assets and making them easy to search, find and be put to use tomake data-driven decisions.
The result is efficiency that increases speed-to-insight.
In addition, data catalogs can enable collaboration among business users working together or on similar projects, and governance to set limits on who can view and use what data.
What exactly is a data catalog?
Wayne Eckerson: It’s not unlike a card catalog in a library, except it’s all digital. What it’s doing is collecting metadata from all the sources of information throughout your enterprise, then pulling that metadata together in one place and indexing it so it can be easily searched. It provides a lot of descriptive characteristics of the data so people can profile it and see who else has used it and any annotations they have left. It’s a great way to give people one-touch access to the information assets that are available to them in their organization instead of having to hunt around and ask people.
They’re definitely used to facilitate self-service BI, and they’re also a good way to curate data and manage access to it. Once you put all the metadata in one place, you can have a data curator go in and define who gets access to what metadata, and hence the actual data as well.
What do data catalogs enable organizations to do that they otherwise can’t — why are they important?
Eckerson: I mentioned self-service before — they really facilitate finding and profiling data so you can accelerate your time-to-insight. Along those lines, the powerful thing about data catalogs is that they can capture the tribal knowledge of an organization that is usually in the heads of just one or two people who function as human data catalogs, and you have to know them to [access to their knowledge]. It’s not a very streamlined process. The idea is that we can collect all this metadata, and then people can not only search but also annotate what they found and how they used the data and leave breadcrumbs along their trail so that others who follow them — whether the next day, the next month, the next year or the next decade — have access to that tribal knowledge. That can really help improve the usage of those assets.
On the other side, they enable data governance and data curation. Not all data is equal from a privacy and security perspective so you want to have access policies, and those can be enacted in the data catalog. Meanwhile, data stewards and curators can use data catalogs as a way to clean up the data and identify things that are missing, inconsistent, duplicative and then fix them.
When were data catalogs first introduced and how have they evolved since then?
Eckerson: You started to see the first couple of catalogs around 2015.


