Why You Need a Data Catalog and How To Select One

In a digital world where data lives everywhere, enterprise data catalogs are an invaluable asset in your information architecture. Over the past two years, I mentioned data catalogs for enhancing self-service BI governance, improving data quality and preparing for the GDPR . In this article, I’ll share why data catalogs have evolved from a “nice-to-have” to a “must-have”. I’ll also share tips on how to select the best data catalog for your organization from research gathered by data catalog industry leader, Waterline Data .
Why Buy a Data Catalog?
Ten years ago, reporting tools used to be complex. Only a few technical resources were granted access to known data sources for developing reports for the masses to consume. Data organization, governance and security was controllable.
Today reporting tools are simple to use and widely available. Data-driven cultures empower the masses with unprecedented self-service data source access. Concurrently, growing numbers of cloud apps, IoT and digital transformation of processes exponentially increases available data sources. Privacy regulation has also been evolving, making it progressively more difficult for organizations to effectively secure and govern their data.
To address modern data management needs, data catalog solutions like Waterline Data have become a true necessity. While it might have been considered a nice-to-have in the past, you can’t afford not to have a data catalog now due to the wide variety of data compliance regulations including the upcoming GDPR enforcement that begins in May 2018. The fines for violations can be the greater of €20 million or 4% of annual global turnover (revenue). The clock is ticking.
How to Select a Data Catalog
As you begin evaluating data catalogs , you’ll find a wide variety of solutions that resemble a data catalog but do not fulfill common requirements. You’ll also see solutions that ideally would be integrated with a real data catalog. To help you decipher what is and what is not a data catalog, here is an overview.
How a Data Catalog Works
A good data catalog serves as a searchable business glossary of data sources and common data definitions gathered from automated data discovery, classification, and cross-data source entity mapping. Automated data catalog population is done via analyzing data values and using complex algorithms to automatically tag data, or by scanning jobs or APIs that collect metadata from tables, views, and stored procedures.
Data Catalog Tagging
Automatically populated data catalogs allow crowdsourcing, human contributions of ratings, annotations, versioning, and documentation. Data catalog administrators should be able to assign data source owners, subject matter experts, stewards and consumers using role-based policies. Governed, audited data catalog access is also needed.
Data Catalog Ratings and Reviews
Data catalog solutions foster search and efficient reuse of existing data in popular business intelligence, self-service data preparation and data discovery tools. They also provide insight into end-to-end data lineage.
Data Lineage
Top 10 Data Catalog Capabilities
Waterline Data interviewed data catalog customers to learn what capabilities were most beneficial. The top 10 key capabilities list below represents the findings from that research.
Automated, intelligent population of the catalog
To efficiently scan and load metadata gathered from hundreds or thousands of potential data sources, automation is a necessity. Good data catalogs also use artificial intelligence and machine learning to intelligently profile, tag and populate objective metadata about the quality of the data.


