Data Catalogues might be the new Black, but metadata discovery to provision them can be tricky

Data catalogues have become a fast growing solution to a critical requirement of an enterprise information management strategy. This is the need to document and understand the types and uses of data across the enterprise application landscape.
The importance of data catalogues to a data management strategy is emphasised in Gartner’s 2017 Report: ‘Data Catalogs Are the New Black in Data Management and Analytics’.In it they suggest that “Data catalogs enable data and analytics leaders to introduce agile information governance, and to manage data sprawl and information supply chains, in support of digital business initiatives.”
They provide both a place to store information about an organisation’s data assets as well as mechanisms for utilising, enriching, managing and valuing that information. “Metadata is the core of a data catalog. Every catalog collects data about the data inventory and also about processes, people, and platforms related to data. Metadata tools of the past collected business, process, and technical metadata, and data catalogs continue that practice.” Eckerson Group:‘The Ultimate Guide to Data Catalogs.’
Most large organisations have a wide variety of applications, file stores and home grown systems that they have acquired over the years. These contain data vital to the business. Whilst more enlightened businesses have striven to optimise their understanding and use of data in the past, it is only relatively recently that trying to document and categorise this data has become more mainstream. Much of this is driven by increasing compliance requirements (e.g. the EU GDPR) as well as by the more widespread acceptance of data as an asset whose use can be better optimised.
Rather than contain actual data, a data catalogue delivers value by providing a mechanism for technical and business users alike to make use of the metadata, or data structures which underpin their source systems.
This means that one of the early tasks to undertake during the implementation of a data catalogue is to identify the sources of that metadata and to start the process of importing it. Any good catalogue solution has a range and of scanners and connectors for identifying and mapping metadata from many different sources.
In the Eckerson Group’s, Ultimate Guide to Data Catalogs: “The initial build of a data catalog typically scans massive volumes of data to collect large amounts of metadata. The scope of data for the catalog may include any or all of data lakes, data warehouses, data marts, operational databases, and other data assets determined to be valuable and shareable. Collecting the metadata manually is an imposing and potentially impossible task. The data catalog automates much of the effort using algorithms and machine learning to accomplish the following:
There are a variety of other methods of capturing metadata, for example via crowd sourcing or catalogue curator based activities.
However there are also various classes of system whose valuable metadata is not accessible or usable by those normal methods. In Eckerson’s Ultimate Guide to Data Catalogs they refer to these as “ challenging data sources” Amongst the most difficult of those are the large, complex and often highly customised ERP and CRM packages from SAP, Oracle, Microsoft and others. They hold large amounts of transaction data critical to the success of many thousands of businesses. It is important therefore that their metadata should be included in the data catalogue.
The problem they pose is that there is no meaningful metadata, by which I mean ‘business names’ and descriptions for tables and attributes and no table relationships defined in the database system catalogue.


