Implementing a real-time catalog of enterprise data assets

Good data governance is important not only for compliance, but also for data-driven competitive initiatives that business leaders care about. In the wake of Sarbanes-Oxley, few companies have been able to thrive without at least a serviceable culture of corporate governance. The same is true now for new data protection and privacy laws such as General Data Protection Regulation.
While every company is time-bound to meet compliance requirements, the most successful ones view effective governance as a business imperative, and they tend to do one thing right. In addition to governing their data and securing it such that no data, especially personally identifiable data, falls into the wrong hands, they also ensure that the right data is accessible to the right person at the right time.
Beyond defining access privileges, effective data governance means that companies need to be able to document or label their data assets, much like the labels on pharmaceuticals. Business users must be able to search across data assets, but before they access a given data set, they should be able to read about it to know what it contains, how trustworthy it is, who the data is intended for, when and how the data can be used, and other critical pieces of information. Finally, effective data governance means that data is labeled by common, standard definitions, enhanced by usage and user information.
What companies need is a centrally accessible catalog or marketplace for data assets from both within and outside the enterprise, updated in real time. But what is the best way to implement that and what is holding organizations back?
Clearly, a catalog such as this would have to overcome some challenges or risk failure: as breadth of sources, unified views, dynamic updates, and direct link to actual data delivery platforms are all potential barriers. If it were a BI dashboard, for example, reporting on a data warehouse, it would not necessarily have access to the latest data, depending on the scheduling of the appropriate ETL updates.
Also, several data governance tools would not necessarily be able to access all data sources, such as cloud repositories, transactional systems, or unstructured Big Data sources and social media feeds. A data catalog that is disconnected from data delivery platforms would also be a passive window on the infrastructure, unable to provide usage statistics, suggest dynamic uses, or prevent multiple individuals from creating multiple, ad hoc definitions for the same data sets.
Today, organizations are creating these catalogs in just a few clicks, using advanced data virtualization technology that provides unified data delivery. Data virtualization is a modern data integration approach that creates real-time, integrated views across a myriad of disparate sources, without replicating any data. It works with existing data warehouses to establish logical data warehouses or data services, which can also access data scattered across multiple sources including cloud-based and transactional sources.
Data virtualization enables a single layer from which to access all data across the enterprise for analytics, digital applications, and directly to users for information self-service.


