5 Guidelines for Building a Successful Data Catalog

Compliments of Zaloni: Download free eBook “Architecting Data Lakes” to learn the key to building and managing a big data lake, brought to you in partnership with Zaloni.
At times, the search for a perfect data catalog can seem like finding the needle in a hay stack. Each stakeholder has equally demanding and disparate sets of requirements for success. Where business analysts want a slick, refined, and easily navigated UI with easy export capabilities, data scientists might refuse to accept anything that does not allow custom-tailored queries, connections to their favorite notebook, and unburdened access to all of the data that has ever existed in the data lake. Meanwhile, the security group wants none of this! Exposing the data at all is a non-starter.
This leaves you — the tech visionary who has a stable of cutting-edge vendors at the ready and a five-year rollout plan to go with them — stuck in neutral.
Before you resort to breaking out the floppy disks in protest, we have five guidelines for building a successful data catalog and how you can help your business succeed without compromising your stakeholders.
Open access is the foundation of any successful data catalog. The desire for a more intuitive, less burdensome method of accessing available data is when the demand for a data catalog commonly emerges. Any solution that does not address this fundamental point is bound to run into difficulties in gaining business support.
For users, open access provides the value that is sorely lacking from traditional data lakes: efficient, accurate, and personalized access to data, no matter where that data may originate.
For administrators, open access can dramatically reduce the overhead that comes from routine requests and maintenance of audit histories and ticketing systems. A good catalog should automate and manage these functions.
Value: Open access to any data owned or used by consumers of the catalog creates fast time-to-value for users and greater efficiency for administrators.
Security in the data catalog often creates a significant conundrum. While providing open, transparent access to data is tantamount to success, this access can also create security threats that will quickly shut the project down. Although security groups at your organization may not have an active role in using the data catalog, they certainly have a vested interest in keeping the organization protected from both external and internal threats. How can you balance these two seemingly opposing forces?
Fortunately, modern Hadoop technologies provide a breadth of options. Projects such as Ranger, Knox, and Sentry provide a new level of protection for at-rest data on the cluster, while more tools such as NiFi and Kafka either include support for external security protocols (i.e., Kerberos) or built-in security layers. Assuming we focus on Hadoop-based data catalogs, these tools will take care of most facets of security within the data lake. Any remaining holes will come from external systems, the burden of which should fall on those systems.
Value: Security, whether applied by the catalog or by underlying systems, is integral to continued operation at an enterprise level. Hadoop has several components to address this requirement.
To leverage Hadoop as a core component of the data catalog, the catalog itself must have a foundation in Hadoop.


