Stop Treating Your Data Engineer Like A Data Catalog

Data trust starts and ends with communication. Here’s how best-in-class data teams are proactively certifying tables as approved for use across their organizations.
Say it with me: data engineers are not data catalogs.
You would be hard pressed to find “answering multiple Slack messages every week about which tables are good to use for this report,” in their job description, but it happens nonetheless.
Data analysts aren’t psychic. Yet, they are often placed in the position of having to intuit if the data being piped is trustworthy.
This misalignment has arisen as data teams are pushed to move faster, weave themselves across the data mesh , and enable increasingly self-service data platforms.
Without a data certification program in place, data teams don’t often know where to go to get the best data. Image courtesy of Chokniti Khongchum on Shutterstock, purchased for use with Standard License.
It’s the data team’s equivalent of the classic document version control issues that have plagued knowledge workers for decades. What starts as a tight pitch deck evolves into:
A million people making and sharing ad-hoc slides;
Massaging content on those slides until it becomes an echo of its original intent; and
Creating copies labeled V6_Final_RealFinal .
The same thing happens across the data team. Everyone is trying to do the right thing (i.e., support your stakeholders, generate insights, pipe more data, etc.), but everyone is also moving fast.
One day you look up and notice you have 6 different models with slight variations essentially doing the same thing…and no one knows which one is most up-to-date or even which field to use.
This creates real operational problems downstream including:
Inefficient cycles of redundant “traffic control;”
Increased data downtime
When you don’t trust your data or you have lower data reliability, organizations often pad the margins of error in their forecasts.
As highlighted by Peleton’s recent production halt , poor forecasting can be especially problematic during the pandemic when uncertainty across demand, supply chains and the overall business environment is at an all-time high.
More data discovery, more problems
We’ve written previously about data discovery , a new approach to understanding the health of your distributed data assets in real-time, and it’s an essential part of the solution.
Data discovery provides distributed, real-time insights about data across different domains, all while abiding by a central set of governance standards. Image courtesy of Barr Moses.
Data discovery provides a domain-specific, dynamic understanding of your data based on how it’s being ingested, stored, aggregated, and used by a set of specific consumers.
As with a data catalog, governance standards and tooling are federated across these domains (allowing for greater accessibility and interoperability), but unlike a data catalog, data discovery surfaces a real-time understanding of the data’s current state as opposed to it’s ideal or “cataloged” state.
It is especially useful when teams take a distributed approach to governance that holds different data owners accountable for their data as products, which allows data-savvy users throughout the business to self-serve from those products.
But as data becomes more accessible, how can downstream stakeholders determine what data sets have been served, transformed, and approved by a given domain’s data team?
How can one domain be sure a common set of data quality standards, ownership, and communication processes are being upheld across the organization?
One of my customers, a leading media company with a mature data organization, was facing these exact questions. As a result, we have been working with them and several others to implement a data certification program.


