DataOps: Navigating the Data Wastelands

After years of defining (and redefining) the term data-driven, companies are beginning to realize the stark truth: we’re rich in data and poor in actionable insights.
Blaming everything on bad data or misaligned data science teams is missing the point. Most of the time, the gems of insights are locked in bits and bytes. While it is the job of data scientists and analysts to sieve through the river of data to find data nuggets, traditional data management makes it difficult. Add the multitude of data teams with different goals, and you can imagine the complexity.
Andy Palmer, chief executive officer and co-founder of Tamr, points to Wikipedia offering a comprehensive definition: “DataOps is an automated, process-oriented methodology, used by analytic and data teams, to improve the quality and reduce the cycle time of data analytics.”
DataOps is not the data science version of DevOps. Nevertheless, the loose analogy puts its value proposition in context.
“If the purpose of DevOps was to increase feature velocity for the large internet companies, then the purpose of DataOps is to increase analytic velocity for the large enterprise,” Palmer explains.
The significant difference with DevOps is that the core artifact here is not software but data. Also, data is generated and consumed very differently, Palmer adds.
DataOps acknowledges that we need to respect data better.
“For decades, data has been treated as an exhaust from operational systems instead of as a strategic asset. We’ve all got a ton of work to do in order to build the next generation of modern data engineering infrastructure in large enterprises — there are huge gaps and tons of legacy to be replaced,” says Palmer.
Doing such can free data scientists and analysts to focus on their job and spend less time preparing the data. It also allows data engineers to weave and manage different data pipelines in a more standardized fashion.
It is not a technical challenge that DataOps is trying to solve. For Palmer, it is addressing the human challenge in data science.
Tamr takes it a step further by using machine learning to help humans cope with vast workloads. “At Tamr, we work with our customers to help them realize the power of using modern machine learning techniques to master their data and deliver on the unfulfilled promises of traditional MDM.”
Still, there are gaps. One huge one is our approach to data governance. Currently, it is “source-based,” but Palmer believes it should be “consumption-based.”
He draws an analogy with water management. Many data governance programs are focused on where the sources are coming from. For water, such sources are rivers, lakes, or rain clouds. While this is always good information, “the most important question for data as with water is what’s coming out of the faucet and do you want to drink it,” Palmer argues.
Consumption-based governance focuses on the data consumer. It adds a framework and monitors using the “information access policy of the company and the requisite information access [across] roles/personas.”
“Most companies are so distracted with source-based governance that they don’t have the time or resources to govern the data which is being consumed — which I believe is by far the more important question,” said Palmer.


