DataOps – It’s a Secret

4 min read

Summary:  DataOps is a series of principles and practices that promises to bring together the conflicting goals of the different data tribes in the organization, data science, BI, line of business, operations, and IT.  What has been a growing body of best practices is now becoming the basis for a new category of data access, blending, and deployment platforms that may solve data conflicts in your organization.

As data scientists we all potentially suffer from that old hammer and nail meme as we tend to view the world of data through the filters of the ways we use it, predictive analytics, visualization, maybe deep learning and AI.  But the more successful our profession becomes and the wider the adoption of our point of view about using data to predict best future actions, the more likely we are to be rubbing elbows with other data-using folks who don’t necessarily share our focus.

I’m talking about all those others folks in the enterprise who use data for BI reporting of past conditions, data analysts assigned to keep those line of business users up to date with their required reports, the DevOps guys and gals building applications not necessarily related to data science, the citizen data scientists who sometimes view us as a bottleneck, and oh yes, those long suffering folks in IT on whose shoulders a lot of this rests.  Think of all these as separate tribes competing for those data resources.

The problems this competition for data and its supporting resources creates aren’t new.  Missed SLAs, confusion over a common version of the truth, and oh those squirreled away spreadsheets with data blended and sculpted into those analyses that no one else can access.  All this has a business cost as well as being a source of stress for all involved.  This cries out for a solution.

In data science the availability of end-to-end analytic platforms has brought with it the ability to access almost unlimited internal and external data sources, to blend, cleanse, analyze, model, visualize, and deploy.  But only if you’re a data scientist or citizen data scientist/analyst.  And while these platforms can create repeatable workflows, they don’t work to the benefit of the BI or DevOps teams.

Then there’s Master Data Management (MDM), aka Data Governance.  These are controls, procedures, and standards programs largely created by IT in an attempt to corral the anarchy, with an emphasis on control.  A relatively small percentage of organizations have successfully implemented MDM.  Its focus was always ‘master data’ and it didn’t solve all the other associated problems.

Each of these data-using tribes has its own tools and work arounds but a whole-organization perspective has been missing.

DataOps by itself isn’t new though it is undergoing a major expansion in scope.  In 2012 and 2013 mentions of DataOps were found mostly in digital marketing confines having to do with getting data from a large number of sources including some unstructured into a RDB for future analysis and operational use, basically ETL and blending with unstructured data

In the following few years IoT and other types of data in motion are often mentioned and since this is all now Big Data, the folks doing the promoting are often the NewSQL crowd of databases that can ingest unstructured data but retain some of the capabilities of RDBMS.

DataOps really blossomed following the formalization of DevOps in about 2013 because it was recognized that the same problems solved and principles used in DevOps could be adapted to data availability.

However, finding a consistent definition and scope for DataOps is a lot like the five blind men and the elephant, the definition has a lot to do with the professional perspective of the writer, be it BI, operations, or IT.  Today, DataOps is a series of data management principles organized into a kind of loosely agreed manifesto.  Here’s a reasonably good definition from Jack Vaughan in a 2017 techtarget post.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datasciencecentral.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.