Leaky Pipelines and the Business Case for Data DevOps

3 min read
Curated from datanami.com →

As enterprises embrace digital transformation and migrate critical infrastructure and applications to the cloud as a key component of those efforts, what can be called “data clouds” have started to take shape. Built on multi-cloud data infrastructure, such as Databricks’ or Snowflake’s data platforms, these data clouds enable businesses to break free from application and storage silos, sharing data throughout on-premises, private, public, and hybrid cloud environments.

As a result, data volumes have skyrocketed. Big data-enabled applications increasingly generate and ingest more and different types of data from a variety of technologies, such as AI, machine learning (ML), and IoT sources, and the nature of data itself is radically changing in both the volume and shape of data sets.

As data is freed from constraints, visibility into the data lifecycle gets fuzzy, and traditional quality control tools quickly become obsolete.

For the typical enterprise, data monitoring and management is still handled by legacy tools that were designed for a different era, such as Informatica Data Quality (released in 2001) and Talend (released in 2005). These tools were designed to monitor siloed, static data, and they did that well. However, as new technologies entered the mainstream — big Data, cloud computing, and data warehouses/lakes/pipelines — data requirements changed.

The legacy data quality tools were never designed (or intended) to serve as quality control tools for today’s complex continuous data pipelines that carry data in motion from application to application, and cloud to cloud. Yet, data pipelines frequently feed data directly into customer experience and business decision making software, which opens up massive risks.

A good example of how bad data can escape notice and undermine business goals is “mistakeairline fares.” Erroneous currency conversions, human input errors, and even software glitches generate mistake fares so often that some travel experts specialize in finding them.

This is just one example. Bad data can result in incorrect credit scores, shipments sent to the wrong addresses, product flaws, and more. Market research firm Gartner has found that “organizations believe poor data quality to be responsible for an average of $15 million per year in losses

As developers rush to catch up to the challenges of maintaining and managing data in motion at scale, most first turn to the DevOps and CI/CD practices they used to build modern software applications. To port those practices to data, however, there is a key challenge: developers must understand that data scales differently than applications and infrastructure.

With applications increasingly being powered by data pipelines from cloud-based data lakes and warehouses and streaming data sources (such as Kafka and Segment), there needs to be continuous monitoring of the quality of these data sources to prevent outages from occurring.

Organizations must ask, who is responsible for inspecting data before it hits data pipelines, and who cleans up the messes when data pipelines leak or feed bad data into mission-critical applications? As of now, the typical business’ approach to pipeline problems and outages is a purely reactive one, scrambling for fixes after applications break.

If today’s typical multi-cloud, data-driven enterprise hopes to scale data platforms with agile techniques, DevOps and data teams should look back on their own evolution to help them plan for the future, specifically taking note of a missing component as DevOps is ported to data: Site Reliability Engineering (SRE).

DevOps for software only succeeded because a strong safety net, SRE, matured alongside it. The discipline of SRE ensured that organizations could monitor the behavior of software after deployment, ensuring that in-production apps meet SLAs in practice, not just theory.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datanami.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.