What Is Data Observability and Why Does It Matter?

In 2021, most businesses’ websites, pitch decks, and job descriptions lay claim to their identities as data-driven companies. Behind this desirable label, you’ll usually find an admirable goal—and not much else. Well, beyond an awful lot of unused data.
Well-meaning companies are collecting and storing data, but don’t allow it to actually drive their operations. We’ve found this is often due to a complete, and not unwarranted, lack of trust in the data itself.
As companies gather seemingly endless streams of data from a growing number of sources, they begin to amass an ecosystem of data storage, pipelines, and would-be end users. With every additional layer of complexity, opportunities for data downtime — moments when data is partial, erroneous, missing, or otherwise inaccurate — multiply.
Forrester estimates that data teams spend upwards of 40% of their time on data quality issues instead of working on revenue-generating activities for the business. And too many data engineers know the pain of that panicked phone call — the numbers in my dashboard are all wrong. I have a presentation in an hour. What happened to our data?
At this point, unreliable data is an expected inconvenience. It’s a pain, sure, but it’s also a problem for almost every company on a regular basis. And it’s hard to let something you don’t trust drive your company’s decision-making, let alone your product.
Twenty years ago, when a company’s website or application went down or became unavailable, it was considered par for the course. Today, if Slack, Twitter, or LinkedIn suffers an outage, it makes national headlines. Online applications have become mission-critical to almost every business, and companies invest appropriately in avoiding service interruptions.
That’s how far application engineering has come in the last several years. And that’s precisely where data engineering needs to go.
We predict that as companies increasingly rely on and invest in data as a key business driver, instances of broken pipelines, missing fields, null values, and the like will go from a tolerated inconvenience to inconceivably rare. Executives won’t have their data team leads on speed-dial for troubleshooting when an important dashboard looks out of whack — because it won’t happen. Over the next decade, we’ll reach a new level of data reliability. And we’ll get there by following in the footsteps of application engineers to develop our own specialty of DataOps.


