Building a Data Platform: 7 Signs It’s Time to Invest in Data Quality

One of the most frequent questions I get from customers is: when does it make sense to invest in data quality and observability?
My answer is, it depends.
The reality is that building a data platform is a multi-stage journey and data teams have to juggle dozens of competing priorities. Data observability may not make sense for a company with a few dashboards connected to an on-premises database.
On the other hand, many organizations Ive spoken with have increased their investment in developing their data platform without seeing a corresponding increase in data adoption and trust. If your company doesnt use or trust your data, your best laid plans for data platform domination are a pipe dream.
To answer this question, Ive outlined seven leading indicators that its time to invest in data quality for your data platform.
From speaking with hundreds of customers over the years, I have identified seven telltale signs that suggest your data team should prioritize data quality.
Whether your organization is in the process of migrating to a data lake or between cloud platforms (e.g., Amazon Redshift to Snowflake), maintaining data quality should be high on your data teams list of things to do.
After all, you are likely migrating for one of three reasons:
Regardless of why you migrated, its essential to instill trust in your data platform while maintaining speed.
You should be spending more time building your data pipelines and less time writing tests to prevent issues from occurring.
For AutoTrader UK, investing in data observability was a critical component of their initial cloud database migration.
As were migrating trusted on-premises systems to the cloud, the users of those older systems need to have trust that the new cloud-based technologies are as reliable as the older systems theyve used in the past, said Edward Kent, Principal Developer, AutoTrader UK.
The scale of your data product is not the only criteria for investing in data quality, but it is an important one. Like any machine, the more moving parts you have, the more likely things are to break unless the proper focus is given to reliability engineering.
While there is no hard and fast rule for how many data sources, pipelines, or tables your organization should have before investing in observability, a good rule of thumb is more than 50 tables. That being said, if you have fewer tables, but the severity of data downtime for your organization is great, data observability is still a very sensible investment.
Another important consideration is the velocity of your data stack growth. For example, the advertising platform Choozle knew to invest in data observability as it anticipated table sprawl with their new platform upgrade.
When our advertisers connect to Google, Bing, Facebook, or another outside platform, Fivetran goes into the data warehouse and drops it into the reporting stack fully automated. I dont know when an advertiser has created a connector, said Adam Woods, CTO, Choozle. This created table sprawl, proliferation, and fragmentation. We needed data monitoring and alerting to make sure all of these tables were synced and up-to-date, otherwise we would start hearing from customers.


