Data quality and reliability for the cloud

The holy grail of “trust in data” from data to insight journey of enterprises is not entirely new. Since BI and analytic workloads are separated from data warehouses, the chasm has widened.
There’s an even larger gap between what business needs, business operations supported by the IT application landscape, and the reliability of the data accumulated in the data warehouses for the business teams.
In this mayhem, data quality solutions and tools were buried deep in MDM and data governance initiatives. Still, two challenges existed – The first was to look into the past while asking whether data was trustable.
Second, ‘quality’ was measured with respect to the golden record and master data – standardization, which itself was constantly evolving.
While the big data hype started with Hadoop, concerns with volume, velocity, and veracity were tackled, this remained an enterprise play.
True innovation kick-started with MPP systems like Redshift on AWS built cloud natively, which guaranteed a higher performance to handle massive datasets with good economics and a SQL-friendly interface.
This, in turn, spurred a set of data ingestion tools such as Fivetran, which made it easier to bring data onto the cloud.
Today, data is getting stored in data lakes on cloud file systems and cloud data warehouses, and we see this reflected in the growth of vendors like Databricks and Snowflake.
The dream of being data-driven looked much closer than before.
Business teams were hungry to analyze and transform the data to their needs, and the BI tool ecosystem evolved to create the business view on data.
The facet that changed beneath and along this evolution is that data moved from a strictly controlled and governed environment to the wild west as various teams are transforming and manipulating data on the cloud warehouses.
It’s not just the volume and growth of data. The teams hungry for data (data consumers) have also exploded in the form of BI teams, analytic teams, and data science teams.
In fact, in the digital native organizations (which were purely built on the cloud), even the business teams are data teams. E.g., a marketeer wants real-time information on product traffic to optimize campaigns.
Serving these specialized and decentralized teams with their requirements and expectations is not an easy task.
The data ecosystem responded with a clever move, marking the beginning of data engineering and pipelines as a basic unit to package the specialized transformations, joins, aggregations, etc.


