Data Virtualization is the CDO’s Best Friend

According to CIO magazine, the first chief data officer (CDO) was employed at Capital One in 2002, and since then the role has become widespread, driven by the recent explosion of big data.
The CDO role has a variety of definitions, but Wikipedia’s is as good as any: “a corporate officer responsible for enterprise-wide governance and utilization of information as an asset, via data processing, analysis, data mining, information trading, and other means.”
This definition also illustrates a real problem in the name. The CDO is responsible for information as an asset but could not be called the chief information officer (CIO) because that title was already taken. The distinction is important. Data is best thought of as information from which the context has been stripped, as I described in my book, Business unintelligence. That context includes meaning, usage, ownership, and more—all significant considerations for success as a CDO.
In its raw form, big data, coming mainly from external sources, is mostly loaded into data lakes—loosely structured and lightly defined data stores based on the extended Hadoop ecosystem. Often, data lakes are poorly managed, filled with duplicate and ill-defined data, leading to the moniker “data swamps,” used descriptively by Michael Stonebraker as far back as 2014.
In the same time period, businesses began to recognize the value of running analytics on social media and Internet-of-Things data—if only the data in the lake could be trusted. Suddenly, data governance, which data warehouse experts had been recommending for years, gained respectability and value. Not only that, but businesses finally accepted that responsibility for data quality rested with them, and that they needed focus and drive from the executive level. Within a few years, the CDO role had become very desirable (although, of course, it can never match the sexiness of the data scientist!)
However, appointing a CDO is only the start of a long journey to good data governance. A significant proportion of that journey will be consumed with organizational issues, methodology questions, and a struggle to keep (or gain) business commitment to the value of quality data, even when that conflicts with shorter-term financial gain. However, even with such a focus, CDOs need tools and technologies to support the undertaking, if they want to have any hope of success. But which tools and technologies will provide the most effective support?
Getting a handle on data quality in a data lake that is becoming more swamp-like with every new onboarding of external data is a difficult challenge. A data warehouse, in which governance is traditionally well-embedded, would be a better starting point from which to expand efforts.


