Why data virtualization offers a better path to decision making

It’s hard to ignore the clamor inside and outside of government circles for agencies to harness data more effectively.
One of underlying lessons revealed thus far by the ongoing pandemic has been the challenges federal and state leaders have faced trying to obtain current, reliable data in roughly real time in order to make critical public policy decisions.
Within government agencies, however, there’s a deeper issue: How to manage vast repositories of data — and develop more cohesive data governance strategies — so that whether you’re an analyst, a program manager or an agency executive, you can get the information you need, when you need it and the way you need it.
That’s why data virtualization is emerging as a pivotal solution in the eyes of a growing number of IT experts.
One of the ironies of the Big Data movement over the past decade is that while it helped organizations come to grips with the explosion of data being generated every day, it also caused a bit of damage by overpromising and underdelivering. The evolution of the cloud and software like Apache Hadoop was supposed to help government agencies, for instance, pour their siloed data sets into vast data lakes, where they could merge, manipulate and capitalize on the previously unseen value lying dormant in all those databases.
A lot of CIOs said, “Great, we have this gigantic data lake, and we can dump everything into them, and that will resolve all or most of our data sharing problems.”
Yet, over time, what many agencies discovered they had created was more of a data swamp.
Part of the problem stems from the fact that organizations hadn’t always taken the necessary steps to resolve and apply adequate data governance structures to their digital assets. Another factor is the extent to which agencies still need to complete the work of inventorying and properly cataloging those assets.
But perhaps the biggest issue is the assumption that data lakes make sense in the first place.
For many organizations, data lakes make data readily discoverable and easy to analyze. But in government, where data tends to remain federated, there are inherent inefficiencies in physically pooling data. It’s not that people don’t want to share data — although some are more willing to share than others — it’s just that it’s complicated to do so.
It often takes a dozen or more applications to properly locate, identify, authenticate and catalog data before migrating it to a data lake or the cloud; and depending on the type of data, it can also require a suite of other products to harmonize it with other data and make it useful to different stakeholders.


