Data Virtualization: Thinking Outside the Bowl

4 min read

Across all vertical markets, organizations adopt data virtualization as a core part of their IT infrastructure because data is becoming more distributed, more heterogeneous, and larger in volume. The pressures on business to be more competitive are continually growing, so it is more important than ever before to get fast access to relevant data.

The time is right to realize the potential of data virtualization outside of the logical data warehouse. While a powerful and flexible architecture, the logical data warehouse is a bit like a fishbowl, there is a larger world waiting outside of it, including a wide variety of different fishbowls. The time is right to take that first leap into that larger world.  

Gartner’s Mark Bayer came up with the term logical data warehouse 10 years ago to describe the need to logically expand existing data warehouse architectures that were supposed to include all enterprise information but didn’t. Since then, terms like “logical data warehouse” and “hybrid architectures” have been widely used to represent the natural evolution of analytical systems, and data virtualization has played an important role in this process.

However, Rick F. van der Lans illustrated in his post “Unifying the Data Warehouse, Data Lake and Data Marketplace”  that data virtualization can act as a “unified data delivery platform” that brings together not only the data warehouse but also data lakes, data marketplaces, streaming data, and any other data delivery system, helping to tear down data silos to support a wider range of use cases.

If we consider the capabilities of data virtualization outlined in the diagram, I think it’s clear how data virtualization can form an important part of the unified data delivery platform described by Rick.

Key aspects of data virtualization are its ability to abstract the location and implementation of the data, provide increased speed of data delivery (since data doesn’t need to be replicated) and enable users to consume data in a self-serve manner through the discovery of a published logical data model. Data consumers come to the data virtualization layer to access a common, meaningful model using familiar technologies like SQL or REST, but the data remains in-situ and the data virtualization layer handles the integration of the data on-demand.

This abstraction loosens the coupling of IT and the business, enabling IT stakeholders to work at their own pace to provision data using their most appropriate processing platform, while consumers are provided access to the data wherever it resides through the logical layer. This removes silos and enables consumers to access the data using the applications of their choice, be they reporting tools, data science tools, enterprise applications, or mobile or web apps.

Data virtualization therefore forms the backbone of a unified data delivery platform, supporting the modernization of the data landscape behind it.

With this post, I’d like to get you to think differently about data virtualization, so you can imagine where else a unified data delivery platform can add value to your organization. To this end, I’d like to highlight three architectures in which data virtualization can play an increasingly important role.

The first category we’ll look at is the use of data virtualization as a data services layer.

In this scenario, data virtualization publishes certified data services that can be used by any development team, which avoids the problems of project teams creating their own siloed data sets or using non-certified or non-maintained data sources.

With a data services layer, there’s no direct access to the data sources. Consistent data sets are published and made discoverable through the use of data virtualization, which enables any developers, not just BI teams, but developers of operational applications, web portals, back or front office systems, to access and reuse the data.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datavirtualizationblog.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.