Simplifying machine learning through data virtualization

Machine learning, in which systems solve problems by recognizing patterns rather than by following specific sets of instructions, has recently been making headlines. Machine learning is supporting diverse domains such as business intelligence, earth science, and online personalization. Organizations are using machine learning to facilitate complex use cases such as speech recognition, fraud detection, and demand forecasting. Enabling machine learning requires sophisticated infrastructures that can quickly integrate and process large amounts of data from disparate sources, often involving multiple data platforms, tools, and processing engines. However, establishing and maintaining such infrastructure can be complex and costly, but data virtualization can greatly simplify the data integration process while saving costs, thus accelerating machine learning initiatives.
To support machine learning, many organizations leverage data lakes, as they can collect large volumes of data from multiple sources, including structured and unstructured sources, and store the data in its original format. However, storing data in different formats does not necessarily facilitate discovery, as data in different formats must first be integrated before it can be leveraged for machine learning. Due to the increasingly distributed nature of data infrastructure in today’s enterprises, data integration gets more complex. Data scientists can spend up to 80 per cent of their time on these tasks, suggesting that it is time for a new approach.
In addition, the slow, costly replication of data from systems of origin can mean that only a small subset of the relevant data will be stored in the data lake. Companies may have multiple data repositories distributed across a number of different cloud providers and on-premises systems.
The burden of adapting the data for machine learning then falls on data scientists who, while able to access the necessary processing capacity, tend not to have the skills required for integration. The past few years have seen an emergence of data preparation tools designed to help data scientists carry out simple integration tasks, but many tasks require more advanced skills. An organization’s IT team may be called in to create new data sets in the data lake specifically for machine learning purposes, but this can significantly slow down the overall initiative.
If organizations are to unlock the full benefits of data lakes, and other diverse sources, to support machine learning, new technologies are needed.
The Many Benefits of Data Virtualization
Rather than moving data from multiple data sources to a new, centralized repository, data virtualization creates real-time, consolidated views of the data, leaving the data in its original locations.


