Data Science for the Modern Data Architecture

Our customers increasingly leverage data science and machine learning to solve complex predictive analytics problems. A few examples of these problems are churn prediction, predictive maintenance, image classification, and entity matching.
While everyone wants to predict the future, truly leveraging data science for predictive analytics remains the domain of a select few. To expand the reach of data science, the modern data architecture (MDA) needs to address the following four requirements:
The below diagram represents where data science fits in the MDA.
The end-users consume data, analytics, and the results of data science analytics via data-centric applications (or apps). A vast majority of these applications today don’t leverage data science, machine learning, or predictive analytics. A new generation of enterprise and consumer-facing apps are being built to take advantage of data science/predictive analytics and provide context driven insights to nudge end-users to next set of actions. These apps are called data-smart applications.
Writing data-smart apps is hard. The app developer needs to write not only the traditional app logic but also the logic to invoke predictive analytics. These data-smart apps also face a set of common problems such as entity disambiguation, data quality analysis, and anomaly detection. Since today’s data platforms don’t provide these functionalities, the app developers are responsible for solving these problems.
We have seen this issue before, and frameworks such as JavaEE & Spring Framework evolved to addresses common application concerns. Now we need the next generation application framework to make writing Data Smart Applications easier. We are starting to see this evolution. Salesforce Einstein is helping applications in Salesforce Cloud become smarter, but similar functionality is yet to be available in open source.
The Internet of Things is rapidly expanding and the market size estimates are huge. IDC estimates global IT spending on IoT-related items will reach $1.29 trillion by 2020. Edge intelligence has the potential to deliver insights and predictions where it is needed most, at a faster speed, without requiring a persistent network connection. What is needed is to deliver predictions at the edge, but predictive models need not be created at the edge. Today, model training at the edge is painfully slow and we can create better models faster in the data center. What is needed is to deliver these models to the edge where they can provide predictions even while being disconnected from the data center.


