Streamlining Data Science and Analytics Workflows for Maximum ROI

Initially, the relationship between data science and analytics is causal. Accurate analytics is one of the outputs of data science; the more effectual the latter, the more so the former. Nonetheless, there’s a marked tendency for most organizations to isolate each of these functions, which inevitably circumscribes their value.
“A lot of times, business intelligence, analytics [and] reporting all live in one silo, and data science lives in a totally different silo,”LookerChief Data Evangelist Daniel Mintz observes. “Each of them builds out their own totally independent, parallel workflow. A lot of those workflows, they branch off and do their own thing, at a certain point. But their beginning steps are all the same.”
Organizations can increase the business value of both data science and analytics by streamlining their fundamental underpinnings ofquality, accurate datain a format meaningful to business end users. Those successful in this endeavor not only decrease the impact of silo culture in data-centric firms, but also improve the speed, usability, and ROI of predictive analytics—and the data scientists empowering it.
Despite the multitude of tasks associated with the data science position, its basic workflow (in terms of analytics) is readily codified into three steps. The first is data preparation or data wrangling; where the data scientist starts with raw data and “just tries to make sense of it before they’re doing anything real with it,” Mintz explains.
“Then there’s the actual model building when they’re building a machine learning model. Assuming they find something valuable, there’s getting that insight back into the hands of the people who can use it to make the business run better.”
Typically, data scientists approach building a new analytics solution for a specific business problem by accessing raw data from what might be a plethora of sources. Next, they engage in a lengthy process to prepare the data for consumption. “So much time and energy goes into that,” says Mintz. “You look at the surveys of data scientists and they say 70-80% of my time goes to data cleaning.”
A much more efficient alternative is for data scientists to get their initial data from platforms designed for universal business access to data—regardless of where data physically reside.
The benefits of this method are multifaceted. Data scientists can decrease the complexity of the preparation process by using data already deployed for specific business purposes, which “schematically syncs up the workflow and…also means they’re more likely to come up with something useful for the business,” Mintz remarks.
Moreover, accessing highly distributed data from a centralized platform also decreases the time spent on data preparation, since it’s streamlined according to both uses: those for the business and for data science. The result is efficient data wrangling, enabling organizations to “unify that, leverage everybody’s expertise, do that once rather than twice, and then go and do their own thing with that data once it’s clean,” Mintz says.
Perhaps the most important aspect of a data scientist’s job is translating insight from advanced analytics into tangible business value.

