Why So Many Data Science Projects Fail to Deliver

4 min read

More and more companies are embracing data science as a function and a capability. But many of them have not been able to consistently derive business value from their investments in big data, artificial intelligence, and machine learning.1 Moreover, evidence suggests that the gap is widening between organizations successfully gaining value from data science and those struggling to do so.2

To better understand the mistakes that companies make when implementing profitable data science projects, and discover how to avoid them, we conducted in-depth studies of the data science activities in three of India’s top 10 private-sector banks with well-established analytics departments. We identified five common mistakes, as exemplified by the following cases we encountered, and below we suggest corresponding solutions to address them.

Hiren, a recently hired data scientist in one of the banks we studied, is the kind of analytics wizard that organizations covet.3 He is especially taken with the k-nearest neighbors algorithm, which is useful for identifying and classifying clusters of data. “I have applied k-nearest neighbors to several simulated data sets during my studies,” he told us, “and I can’t wait to apply it to the real data soon.”

Hiren did exactly that a few months later, when he used the k-nearest neighbors algorithm to identify especially profitable industry segments within the bank’s portfolio of business checking accounts. His recommendation to the business checking accounts team: Target two of the portfolio’s 33 industry segments.

This conclusion underwhelmed the business team members. They already knew about these segments and were able to ascertain segment profitability with simple back-of-the-envelope calculations. Using the k-nearest neighbors algorithm for this task was like using a guided missile when a pellet gun would have sufficed.

In this case and some others we examined in all three banks, the failure to achieve business value resulted from an infatuation with data science solutions. This failure can play out in several ways. In Hiren’s case, the problem did not require such an elaborate solution. In other situations, we saw the successful use of a data science solution in one arena become the justification for its use in another arena in which it wasn’t as appropriate or effective. In short, this mistake does not arise from the technical execution of the analytical technique; it arises from its misapplication.

After Hiren developed a deeper understanding of the business, he returned to the team with a new recommendation: Again, he proposed using the k-nearest neighbors algorithm, but this time at the customer level instead of the industry level. This proved to be a much better fit, and it resulted in new insights that allowed the team to target as-yet untapped customer segments. The same algorithm in a more appropriate context offered a much greater potential for realizing business value.

It’s not exactly rocket science to observe that analytical solutions are likely to work best when they are developed and applied in a way that is sensitive to the business context. But we found that data science does seem like rocket science to many managers. Dazzled by the high-tech aura of analytics, they can lose sight of context. This was more likely, we discovered, when managers saw a solution work well elsewhere, or when the solution was accompanied by an intriguing label, such as “AI” or “machine learning.” Data scientists, who were typically focused on the analytical methods, often could not or, at any rate, did not provide a more holistic perspective.

To combat this problem, senior managers at the banks in our study often turned to training. At one bank, data science recruits were required to take product training courses taught by domain experts alongside product relationship manager trainees.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at sloanreview.mit.edu →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.