Stop Putting the AI Cart Before the Data Horse

4 min read
Curated from linkedin.com →

Flush with Big Data and an accelerated way to capitalize on it (AI), many large enterprises are making a classic mistake. They’re assuming that their data is good enough to leverage as an asset. It is not. 

In a recent article in The Wall Street Journal, IBM executive Arvind Krishna said that data-related challenges were a top reason for IBM clients halting or cancelling AI projects. Further, he said that about 80% of the work with an AI project is collecting and preparing data. This number jibes with long-standing industry estimates for the amount of time data scientists and analysts typically spend on finding and preparing data before they get to the actual science/analytic part. 

While the admission may have been startling, the problem is far from new. I started working in the AI industry back in the 1980s and the same was true then as now with respect to any algorithms (AI or other): “Garbage In, Garbage Out.” Without great data, the math isn’t very useful.

In a June 18, 2019 article by my colleague/co-founder Ihab Ilyas (of the University of Waterloo) and O’Reilly Publishing’s Ben Lorica, the authors noted that a recent O’Reilly survey on the adoption of AI by enterprises “found that those with mature AI practices (as measured by how long they’ve had models in production) cited ‘Lack of data or data quality issues’ as the main bottleneck holding back further adoption of AI technologies.” Meanwhile, worldwide spending on artificial intelligence (AI) systems is predicted to hit $35.8 billion in 2019–a 44% increase over that spent in 2018, says IDC.

Translation: The AI cart is getting loaded up pretty full: we’ve gotta make sure that the enterprise data horse has the legs–and is actually in front of the cart 

If traditional methods of data preparation don’t work for AI, they really won’t work for today’s data scientists and analytically empowered front-line workers. The number of data scientists has exploded in the last decade, all of them wanting to do advanced modeling and analytics, many of which also involve AI techniques. At the risk of mixing transportation metaphors: given the high cost of data scientists–with entry-level salaries of ~$100k–ramping up these teams before having your data in order is like buying a Ferrari then putting raw crude oil in the gas tank. Enterprise data needs to be “refined” to make your data science, AI and next-gen analytics purr like a kitten.

If you’re betting the future of your company on the predictive-analytics power provided by AI and your analytically empowered workforce, you need great data. Period.

By “great data,” I mean data that’s been rationally organized, unified and curated; is continuously updated; and is readily accessible by everyone who needs it, including data scientists and their models as well as the average person on the front line of your business. This is not so easy, however. The core systems we build reflect the idiosyncrasies of the time and context in which they are built: the more systems, the more idiosyncrasies. We’re coming off of 50+ years of building systems in the enterprise–that’s five decades of idiosyncrasy. Most large companies have thousands of systems that generate data that could be an asset, if it weren’t for the data’s radical heterogeneity. 

As large companies change, their business processes and systems change MUCH more slowly, in effect creating a “data-consistency drag coefficient” that results in database decay. M&A, changing corporate strategies, and other facts of corporate life make this decay even worse. The quantity and rate of decay vastly outstrip humans’ ability to organize it for strategic uses like predictive analytics. 

Traditional data integration/unification methods (Data Warehousing, Data Lakes, Master Data Management) have not kept up with the pace of variety of data in the enterprise. This is because humans are still in the critical path, at multiple levels.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at linkedin.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.