5 Things to Keep in Mind When Using Data for Artificial Intelligence

3 min read
Curated from entrepreneur.com →

Data is one of the most important strategic assets for companies in the emerging data-driven and AI-powered economy. Data is needed to measure the efficiency of business strategies and draw insights from its operations but also to train machine learning algorithms. Getting data is not a problem for companies, the question is can they get the right kind of data and can that provide them with a much desired competitive advantage.

Many companies do not realize that they are sitting on a pile of bad or dirty data. This data contains a lot of missing fields, has wrong formatting, numerous duplicates, or is simply irrelevant information. IBM research estimated that the annual cost of bad data for the U.S. economy is a whopping $3.6 trillion. Still, many managers have certainty that they are sitting on a goldmine of data when in reality they have nothing valuable.

I interviewed Sergey Zelvenskiy, who is an experienced machine learning engineer over at ServiceChannel, where he automates facilities management processes using artificial intelligence. We talked about common misconceptions when it comes to the good/bad data dichotomy and what companies should be focusing on when building AI products.

As Zelvenskiy says, “The data that companies have may not necessarily be bad, it is just likely incomplete to solve the problem. There is a chicken and egg problem here. The original system is usually built to collect the data needed for human-driven solutions and moving it to an AI driven solution might require filling of the gaps. While a human can quickly assess these and fix the problem, the automated system needs automated ways to wrangle the data.”

Finding good data should start with a product itself. To get good data, companies should design products that provide the right incentive for the users to contribute their data. Good usability and user experience will encourage users to contribute valuable information.

You can always strive for the user-in-the-loop model, in which users have to give away their data in order to use the features of your product. This is precisely how Google and Facebook get tons of data in exchange for their services. Users are not even aware that they are giving away their data absolutely for free to power advanced machine learning algorithms and continually improve the software.

The best way to build a great product is by delivering iterative improvements while gathering the much-needed data. As Zelvenskiy says, “You can see this with the evolution of Amazon Alexa. The team behind it realized the difference between general speech recognition and the ability to recognize a simple set of predefined commands. While many other companies struggled with the adoption of general speech recognition and the capability to maintain the conversation, Alexa team focused on a simple set of commands and simple scripted dialogues.”

The Alexa team did it right by shipping a very simple solution at a low price and conquered the market. Focusing on the specific simple use case and perfecting it wins the end game.

Let’s take the company that wants to build a robot that will automatically put library books on the shelves. It has plenty of data about the actual book content, it knows the names of the authors and the year the book was published. But, in reality, this data is not sufficient for an automated arrangement of the books.

The robot can use the existing data only to find the proper shelf for the book.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at entrepreneur.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.