Looking for AI Success? It’s All About the Data

3 min read
Curated from enterpriseai.news →

As the number of industries integrating artificial intelligence (AI) into their operations continues to grow, more organizations are scrutinizing their AI system design workflows, including the roles of modeling and data. Within the workflows, organizations are finding and confirming that starting with good data plays the largest role in producing accurate insights.

That is because when the data is fed into a model, it shapes how the model analyzes, learns, and arrives at its decisions. If that model is forced to analyze substandard data, its insights will be substandard. Conversely, if the model is fed the most accurate and useful data available, its insights will be useful.

So how does one ensure the data fed into an AI model is the best available?

Here are three tips that show how to use a data-centric approach to improve the effectiveness of AI-built systems.

Tip 1: What to Do When There is Not Enough Data

Many organizations implementing AI systems often start with the question, “Do I have enough data to build a successful model?”

One scenario where this is commonly asked is when developing a predictive maintenance application for detecting critical failures. Such failures are often destructive, costly and rare, so gathering enough failure data needed to accurately train an AI model capable of detecting real-world equipment failures can be a difficult task.

Fortunately, several data simulation techniques can be used to generate accurate, realistic input data that can be used for such training.

The first method is to generate realistic synthetic data using a digital twin. Realistic digital twins can be made using methods such as model-based design, where all the components of a large physical system, such as an autonomous car or wind turbine, are combined into a single model. This allows the system to be simulated in multiple scenarios.

Traditionally, model-based design is used to design and simulate systems that have not included AI. When using it with AI, however, it solves several challengers. It allows engineers to run thousands of simulations covering all cases a system will operate in. That generated data can then be used for training an AI model.

In addition, the trained model can then be included back in the original system model, which ultimately enables it to be tested on how it performs under operation in various simulated scenarios. For example, creating a simulation model of a pump that can generate faults and healthy data for building a predictive maintenance application. The data created can then be used to train a multi-class classifier to detect different combinations of faults.

Another type of digital twins is one where an AI model is created based on historical input and output data. For example, an industrial tools and equipment manufacturer is using the digital twin approach for obtaining the data it needs, and then is using the data for building predictive maintenance models that will ultimately be deployed to more than hundreds of thousands air compressors manufactured and operated in their global manufacturing plants.

Synthetic data can also be generated using deep learning. If you don’t have access to a system-level model, techniques such as generative adversarial networks (GANs) can be used.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at enterpriseai.news →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.