Why successful AI needs fast data access

Sponsored Spending on artificial intelligence systems will grow from $37.5bn worldwide in 2019 to $97.9bn in 2023, according to IDC. And use cases cover everything from ERP, manufacturing software and content management to automated customer service agents, threat intelligence and fraud investigation.
The number of smaller data centres in operation is also likely to increase, IDC analysts say. This will enable enterprises to navigate local data protection regulation and host AI workloads in edge computing environments to optimise application performance.
However, respondents to Gartner’s 2019 CIO Agenda survey indicate that data scope and quality continue to be major challenges to AI adoption. Also, they expressed concerns that organisations cannot store, process and analyse sufficiently large quantities of good quality information to build successful use cases. Gartner notes that many organisations struggle to scale their AI pilot projects into enterprise-wide production implementations, which inevitably limits the technology’s business value.
The process intensive nature of AI means that access to high performance, scalable IT infrastructure is critical to its success in real world deployments. This entails the ability to analyse large data sets quickly and accurately enough to produce actionable insight or provide the basis for automated responses within particular applications or services.
For example, chatbots need to ingest and process language and provide a sufficiently fast and accurate response to the human participant to avoid conversational disruption caused by undue latency. Speed is also the essence for fraud analysis and investigation applications, which must feed back risk scores almost immediately to enable financial institutions to conducting safe transactions in real time.
But building, testing, optimising, training, inferencing and maintaining the accuracy of the AI models depends on having large quantities of data to train data in the first place. This in turn is sourced from a diverse array of high quality and dynamic data inputs. Gathering all of that information and storing it ready to feed into adjacent processing engines is a big undertaking in itself. And to do that at optimum speed requires an AI data pipeline that provides ready access to reliable and well-structured data sets for analytics purposes. However, most of the underlying legacy network and storage architecture is ill-equipped to provide this.
The storage architecture also has to handle extreme fluctuations in the volume and type of data being processed and analysed by the AI workload. This can range from petabytes at the ingestion stage to gigabytes of structured and semi-structured during training and turning into just kilobytes when data is subsumed into the final trained model.
Workloads vary in terms of read write operations too, beginning with 100 per cent writes at the ingest stage which progress to 50/50 read write during preparation before they shift to 100 per cent reads during training and inference. And high throughput and low latency is a must to support fast execution of compute intensive training and inference processes, irrespective of the IO demands or the type of data being processed.
Due to the performance limitations of spinning disk technology, simply adding more hard disk drives (HDDs) to existing storage infrastructure is unlikely to deliver the foundation needed to build an AI data pipeline. But simply upgrading to SATA-based flash memory solid-state drives (SSDs) probably won’t solve the problem either, even when latency is improved using faster interfaces based on the non-volatile memory express (NVMe) protocol.
The improved performance and low latencies offered by Intel® Optane™; SSDs can reduce the time it takes to build and train those AI models. The longer it takes for different system components to get the data, the slower the whole computing process goes.


