IoT boom will change how data is analysed

As more devices come online, each generating data, the way information is analysed and used to facilitate machine learning will have to change.
Data is key to improving the accuracy and predictions of machine learning and artificial intelligence (AI) systems–the more images of apple and oranges it is fed, the better it will be at distinguishing the two.
The “smartness” of a machine was data-driven, said Goh Eng Lim, vice president and CTO of high performance computing and AI, Hewlett-Packard Enterprise (HPE), who was speaking to media at the vendor’s Reimagine Summit in Singapore.
However, he stressed, companies should not be focusing only on data retention as establishing volume alone was not adequate. Data also needed to be curated, labelled, and federated.
Goh noted that, too often, organisations operated data in silos, with the HR department generating data that did not integrate with data sitting with the sales team. In order to make better predictions, machine learning systems needed to be able to seamlessly pull data across the company.
They also needed to work on data that was properly labelled and curated to ensure decisions were made on accurate, quality data, he said.
Asked if the anticipated boom of IoT then would introduce further complications, he acknowledged the likelihood, adding that it could indeed worsen the situation.
The number of connected devices was estimated to have outpaced the global population last year at 8.4 billion, and would continue to climb to 20.4 billion by 2020, predicted Gartner.
For one, Goh noted, it would not be feasible to push all the data generated by every connected device back to the data centre to be analysed.
“The network can’t keep up,” he said. “Therefore, you’ll need the IoT device to be smarter so it can make smart decisions at the edge, for instance, only sending back information that’s needed back to the network.”
The IoT device could ascertain if the data was of high quality and should be pushed back to the network to facilitate deep learning, or to process the learning at the edge and send back only the knowledge–rather than pure data.


