IBM Unwinds Tangled Data for Enterprise AI

4 min read
Curated from nextplatform.com →

These days, organizations are creating and storing massive amounts of data, and in theory this data can be used to drive business decisions through application development, particularly with new techniques such as machine learning. Data is arguably the most important asset, and it is also probably the most difficult thing to manage. Well, excepting people.

Data is tangled mess. It can be structured or unstructured, and it is increasingly scattered in different locations – in on-premises infrastructure, in a public cloud, on a mobile device. It is a challenge to move, thanks to the costs in everything from bandwidth to latency to infrastructure. It has a zillion different formats, sometimes chunks of data are missing, and usually it is unorganized and alarmingly often ungoverned. So for those organizations looking to leverage machine learning for research and application development, the challenges associated with managing and processing the data can be a hurdle.

“Every client is on journey towards AI, which means they have some level of sophistication around machine learning, some level of sophistication around analytics and, ultimately, the right data and information architecture to support those,” Rob Thomas, general manager of IBM Analytics, tells The Next Platform. “Think of those as kind of the building blocks for AI, which are often a deterrent or an obstacle to getting to AI if that is not in place. What we’ve identified as the biggest inhibitor is basically the integration – stitching things together, getting systems to talk to each other, getting the data architecture normalized. The biggest thing that holds up developers today is, ‘I can’t get the organization to get me the data I need, or they’ll only give me a subset, or I can only get access to the on-premises data, I can’t get access to the cloud data.’”

All of this falls close to IBM’s heart. The company has put down a big bet on AI, through its focus on cognitive computing and its efforts to drive that with its Watson technology. At the same time, Big Blue has ramped up efforts to build out its public and private cloud capabilities as it looks to better compete with Amazon Web Services and Microsoft Azure, and to make it easier for businesses and other organizations to collect and analyze the data they are creating. A recent example was the launch in the fall of the Integrated Analytics System, a platform to enable data scientists and developers to use advanced analytics with data regardless of its location – including private, public and hybrid clouds – and to move workloads among data stores. As we noted, it also uses machine learning techniques and data science tools to automate many of the tasks involved with data analytics. Soon after, the company unveiled its IBM Cloud Private, a software platform built on Kubernetes that enables enterprises to develop private clouds, embrace Docker containers, more easily move workloads between private and public clouds and accelerate application development.

This week, in the runup ahead of its Think 2018 conference, IBM is expanding its capabilities around data management, artificial intelligence, and the cloud. Among them is its new Cloud Private for Data, a platform designed to enable businesses to more quickly take in and analyze large amounts of data that is coming in from such areas as IoT sensors and mobile devices. It’s an application layer launched on the IBM Cloud Private platform that is based on Kubernetes and leveraging microservices. It is powered by a fast in-memory database that can ingest and analyze huge amounts of data. The database, built over the past two years, uses an open source Spark engine and the Apache Parquet data format, according to IBM’s Thomas.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at nextplatform.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.