Unifying Big Data And Machine Learning, Cisco Style

3 min read
Curated from nextplatform.com →

It doesn’t take a machine learning algorithm to predict that server makers are trying to cash in on the machine learning revolution at the major nexus points on the global Internet. Many server makers rose to satisfy the unique demands of the initial dot-com buildout back in the 1990s, and a new crop of vendors as well as some incumbents are trying to engineer some differentiation into their platforms to appeal to the machine learning crowd.

This is particularly true for servers that are used to train neural nets, which require lots of very beefy GPU accelerators, almost universally those from Nvidia, as well as a few hefty CPUs, lots of main memory and usually fast networking, too. Cramming this plus enough storage to be useful into a single node that is then clustered to scale out performance in a parallel fashion (like transitional HPC workloads) is a challenge. But the opportunity is large enough – and profitable enough – that after a bunch of customers started asking for a server that plugged into its Unified Computing System framework, could be managed by the UCS Manager stack, and integrate into the UCS network fabric. This means that UCS customers that want a single vendor and management scheme for their infrastructure don’t have to go outside of the UCS fold to do it.

They want one more thing, Todd Brannon, senior director of data center marketing at Cisco, tells The Next Platform, and that is to unify the platforms that are doing more traditional data analytics, based on tools such as Hadoop and Spark, with those that are doing machine learning training. This way, the iron can be used for multiple purposes while the data needed for both workloads is located on the same clusters. Moving data between a Hadoop cluster to a TensorFlow machine learning cluster is such a hassle that it makes economic as well as technical sense to put a lot of data storage on a machine learning node and have some of those GPUs go dark silicon some of the time. Machine learning algorithms work best with more and more data, so moving data between clusters becomes an inhibiting factor in trying to increase the accuracy of machine learning algorithms.

“The UCS installed base, particularly government agencies and enterprise customers getting started with machine learning, wants a UCS machine learning system,” says Brannon. “But there is more going on that that. Datacenters and data sources as increasingly distributed, even out to the edge, and data scientists are pushing the IT department to the bleeding edge. And making it even more complex, the predictability of applications is going away.”

Cisco doesn’t control its own big data or machine learning stacks, but it is relying on the tools from Cloudera and Hortonworks, which mash up the HDFS distributed file system from Hadoop with various machine learning frameworks. Cisco is also endorsing the Kubeflow containerized variant of Google’s TensorFlow machine learning framework, which as the name suggests packages up TensorFlow in Kubernetes containers for it to be deployed on clusters.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at nextplatform.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.