Cloudera Gives a Peek at Future ML Platform

Cloudera continued its evolution away from Hadoop today by announcing a technical preview for Cloudera Machine Learning, its new data science and data engineering platform that’s based on Kubernetes, which enables it to run in the cloud and on premise.
Hadoop’s influence has waned, in many respects, in direct proportion to the rise of public cloud platforms. Instead of taking the time to build and manage Hadoop clusters to store big data and run analytics on them, companies are turning to cloud providers like Amazon Web Services, who can offer cheap object storage and scalable compute resources.
This changing market dynamic has helped to drive 46% year-over-year revenue growth for AWS, and Microsoft Azure and Google Compute Platform are growing even faster. In many ways, the October merger of Cloudera and Hortonworks was a response to this dynamic. But Cloudera thinks it has an advantage that the cloud vendors can’t touch: the capability to support multi-cloud and on-premise computing in a hybrid manner.
That’s the background behind today’s announcement of Cloudera Machine Learning, which will combine data engineering and data science capabilities in a cloud-friendly package. As the company points out, the new offering will work “on any data, anywhere.”
Thanks to Kubernetes, Cloudera Machine Learning can run on Hadoop clusters that companies have installed in their own data centers. But it can also run on public and private clouds – provided they offer Kubernetes (which most do). The new offering supports data stored in HDFS and cloud object stores offered by AWS, Azure, and GCP.
Cloudera Machine Learning is separate from its Hadoop distribution, CDH, but it can utilize the services of CDH in an on-premise or cloud environment, Cloudera says.
“Cloudera Machine Learning is self-contained and manages its own distributed compute, natively running workloads – including but not limited to Apache Spark – in containers on Kubernetes,” the company says. “It can also connect to an existing CDH cluster, on-premises or in the cloud, to leverage its distributed compute (e.g. Spark-on-YARN, Impala), data, or Shared Data Experience (SDX) metadata (Kerberos, HMS, Sentry, Navigator) for full enterprise security, governance, and management.”
As Hilary Mason, the general manager of machine learning for Cloudera, recently told Datanami, Cloudera is executing on a strategy to build a new platform around data science and machine learning. The strategy began with Cloudera Data Science Workbench (CDSW), which is designed to help data scientists and data engineers collaborate on the building of machine learning models. With Cloudera Machine Learning, the company gains more capability to push the learning and inference out to clusters besides Hadoop.
“We really see this as building a platform on a platform,” Mason said in an interview last week. “Specifically we see a suite of products in data science [and] machine learning…that are cloud native that are based on Kubernetes and container-based capabilities that allow for data engineering and data ingest, and that allows for data science modeling and machine learning model management.


