The Drive for Big Data in the Cloud

Alan Clark of OpenStack describes how the OpenStack open source IaaS provides a solution for big data in the cloud and how it can offer an attractive Hadoop deployment strategy.
We’ve all seen statistics demonstrating the amount of data being generated and gathered daily with average amounts in the Petabyte and Exabyte range. Processing such large amounts of data is what Big Data is built to do. It’s no wonder that the industry has become so large with the amount of data and the potential business around it.
While there are multiple solution sets targeting big data analytics, we have traditionally focused on maximizing dedicated hardware and large processing power. As of recent there is a fast growing convergence of Big Data and cloud, particularly where the data sets are unstructured with simple data models — an area of specific focus for the Apache Hadoop technology.
The mission of OpenStack is to produce a ubiquitous open source cloud platform that will meet the needs of public and private cloud providers. OpenStack is about Infrastructure-as-a-Service (IaaS), an open source project, community and ecosystem that has dramatically grown to over the past five years. Today the project hosts over 27,000 individual members, 2,000 contributors and 500 supporting companies. The number of components has grown from the original two to over 25 today.
These statistics convey the significance of open source and the transformation of the cloud over time. Users within the open source project can see the power of open source and appreciate its ability to embrace new ideas and market needs including the convergence of Big Data on cloud.
While this transformation is taking place, it is important to note that OpenStack is not recreating Hadoop or any of the other Big Data technologies. The OpenStack effort was created to facilitate the care and management of Big Data within the IaaS infrastructure; sometimes called Analytics-as-a-Service. The technology effort within OpenStack is code named Sahara and was created to provision, launch and manage Hadoop clusters on top of OpenStack, making it simple to deploy and manage Big Data infrastructure and tools.


