Beyond Hadoop

There was a time when Hadoop was synonymous with Big Data, but that time, it appears, is nearing an end. Sri Ambati, the CEO and co-founder of H2O.ai, famously said earlier this year that ‘Hadoop is dead,’ and his is not a lone voice, with Gartner claiming that many organizations are re-examining its role. In Gartner‘s hype cycle earlier this year, the advisory giant even declared Hadoop distributions ‘obsolete’ due to complexity, and stated that ‘the questionable usefulness of the entire Hadoop stack is causing many organizations to reconsider its role in their information infrastructure.’
The reasons for this are complex and the way forward shrouded by would-be giants who want to take its place. Firstly, a potted history. A large corporation 50/60 years ago would buy an IBM 360 Mainframe. This had a redundant power supply, it was properly engineered, had high quality components, and software that played nicely with it. Over time, as you had to deal with more and more transactions, you’d get the next model up, and then the next model, and so forth. Eventually, we reached a point where the largest organizations were – and they still are – operating a number of large data centers, much of which are only there in case something goes wrong. They are good because they are built to last and they rarely fail. They are, however, extremely expensive.
The alternative was developed by Google, who needed a tremendous amount of computing power but in the early days lacked the funds to get it. Legend has it, they would drive the streets of Silicon Valley looking for systems that had been thrown out onto the streets. There were obvious issues with this, so they came up with two technologies. The Google File System, which was based on a white paper released in 2002, and Map Reduce. The Google File System became the Hadoop Distributed File System, and Map Reduce is still Map Reduce. These two technologies are the core of Hadoop.
At the recent Big Data & Analytics Innovation Summit in Sydney, Mike Seddon, Senior Data Engineer at leading integrated energy company AGL, described his issues with Hadoop. According to Seddon, there are three myths around Hadoop that have been perpetuated over the years and contributed significantly to its growth. Firstly, that there is an avalanche of unstructured data being thrown at every organization. There is not. Most data in your organization came from a relational database, argues Seddon. If you want to run video through your Hadoop cluster, the first thing you’ve got to do, and the first thing you are paying machine learning guys to do, is to restructure your data. Essentially, you might have thrown away the structure in some cases, but it is still structured data, and there is no point dumping data on there and hoping that you’re going to be able to sort it out later because it likely won’t happen.
The second myth Seddon notes is that data types don’t matter. Hadoop distributors have somehow managed to convince people they can be left, but they are something you have to get right up front. If not, you’re going to get queries that don’t behave as you’d expect them to, for example if your customer IDs are stored as a string and an integer, they will not behave properly. Analysts shouldn’t have to know this, they should just have data that works. It is also really expensive computationally to calculate types again, back from CSV or whatever you’ve got.
Lastly, Seddon takes issue with the whole premise of Hadoop – that you should move your compute to your data. As it turns out, this doesn’t really matter. Data locality is important, which is why RAM is important, but once you get to a certain sized cluster, the probability of having compute available next to the data drops to a very low probability. As long as you can ship your data to where you need it quickly enough, it doesn’t really matter where it’s stored.
The main problem we have, and that Hadoop has, is that Google built this stuff in 1997.


