10 Key Big Data Trends That Drove 2017

2017 has come and (almost) gone. It was a memorable year, to be sure, with plenty of drama and unexpected happenings in terms of the technology, the players, and the application of big data and data science. As we gear up for 2018, we think it’s worth taking some time to ponder about what happened in 2017 and put things in some kind of order.
Here are 10 of the biggest takeaways for the big data year that was 2017.
2016 closed with a keen interest in AI, and that momentum in AI surged in 2017, thanks to emerging deep learning technology and techniques that provide better and faster results for some machine learning tasks. Teradata, for instance, found that 80% of enterprises are already investing in AI, which backed similar findings from IDC.
Nevertheless, the same old challenges that kept big data off Easy Street also emerged to cool some of the heat emanating from AI. Over the summer, Databricks‘ CEO, Ali Ghodsi, warned about “AI’s 1% problem.” “There are only about five companies who are truly conducting AI today,” Ghodsi said.
The sudden re-emergence of AI also re-kindled an old debate about the meaning of the phrase, as well as the differences between AI and deep learning and machine learning. While some purists may consider AI to refer to a human-like replacement, the software community generally has a much broader view of AI today. On the hype meter, AI has replaced big data as the new “it” term.
By year’s end, we noticed a dip in excitement around neural networks, which really don’t replicate the functioning of the human brain. We also heard how “cracking the brain code” is our best chance at achieving “true” AI, which is something backed by Geoffrey Minton, the father of deep learning.
One of the biggest big data stories of 2017 was the shine coming off Apache Hadoop. We detailed the struggles that some customers have had with trying to get Hadoop software set up and running. Common complaints about Hadoop that we relayed in our March story include: too much complexity, incompatible components, poor support for small files, poor support for non-batch workloads, and just an overall limited usefulness for non-developers.
There was some pushback from technologists. Sure, Hadoop is complicated, they said. It’s not perfect. It’s not “set it and forget it.” But until a better general-purpose distributed computing platform comes along, Hadoop is the best option for organizations that need to store and process very large data sets, they said.
There are different ways of processing this. One view holds that Hadoop’s time in the sun has passed. Another view holds that its difficulties are temporary setbacks that all emerging technologies face before eventually finding its footing as a core component of the enterprise stack moving forward. Interestingly, there seemed to be very little interest in Hadoop 3.0, which we first wrote about in 2016 and covered again in mid-2017. With Hadoop 3.0 nearing GA this month, interest seems to have perked up a bit — but it’s nothing like the frenzy that accompanied the release of Hadoop 2.0 four years ago.
We also saw Hadoop removed from the names of two prominent industry conferences, including Cloudera and O’Reilly’s Strata + Hadoop World (now called Strata Data Conference) and Hortonworks’ Hadoop Summit (now called DataWorks Summit).
Graph databases continue to gather momentum in the industry, thanks to their status as the preferred technology for a set of specific use cases revolving around connected data. In January 2017, we saw the emergence of JanusGraph to continue open source TitanDB development, while in February we described how Qualia uses Neo4j to allow ad tech firms to quickly detect trends in user behavior.
In April we saw IBM, which debuted a Titan-based graph database service in 2016, touting the advantages of the graph approach with its “State of Graph Databases” report.


