Big data tooling rolls with the changing seas of analytics

In the early days of big data that followed the invention of Hadoop at Yahoo, proponents emphasized its potential for replacing bulging enterprise data warehouses focused on business intelligence.
Open source Hadoop data tooling was posed as an alternative to existing systems seen as expensive and ill-suited for the ever-larger volumes of arriving data.
That emphasis has shifted over time to complementing existing data warehouses, and more. Hadoop applications have often come to be called data lakes, and the shifts just keep on coming.
Big data tooling has expanded far beyond mere data warehouses, according to Mike Matchett, an analyst and founder of the Small World Big Data consultancy.
“We are seeing increasing capabilities on the Hadoop and open source side to take over more and more of the corporation’s data and workloads, including BI,” Matchett said.
The original independent Hadoop distribution providers have had to be agile amid such major shifts.
A look at recent efforts of pioneer Hadoop vendors such as Cloudera, Hortonworks and MapR forms a backdrop as many look toward next week’s Strata Data Conference in New York. While big data tooling is still a prominent focus of an event that tends to highlight progress in big data, the conference’s emphasis – like that of original Hadoop boosters — has moved deeply into management of processes, and far beyond just Hadoop.
In August, Cloudera rolled out Workload XM management services for cloud-based analytics. Almost at the same time, the company made a hybrid Cloudera Data Warehouse and a Cloudera Altus Data Warehouse service generally available on both AWS and Azure clouds. The management services seek to bring visibility into diverse data workloads. Workload XM is also designed to help administrators provide dependable service-level agreements for self-service analytics applications, according to Anupam Singh, general manager, analytics, at Cloudera. Meanwhile, Singh said, the cloud warehouse offering enables encryption for data at rest or in motion, and provides a view into the lineage of data sets in analytic workloads. Such capabilities have grown in importance as GDPR and other data privacy initiatives have gained momentum. All these moves play to corporations’ needs to increase use of big data analytics, Singh said. “Customers don’t look at buzzwords like Hadoopand cloud. But they do want more business units to access the data,” he said. Individual business units want to spin up data warehouses on the cloud in no small part because doing so as a capital expenditure is beneficial, Singh said.
Cloud has been a clear preoccupation for Hadoop player Hortonworks, too. In June, the company expanded its Google Cloud presence with Google Cloud Storage support. Improving the management of real-time data analytics on cloud and on-premises has also been a goal. People get excited about AI as if it’s a wonderful glowing orb in the sky, like it’s magic. In August, Hortonworks sought to improve the handling of streaming data, launching Streams Messaging Manager (SMM) to provide administrators better views into Kafka messaging clusters that have become increasingly prevalent in big data pipelines.


