Why Event Stream Processing Is Leading the New Big Data Era

3 min read
Curated from dzone.com →

Big Data is probably one of the most misused words of the last decade. It was widely promoted, discussed, and spread around by business managers, technical experts, and experienced academics. Slogans like “Data is the new oil” were widely accepted as unquestionable truth.

These beliefs pushed Hadoop technologies forward. Its stack, formerly developed by Yahoo! and now owned by the Apache Software Foundation, was recognized as “The” Big Data solution.

Many companies started to offer commercial, enterprise-grade and supported versions of Hadoop until it started to be experimented and adopted on by a large number of industries, ranging from medium-sized companies to Fortune 500.

The possibility to analyze huge amounts of data generated by heterogeneous sources, trying to boost the company’s competitiveness and profitability, was the key reason for the investments made on Hadoop.

Another important point was the idea to replace the expensive legacy data warehouse installations with Hadoop, trying to improve both performances and data availability while reducing operational costs at the same time.

However, during the last few years, a growing number of analysts focused on the Big Data market, started to publish articles, declaring an upcoming fall of the Hadoop world. Their mainly motivation behind such statements can be summarized as follows:

In a few words, the analysts started to declaring that Hadoop was not anymore an innovative technology and to solve the future challenges something different needed to be put on the plate.

On the opposite side, from a more empirical point of view, analyzing our personal past experiences, the solutions based on the Hadoop stack were proven to be really hard and expensive to be developed and maintained. Furthermore, the professionals having the right skills and any proven experience were not easy to be recruited.

As a result, a lot of adopters finally didn’t reach the maturity of its vertical solution developed on top of the technology. The consequence was that moving those systems from PoC and prototypes statuses to real productiveness seemed almost an unreachable finish line.

Those aren’t the unique key reasons for the recent disillusion around Hadoop technologies and in general on the “Big Data” movement. One of the main motivations can be identified on the proposition used by many Hadoop vendors, positioning concepts like Data Lake as central to data management.

While creating a unique, denormalized data repository is still a need for big and complex organizations at least to empower data governance and data lineage practices, the projects aiming to feed Data Lakes usually tend to last years in large enterprise before reaching the maturity. Most of those initiatives finally demonstrated to be really expensive both from an economic and project governance point of view.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at dzone.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.