Why event stream processing is leading the new ‘big data’ era

3 min read

From the downhill of key technologies to new innovative solutions on the plate: Roberto Bentivoglio takes us swimming in the stream, with a technical explanation on event stream processing architectures as a future-proof approach to take full control over your data strategy

Big data is probably one of the most misused words of the last decade. It was widely promoted, discussed and spread by business managers, technical experts, and experienced academics. Slogans like ‘data is the new oil’ were widely accepted as unquestionable truth.  But with this hype, different ideas and solutions have been considered, and then rejected. Event stream processing may provide the answer.

These beliefs helped push Hadoop technologies —an open source distributed processing framework that manages’ data processing and storage for big data. Its stack, formerly developed by “Yahoo!” and lately owned by the Apache Software Foundation, was recognised as ‘The’ big data solution.

Many companies started to offer commercial, enterprise-grade and supported versions of Hadoop, until it was eventually adopted across industries, and by both large, Fortune 500 companies,  and medium-sized companies.

The prospect of analysing huge amounts of data generated by heterogeneous sources in an attempt to boost competitiveness and profitability, provided and alluring reason to use Hadoop.

There was another factor; Hadoop was seen as providing a way to replace expensive legacy data warehouse installations, in the process, improving both performances and data availability while simultaneously reducing operational costs.

However, in recent years, a growing number of analysts who were focused on the big data market published articles suggesting that  Hadoop was not all it had been cracked up to be.

Their critique can be summarised as follows:

● The deployment model is moving from on-premises solutions to hybrid, full and multi-cloud architectures. Hadoop technology isn’t made to be completely cloud-ready. Furthermore, cloud vendors had been selling cheaper and easy to manage and use solutions for years;

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

Machine Learning technologies and platforms are quickly reaching a level of maturity. The Hadoop stack was not designed around Machine Learning concepts even if support in this area had grown over the years

● The advanced and real-time analytics market is rapidly increasing, but the Hadoop stack doesn’t seem to be the best fit to implement those innovative forms of of analytics.

In short, analysts began to fear that Hadoop was no longer an innovative technology.  To solve future challenges, something different was required.

Our own experience chimed with this, we found that solutions based on the Hadoop stack were hard and expensive to develop and maintain. Furthermore, it was hard to recruit professionals with the right skills and any proven experience.

Consequently,  moving systems from proof of concent and prototypes statuses to real productiveness seemed like an unreachable finish line.

There was another issue. Hadoop vendors tended to focus on the data lake. While big and complex organisations need a unique de-normalised data repository, the projects designed to feed data lakes tend to take before reaching maturity. Most of these initiatives turned out to be expensive, from both an economic and project governance point of view.

These complex repositories, or data lakes, are filled with historical data referring.  If you are lucky, they can provide a series of snapshots of the last closing day.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at information-age.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.