5 Reasons Your Data Lake Is Failing

3 min read
Curated from cbronline.com →

The reality is that data lakes are failing to support the time-to-market requirements new analytics-driven innovation requires, and it is safe to say that in many companies, data lakes are widely perceived to be expensive and ineffective. So why is it happening?

In this article, we look at some of the common culprits turning data lakes into data swamps, at the same time delivering some advice based on experience to help companies from experiencing data lake disasters.

Many data lake programmes are suffering from lack of real experience with entire teams or departments exploring and testing Hadoop technologies for the first time. They are often disoriented due to the very different paradigms and approaches, and by the novelty of these tools and frameworks, which have very little in common with well-known traditional technology stacks.

As a result these programmes are very slow, with the implementation becoming complex and difficult, the business objectives becoming quickly obsolete, and the original excitement slowly fading away. At this point, many stakeholders begin to wonder whether a big data solution is ever going to take off and achieve the original goals.

In the endeavour to achieve data lake success, working with thought leaders is key. As with all emerging technologies, it is important to start with confidence and get in touch with the experts. These are the pioneers that have already accumulated a good amount of knowledge and expertise, who have already failed in their first attempts and learned from those lessons to finally identify best practices and successful solutions.

Simply put, most data lakes today suffer from poor design and implementation. The shortage of software engineering talent combined with the lack of Hadoop experience is definitely one of the root causes. Such shortage leads in fact to hiring inexperienced data engineers. At the same time, it’s also hard to identify the right skills required for a good data engineer. The ability to master technologies like Spark, Kafka, HBase, is absolutely positive but could not be enough when it comes to building a complex and well-engineered data lake and delivering a production grade platform.

A lack of engineering skills often leads to data lake implementations with poor architectural design, poor integration, poor scalability and poor testability – all of which can lead to a level of instability that unfortunately only a full rewrite can fix.

Hiring solid software engineers is the best way to minimise this risk or quickly recover from it. So they should invest more in talented software engineers and train them on Hadoop technologies if necessary. It is much easier and faster to skill up a bright software engineer than turn a Hadoop “certified” professional (with no software engineering) into a good software engineer. Companies can also largely reduce the need for Hadoop engineers and accelerate their programmes by investing in a data lake management platform.

Particularly in the initial phase, the typical separation between IT and business can be a big obstacle. Data scientists tend to fall into business silos, while data engineers fall into IT silos. Yet a successful analytic solution relies completely on close collaboration between data scientists and data engineers.

Data scientists must use the tools made available by IT, while data engineers need to productionize and operationalize what is implemented by data scientists.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at cbronline.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.