What wasting data engineering talent really costs you

4 min read
Curated from fivetran.com →

Building a modern data stack from scratch is complex. Using traditional APIs and legacy data integration tools is tedious and expensive. Collecting, cleaning, sorting and analyzing data on an enterprise scale across dozens of internal and third-party sources takes an entire team of data professionals working around the clock, constantly maintaining and updating these complex pipelines. 

But a functioning data stack is a critical need. Data drives business, giving companies real-time insight into market conditions, dynamic supply chains and unpredictable customer expectations so they can make quick, accurate decisions in the moment.

Organizations that are able to radically transform their approach to integrating data by embracing an automated, efficient approach will save valuable engineering time and create critical insights that will result in major business impact in a competitive market.

Unfortunately, many companies are letting data infrastructure inefficiencies sap their engineering resources. According to a new report by Wakefield Research, the average data engineer spends 44 percent of their time maintaining data pipelines – which costs $520,000 per year.

It’s no surprise that, according to the same report from Wakefield, nearly three out of four data engineers feel that their team’s time and talent are being wasted by having to manage these data pipelines manually.

Can you blame them? Refocusing their responsibilities away from manual maintenance to data analytics would go a long way toward improving morale and retention.

So, why ask a highly-trained data engineer to spend their time twisting knobs and pulling levers? Shouldn’t they be, you know, actually working with data?

Unfortunately, legacy data ingestion tools are ill suited for today’s enterprise data needs. First, it can take a data engineer weeks or even months to build a single connector. Multiply that by the number of connectors needed — often dozens or even hundreds within a single company, depending on the data environment — and you’re talking about a multi-year project just to get everything fed into your data warehouse.

Once configured, legacy tools may require engineers to manually set up hundreds or thousands of schemas and tables one at a time to ensure the data is landed in the desired format. There’s also the issue of post-build maintenance: what happens after you’ve created a custom data pipeline. Your new connector can pull your data from a given source, but now you have to continually maintain the code and system infrastructure. Every time your source issues an update to its API or data structures, you’ll need to accommodate new API endpoints and supported fields, resulting in updating your pipeline extract scripts. And that takes resources away from the data engineering team.

An alternative to the DIY route is a managed data integration service to build, manage and update data pipelines. This allows companies to automate the process of integrating data into their data warehouse, eliminating ongoing maintenance and all the hassles of updating pipelines when APIs or data structures change, and giving data engineers valuable time back. 

“My job is to make sure everyone has the information they need to create efficiencies, deliver superior experiences for customers and optimize their efforts,” said Daniel Deng, a data architect at mortgage broker Lendi. “I need to make it as easy as possible to work with and analyze the data from various data sources. The quicker we get the data, the quicker we get the insights that the business needs for making a decision. Then the quicker our business operates and evolves.”

The impact of automating the data pipeline was even more significant for Mark Sussman, head of data analytics for ItsaCheckmate. He was able to stand up an entire data analytics framework in just a few months with minimal engineering help. “I can honestly say that we are now a data-driven company,” Mark said.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at fivetran.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.