Are you paying enough attention to data engineering tasks?

2 min read
Curated from blogs.sas.com →

Data science is hot. Being a data scientist has been described as one of the world’s sexiest careers – though that may say more about the people coining the description than data science – and data scientists are both in high demand and increasingly well-paid. It is, in fact, highly lucrative to be able to use statistical techniques to manipulate data and generate insights. If you can present and explain those insights to others, you are very definitely able to name your own price.

Increasingly, however, those in the data business are noticing another term emerging. Data engineering is mentioned more and more often as an essential adjunct to data science. It may not be quite as sexy, but it is often given equal weight in terms of importance. But what does it mean?

There is no absolutely agreed single definition of data engineering, perhaps because it is such a new discipline, but it seems clear that it is about data management. Data engineers, in other words, deliver the data to data scientists in a form and way that allows the data scientists to generate insights. They are, for example, responsible for data collection, storage and cleaning. They are also, in many organisations, the people chiefly responsible for assets such as data lakes and analytics platforms.

A good platform now offers far more data processing power than has ever been possible before. This means that analytics processes can draw on more data – but only if it is available and stored in a form in which it can be used.

Data lakes need ongoing engineering and management to prevent them from becoming swamps, which can happen extremely quickly. Reliable data set processes and other engineering work ensure that data is available and in a usable form. It is much harder to drain a swamp than flood a reservoir!

Data engineers are the primary people responsible for data quality in the organisation, and at both the macro (data lake) and micro (individual data) levels.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at blogs.sas.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.