Getting Data Scientists and Data Engineers on the Same Page

3 min read
Curated from datanami.com →

Like cats and dogs, data engineers and data scientists often seem like two incompatible species. Scientists love probabilities and experimentation, while engineers live for repeatability and efficiency. They have different responsibilities and dissimilar mindsets, but getting these two personas to work together is a critical step for any organization that wants to succeed with data.

Data scientists emerged as the rock stars of the 2010s thanks to their ability to use machine learning algorithms to detect small differences in big data sets and exploit them for business advantage. As data continued to grow bigger and more complex, the work became more specialized and data engineers emerged as a critical cog in the big data machine.

In today’s larger organizations, you will often find a mix of data scientists and data engineers working with data (and perhaps other related positions, such as the machine learning engineer, which blends characteristics of both). While engineers and scientists, ostensibly, have the same end goal–their organization’s successful exploitation of data–their paths to achieve that goal could not be more different.

One data scientist who has given the problem some thought is Max Boyd. Before joining Kaskada as its data science lead, Boyd worked in numerous Seattle, Washington-area startups where he often worked closely with data engineers. Boyd was dependent upon the data engineers to get the data he needed to build and train his machine learning models, but the relationships were sometimes as dreary as the Northwest’s weather.

“One of the big things that come up and prevents us from working well together is…we think about problems in very different ways,” Boyd tells Datanami. “Our main metrics for success of a problem is done differently.”

As a data scientist, Boyd’s goal was to conduct as many experiments as possible to find the best machine learning model for a given problem. His data engineer colleagues, on the other hand, were focused on building data systems that were maintainable, reliable, and wouldn’t cause them to get a phone call at 2 in the morning because something broke down.

“That affect kind of bleeds through in terms of how we organize our work, how we think about our work, etc.” Boyd says. “Data scientists focus on the experimentation. Data engineers focus on data pipelines. Machine learning engineers focus on model orchestration and productionalization.”

Instead of relegating himself to his silo in the data organization, Boyd reached across the aisle and searched for ways that could give him what he needed as a data scientists, without forcing his data engineering colleagues to compromise on their approach.

Step one is just getting in the same room together.  When data scientists and data engineers are forced to talk with one another, they will (hopefully) find some common ground. If the two personals keep the organization’s end goal in mind, then it will set the stage for future collaboration.

“They do have a lot of similar concerns. They’re both thinking about data and algorithms that process data,” Boyd says. “These are the only two disciplines that are thinking about them in conjunction with each other. They’re just thinking about them in different ways, so it’s great for them to unite and find common ground.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datanami.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.