Understanding the Changing Position Roles in Data Science

4 min read

Summary:  Is everyone a ‘data scientist’?  What about ‘data engineers’ and the junior versus senior, or skill level distinctions?  We do seem to need some agreement about titling.  Data Scientists is still the prestige title but there are some folks lobbying to take that title away.

Over just the last 12 to 18 months there has been more and more written about new job titles and roles in our data science profession.  It was perhaps only three or four years ago where if you worked in data science you were called a “Data Scientist”.  Regardless of degree, regardless of the specific tasks you performed, or the level of experience you possessed this was the title. 

If an employer wanted customer predictive analytics, a recommendation engine for ecommerce, an expert in (then brand new) deep learning, or even someone to set up a Big Data data lake or streaming system architecture, the ad still read “Data Scientist”.  Increasingly though we see mention of new titles like “Data Engineer” or “Predictive Analytics Professional” that seek to carve off portions of the broader Data Scientist title. 

There has always been a little under-the-covers dissent about exactly who should be allowed to call themselves a Data Scientist.  For the most part a lot of that came from folks with Ph.Ds. who no doubt thought, perhaps correctly, that only those who had invested as heavily as they had earned that title. 

The whole issue of titling in our profession has never been agreed or expressed with any clarity.  Not even between junior and senior folks performing the same tasks.  So perhaps a little clarity would be welcome.

You Can Still be a Data Scientist

Make no bones about it.  Everyone in our field still wants to be called a Data Scientist.  There’s money and prestige in the title and we all worked hard to get here.  The titles and job descriptions we’re going to discuss here are not widely agreed but they do come with some compelling logic from some credible sources.

If there’s a problem it’s that these titles want to restrict the title ‘Data Scientist’ to just a few of us and that’s going to make a lot of practitioners unhappy. 

More importantly, employers will probably continue to be ignorant of these distinctions for quite a long time.  It’s still as common now as four years ago to see an ad for ‘Data Scientist’ applied to a pretty much SQL-on-static-data job, and the term ‘Analyst’ applied to a fully skilled predictive analytics job.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

In the end, regardless of what is written here, it’s what your employer agrees to call you.

The first subset and the one that makes the most sense to me is the fairly new term “Data Engineer”.  As it’s used, this is intended to describe a person predominately skilled in CS who plays an important and supporting role to Data Scientists.

Where the Data Engineer used to be called an Analyst or some even less descriptive title within IT, the core of the task is the ETL and blending of the data from various sources that feed up to the Data Scientist. 

Typically one Data Engineer could support several Data Scientists, and some vendors like Qubole have made a business trying to make the relationship between Data Engineer and Data Scientist as efficient as possible.  You could also argue that blending platforms like Alteryx serve this same relationship.

The term has also been expanded to those who do the planning and execution of Big Data architectures.  This ranges from the setup of NoSQL DBs like Hadoop, to establishing streaming and IoT architectures like Spark Streaming, and includes projects of intermediate difficulty like setting up data lakes.

It seems clear that the path to entry level junior level positions involves good SQL skills but that the path to seniority passes through mastery of increasingly sophisticated data provisioning platforms.

Does this require R or Python mastery?  Not in the sense that the Data Engineer must produce statistical or machine learning models.

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.