How To Grow As A Data Scientist

4 min read
Curated from kdnuggets.com →

In order for a data scientist to grow, they need to be challenged beyond the technical aspects of their jobs. They need to question their data sources, be concise in their insights, know their business and help guide their leaders.

The role of a data scientist still varies from company to company and even team to team. This makes it much harder for companies to create a standardized growth plan for their data scientists.

Without having a clear growth plan, there is a risk that these talented computer wizards will get stuck. They might provide good insights, but they will never really grow and provide the true ROI they have to offer a business or more importantly themselves.

With this in mind, our team talked to managers around Seattle working at the top tier tech companies to find out what they wanted and expected from their senior data scientists. We wanted to share the information we learned both to help data scientists grow as well as help managers who are trying to challenge their new data scientists to grow.

Based off our discussions we found that it wasn’t about programming, or designing algorithms (that was a baseline for a Jr. data scientist). When we asked these managers what they wanted to see from their more senior data scientists they informed us that they wanted driven individuals who can communicate concisely, who were able to think for themselves, who have a solid understanding of the business and who are capable of managing up.

In order for a data scientist to grow, they need to be challenged beyond the technical aspects of their jobs. Data scientists have the opportunity to sway company decisions. They have a lot of responsibility on their shoulders. That means they need to take ownership of the work they do. They need to question their data sources, be concise in their insights, know their business and help guide their leaders.

  A senior data scientists won’t just trust their data after receiving it. They will poke and prod it for things like bias, missing data, duplicate data, etc.

Data is bound to have quirks. For those who spend hours and hours in data, you know what I am saying. While scrolling, or graphing data you see those strange patterns that make you stop and say “I wonder why x looks like z”. Younger data scientists will often be too focused on finishing the project. They haven’t learned how to stop and really analyze these strange patterns. These patterns can be caused by systems that default specific data outputs like -1 or 1 or maybe even biased data caused by purchasing bots that might skew what customers actually are buying on an e-commerce site, and a thousand other plausible causes of misleading data.

These patterns are not necessarily incorrect or bad data. Even when data is accurate there will always be operational quirks. When designing reports, algorithms and metrics, these need to be considered. An experienced data scientist will not only look for these data quirks, they will expect them.

The term source of truth gets thrown around a lot in data teams. It refers to the original data source that multiple teams have decided is correct. I was very naive when I started out as a data scientist. On one of my first projects, I was informed about a data source that our team had labeled as the source of truth. For months I worked on our “Source of Truth” developing analytics and applications to help over 200 managers and directors have access to that data. Of course, it wasn’t too long until there were consistency issues with other metrics. It was then that I realized I had been working on a data source several ETLs from the source of truth.

Talking to tech managers across Seattle. This is a common issue. Young analysts, data scientists and developers are overly trusting of their data sources. Typically, the younger, less experienced employees will be very eager to get the work done. This will inadvertently lead to less understanding of what the data actually is.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at kdnuggets.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.