What IBM looks for in a data scientist

Job seekers sometimes ask how IBM defines “data scientist.” It’s an important question since more and more would-be data scientists are fighting for attention in an increasingly lucrative labor market.
The first step is to distinguish between what we see as true data scientists and other professionals working in adjacent roles (for instance, data engineers, business analysts, and AI application developers). To make that distinction, let’s first define what we mean by data science.
At its core, data science is applying the scientific method to solve business problems.
You can further expand on the definition by understanding that we solve those business problems using artificial intelligence to create predictions and prescriptions and to optimize processes.
The definition demonstrates that to achieve the true potential of data science, we need data scientists with very particular experiences and skills — specifically, we need people with the experiences and skills required to run and complete data science projects:
1. Training as a scientist, with an MS or PhD 2. Expertise in machine learning and statistics, with an emphasis on decision optimization 3. Expertise in R, Python, or Scala 4. Ability to transform and manage large data sets 5. Proven ability to apply the skills above to real-world business problems 6. Ability to evaluate model performance and tune it accordingly
Let’s look at those qualifications in the context of our definition of data science.
This is less about the degree itself and more about what you learn when you get an advanced degree. In short, you learn the scientific method, which starts with the ability to take a complex yet abstract problem and break it down into a set of testable hypotheses. This continues with how well you design experiments to test your hypotheses, and how you analyze the results to see whether the hypotheses are confirmed or contradicted. A determined person can learn these skills outside of academia or via the right mix of online training and practice — so there’s some flexibility around having the actual degree — but direct experience applying the scientific method is a must.
Another advantage of an advanced degree is the rigor of the peer review process and publishing requirements that the degree programs impart. To get published, candidates have to present their work in a way that allows others to review and reproduce it. You must also provide evidence that the results are valid and the methods are sound. Doing so requires a deep understanding of the difference between probabilistic and deterministic factors as well as the value and curse of the correlation. It’s possible to get an abstract sense of those values, but there’s no substitute for the negative and positive reinforcement from mentors or the rejection or acceptance of journals and reviews.
Applying the scientific method to business problems lets us make better decisions by predicting what will happen next. Those predictions are the product of artificial intelligence and more specifically machine learning. For a true data scientist, the core technical skillsets of machine learning and statistics are simply non-negotiable.
In addition, decision optimization (aka operations research) is a fast-growing aspect of data science. Indeed, the goal of data science is to help make better decisions by probabilistically estimating what’s likely to occur in the future.


