6 Tips for Better Data Science in the Cloud

The cloud has transformed what is possible with data science. Data teams now have access to a vast pool of elastic computing power, numerous sources of internal and external data, and managed cloud services that reduce the complexity of building, training and deploying machine learning and deep learning models at scale.
But that doesn’t mean there aren’t challenges as teams adapt from an on-premises infrastructure to a cloud-based model. Data scientists, data engineers and developers are all having to learn and adapt to a new environment, and there is an ever-expanding and rapidly evolving ecosystem of tools and frameworks from which to choose. Many are learning on the job, figuring it out as they go.
The very capabilities that make the cloud so exciting also create potential pitfalls to watch out for. The ease of copying data across diverse systems can create governance challenges if not handled properly. The speed of change means that data teams can bet on the wrong tool or framework and become stranded there. Habits and biases from the on-premises world can limit understanding of what’s possible in the cloud.
After building data management technology for many years, and from frequently talking to organizations of all sizes across all industries, I’ve seen some common pitfalls and misunderstandings that can hold data teams back from doing great work. The cloud opens an exciting frontier to better understand customers, monetize data in new ways and make predictions about the future. So I hope the following tips will allow data teams to capitalize on those benefits, while working in a way that is secure, efficient and effective.
It’s critical to enable iteration and investigation without compromising governance and security. For example, many data scientists intuitively want to copy a dataset before they start working on it. But it’s too easy to make copies, move on and forget they exist, creating a nightmare in terms of compliance, security and privacy. A modern data platform should allow you to work on snapshots, or virtual copies, without needing to duplicate entire datasets, while maintaining fine-grained controls to ensure that only the right users and applications have access to it. Create processes that minimize copies and clean up anything copied; don’t be the person that gets your company in the news headlines for the wrong reasons.
If you’re coming from an on-premises world, you’ll often bring perceptions and biases about infrastructure that no longer apply to modern platforms in the cloud. I’ve often heard data scientists say, “I’d love to retrain my model several times a day, but it’s too slow and will delay other processes.” But that’s not an issue in a world of elastic infrastructure. Approach the cloud from first principles. Start with what you want to achieve, not what you think is possible, and move forward from there. That’s the only way to push the boundaries and take full advantage of this new environment.
Closely tied to data governance is the concept of silos. In the cloud, it’s important not to replicate the fragmentation that’s common in the on-premises world..


