4 Things Data Scientists Can Learn From SoundCloud’s Process

According to a recent blog post from the SoundCloud team, the company has recently restructured and reorganized the way its data scientists and analysts function.
The goal is to help them be more effective, happy and productive members, hopefully improving many internal processes and operations.
If you want to dive into the process in more detail, you can read about it here.
Essentially, it parsed its entire strategy down into smaller segments or steps — beginning with problem definition, to preparation, to solution development, to production-ready deployment, and finally validation and maintenance. It’s a process that helps iron out the kinks and optimizes the solutions it’s deploying.
Looking at this, it’s not out of the question to wonder what it means for the big data world as a whole. How can this be evolved and reproduced elsewhere?
To make sense of it, first you need to understand the foundation and processes that help SoundCloud stand and walk.
Operating on a global scale, the SoundCloud audience uploads about 12 hours of audio content every minute. That is an extremely large amount of data being stored and processed on its servers. For each audio file uploaded, it must be transcoded and stored in varying formats.
This allows the customers or audience to download content in the format they prefer and use it on the device they prefer, such as an iPhone or iPod over a standard MP3 player.
Because SoundCloud is a hub for thousands — if not millions — of artists, bands, podcasters and audio creators, it needs to have an incredibly vast storage capacity to deliver.
Furthermore, all content uploaded through the service can and will be shared via blogs, websites, social networks, mobile apps and chat services.
The company’s audience is active 24 hours a day, seven days a week. That means traffic and performance requirements fluctuate. If and when multiple regions use the platform at the same time — such as the U.S. and U.K. — there are extremely high loading spikes.
That means the team behind the platform needs to be able to collect all this user and performance data and put it to work.
Alexander Grosse, vice president of engineering at SoundCloud, says “if our storage crashed, that would be the end of SoundCloud, we [must] focus on our platform’s core functionality.”
It needs the data its scientists are collecting to improve, enhance and support the product. Furthermore, it needs to be converted and made available for nearly everyone, including those with little to no analytics experience.
Before organizing or making sense of data, you first have to understand the problem and what solutions can be used to solve it. This means you must understand the business needs, identify issues via metrics and narrow the scope so it’s manageable.
For example, one problem SoundCloud might have is performance issues. At the outset, it won’t exactly know where the issues are coming from or what’s causing them.


