Big Data Investment Pays Off for PerkinElmer

3 min read
Curated from datanami.com →

Nearly a decade ago, the folks at PerkinElmer made the decision to rebuild their flagship software offering for managing scientific research. As data was growing and becoming more diverse, the Oracle-based relational technology that underpinned the platform struggled to keep up. The company sought a new technological direction with its cloud-based Signals Research Suite, and found it with emerging open source frameworks, including Spark, Kubernetes, and Elasticsearch.

Based near Boston, Massachusetts, PerkinElmer has built a solid reputation for providing a range of products and services for some of the largest pharmaceutical companies in the world, including names like Bristol-Meyers Squibb, Merck, Pfizer, Johnson & Johnson, and Glaxo SmithKline, among others. The 84-year-old publicly traded company operates in 150 countries, and brings in around $4 billion a year in revenue.

In their quests for creating the next blockbuster drug or treatment, the research and development arms of these pharmaceutical firms are constantly formulating and testing new compounds to determine which candidates to take to the next stage. (The company also serves other manufacturers, such as those that develop paint and other coatings.)

For years, PerkinElmer has developed a software product called an electronic notebook that automates many steps in this process, from gathering raw data from scientific instruments and equipment and crunching the results, to managing the workflow and presenting the results.

PerkinElmer previously developed version of this electronic notebook was an on-prem system based on Oracle’s database. The software was capable of doing sophisticated analysis, such as enabling scientists to perform searches using chemical equations and fuzzy logic. This capability was important as it allowed scientists to search not only for exact matches among chemical compounds, but compounds that are related in some way.

However, as the pharma and manufacturing companies scaled their operations, the Oracle-based offering was starting to hit its architectural limit. Big pharma firms needed to store data about tens of millions of molecules, including tens of billions of test results. Scaling the infrastructure was becoming costly.

According to David Gonsalvez, the director of product portfolio and informatics R&D for PerkinElmer, one of PerkinElmer’s early customers were spending millions of dollars to scale its software on Oracle Exadata machines.

“They just threw money at it,” Gonsalvez said. “The only way they could get the system to scale is by being the forefront. They were the first ones to use the Exadata technology…They had these massive machines to be able to cope with the demands of this large electronic notebook datasets.”

About eight years ago, PerkinElmer made a strategic decision to shift directions. Instead of trying to scale the product vertically atop relational technology like Oracle’s database, PerkinElmer decided to embrace some of the new open-source databases and technologies that scaled horizontally. At the same time, it also decided to move to the cloud and the software-as-a-service (SaaS) delivery model that it enabled.

The company adopted MongoDB to store research documents, Elasticsearch to search across the documents, and Apache Spark to power parallel data processing. It also uses Kubernetes to orchestrate the containerized software on the AWS cloud, and it uses S3 to store all the data.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datanami.com →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.