Data visualisation isn’t just for communication, it’s also a research tool

At the heart of the scientific method lies the ability to make sense from data.
However, this is a challenge in the fast-moving field of biotechnology, where new experimental methods are creating huge amounts of complex data. These data promise to revolutionise healthcare, food and agriculture, but it can be difficult to extract answers to specific research questions from these sets of numbers.
Data visualisation can help. Our eyes deliver information very rapidly to our brains, and then sophisticated pattern recognition abilities take over. Well-designed visualisation tools can reveal discoveries that would otherwise remain buried.
Below we highlight three data visualisation tools we have developed to help life scientists find relevant and useful information amongst the noise. The visualisation principles used in these tools are general and help in many complex data challenges.
Proteins and other molecules in our bodies exist as complex 3D structures that constantly change shape and interact with each other. Mapping out the many possible ways that proteins can be structured helps scientists understand how biological processes work, and may inform drug development and treating diseases such as cancer.
Thanks to decades of research worldwide, we now have reliable, evidence-based 3D structures for tens of thousands of proteins, plus more than 100 million models of protein structures.
These models are useful for learning about life’s molecular processes – such as how RNA and proteins are made – however, the large number of models can make it difficult for scientists to pin down which specific models can help answer a particular research question.
To address this difficulty, one of us (Seán O’Donoghue) and colleagues developed Aquaria, a tool using the visualisation principle of “overview first, details on demand”. By using a technique called “clustering”, Aquaria creates a concise visual overview of all structural models available for any specific protein.
The image above shows this overview for p53, a protein that protects against cancer. Each cluster of related 3D models can be interactively expanded and explored (bottom of the image), helping scientists find the most useful models suited to address a specific research question.
Once a suitable model is found it is shown (top of the image), with dark colouring used to indicate regions where the structure of the model is less certain. In addition, yellow, blue and green are used to highlight different shapes within the structure, which helps scientists understand how the protein is arranged in three dimensions.
Sometimes, we need to look at data from multiple viewpoints. This is particularly true for a field of research known as sequencing. Sequencing involves determining the precise order of the chemical building blocks that make up DNA, RNA and protein. Knowing these sequences and comparing how they vary between individuals can tell us about mutations that cause disease and reveal how we evolved.
One of the most widely used tools for visualising sequences is Jalview, co-developed by one of us (James Procter), which brings together the huge amounts of data that are created through sequencing.

