Data visualisation isn’t just for communication, it’s also a research tool

3 min read

At the heart of the scientific method lies the ability to make sense from data.

However, this is a challenge in the fast-moving field of biotechnology, where new experimental methods are creating huge amounts of complex data. These data promise to revolutionise healthcare, food and agriculture, but it can be difficult to extract answers to specific research questions from these sets of numbers.

Data visualisation can help. Our eyes deliver information very rapidly to our brains, and then sophisticated pattern recognition abilities take over. Well-designed visualisation tools can reveal discoveries that would otherwise remain buried.

Below we highlight three data visualisation tools we have developed to help life scientists find relevant and useful information amongst the noise. The visualisation principles used in these tools are general and help in many complex data challenges.

Proteins and other molecules in our bodies exist as complex 3D structures that constantly change shape and interact with each other. Mapping out the many possible ways that proteins can be structured helps scientists understand how biological processes work, and may inform drug development and treating diseases such as cancer.

Thanks to decades of research worldwide, we now have reliable, evidence-based 3D structures for tens of thousands of proteins, plus more than 100 million models of protein structures.

These models are useful for learning about life’s molecular processes – such as how RNA and proteins are made – however, the large number of models can make it difficult for scientists to pin down which specific models can help answer a particular research question.

To address this difficulty, one of us (Seán O’Donoghue) and colleagues developed Aquaria, a tool using the visualisation principle of “overview first, details on demand”. By using a technique called “clustering”, Aquaria creates a concise visual overview of all structural models available for any specific protein.

Once a suitable model is found it is shown (top of the image), with dark colouring used to indicate regions where the structure of the model is less certain. In addition, yellow, blue and green are used to highlight different shapes within the structure, which helps scientists understand how the protein is arranged in three dimensions.

Sometimes, we need to look at data from multiple viewpoints. This is particularly true for a field of research known as sequencing. Sequencing involves determining the precise order of the chemical building blocks that make up DNA, RNA and protein. Knowing these sequences and comparing how they vary between individuals can tell us about mutations that cause disease and reveal how we evolved.

One of the most widely used tools for visualising sequences is Jalview, co-developed by one of us (James Procter), which brings together the huge amounts of data that are created through sequencing.

Jalview employs two principles – “linking and brushing” and “multiple coordinated views” – to bring together different types of information. Jalview also allows other tools to be connected, enabling scientists to navigate through complex, interrelated datasets.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at datadrivenjournalism.net →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.