How data lakehouses are vital to fuelling AI and the future of medicine

The pandemic has not only highlighted the importance of speed for medical discoveries, but also how data science and artificial intelligence (AI) can aid this acceleration. For example, machine learning in medicine has taken significant strides in recent years, withdrug molecules discovered through AI used in human trials. Despite this, a recent report from the Alan Turing Institute revealed that difficulties with data collection, use, storage, processing and integration with different systems, namely the lack of a robust data architecture, hindered efforts to build helpful AI tools in response to the pandemic.
To tap into the full potential of AI, organisations, especially in healthcare and pharmaceuticals, need to get their data in order. The question is, how?
While great efforts have been placed into the likes of drug and medical discovery, particularly in light of recent events, it can be a lengthy, complex and costly process. Not to mention, its low success rates – only a couple of years ago, the overall failure rate of drug development was reported to sit at 96%. This is where data has stepped in and is beginning to update methods and transform the potential of drug development to bring down that percentage.
Without human data, particularly genomic data, we cannot comprehensively capture all elements of a disorder or disease to gain a wider and deeper picture. This calls for sequencing on a very large scale to be able to discover and validate key genetic variants. More information and insights gathered means organisations can take better-informed steps and counteract a major cause of drug development failure – a lack of efficacy. Creating and establishing machine learning (ML) algorithms with this data is also enabling drug development pipelines to be automated – not only offering greater understanding but also accelerating drug discovery.
As another example, QSAR (Quantitative Structure-Activity Relationships) models are able to improve predictive accuracy on novel chemical structures as well as lower costs and time by reducing the number of compounds synthesised. Predictive analytics can also be used in drug development and manufacturing by transferring knowledge and incorporating learnings from rich historical data. This data can then be used to predict new compounds and accelerate the experiment lifecycle.
AI will and already is playing a big role in drug development, discovery and the clinical trials process. There are opportunities to accelerate clinical research with a modern approach to data and analytics.
Despite these steps forward, all this great data brings its own challenges. With so much biological and medical data now available, pulling out the necessary insights needed – and quickly – is harder than ever. There is no point having all this data if it cannot be properly utilised. Moreover, genomic data in particular requires a huge amount of storage, specialised software to analyse it and raises many data management, data sharing and also privacy and security issues – it is important to remember that this is highly sensitive and private information.


