Adding Artificial Intelligence to Drug Discovery

Scientists face slim odds when trying to turn a molecule into a medicine. Most studies put the batting average at about 0.100—or 1 in 10. Some go a little higher, some a little lower, but the success rate for drug discovery is never “good.” Some scientists believe that the success rate could be improved if drug discovery were to apply artificial intelligence (AI), that is, if it were to use advanced computational tools such as machine learning (ML) and molecular dynamics simulation. Besides leading to more medicines, AI might even allow the creation of better medicines.
Traditionally, a drug discovery project starts with basic research to uncover targets that may be susceptible to attack, such as a disease-related protein receptor on the surface of particular cells. Then, scientists use techniques like high-throughput screening to see which compounds bind the target. (These compounds can come from libraries of tens of thousands or even millions of molecules at large pharmaceutical companies.) After that, various methods of biological and chemical testing are used to fine-tune the structure or test other features, such as a compound’s ability to reach the target in an organism.
To some extent, AI can be used in all these steps. Essentially, AI looks for patterns in data that can be used to sharpen predictions about which compounds will become medicines. Some AI-based tools are already being applied to discovering tomorrow’s medicines. It remains to be seen how ready the industry is for this transition.
When it comes to improving medicine with computational tools, one of the top players is GNS Healthcare. The company’s chief commercial officer Iya Khalil, PhD, says, “AI is being used to really leverage learning from large-scale datasets, and those large-scale datasets are used to get to better novel targets.” But she adds, “We’re still in the era of really trying to leverage what we can learn from genetic data and trying to understand what are the causal drivers of disease.”
The missing information about what drives a biological system—healthy or diseased—explains why a compound often fails in Phase II or III trials. With better knowledge—created from applying AI to whole-genome, phenotypic, and clinical data—scientists will find better starting targets, Khalil believes, “because you’re learning it directly from the human population.” Then, by turning off specific genes to see what changes, scientists may get even more data to dump into computational models.
Many techniques use AI like a black box: It finds patterns in the data, but no one knows what, if anything, they mean. Instead, Khalil prefers causal, or white box, AI. “We use it to get causality right from the beginning, and not just learn patterns,” she notes.
Still, making that work depends on lots of data, what Khalil calls deep data, such as a patient’s genomic and molecular data, and phenotypic and clinical data. Getting that data requires interdisciplinary teams. As an example, GNS Healthcare works with the Multiple Myeloma Research Foundation and others to collect more data. Khalil would like to work with hundreds of thousands of variables per person. It sounds like a lot, adding up to billions of data points, Khalil admits, but “it’s data that we can collect today.”
To put so much data to work in the best ways, Khalil and her colleagues do in-house development and coding, and run the simulations with cloud computing, like Amazon Web Services.
Data is so plentiful that scientists struggle to make sense of it. Alexis Borisy, executive chairman of the board at Celsius Therapeutics, says that the company analyzes tens of thousands of gene transcripts in cells from hundreds of human samples. In all, Celsius scientists work with a dataset with millions of dimensions. “That’s more data and dimensionality than you can look at and understand,” he says. “Machine learning helps you focus on the key things.”
No matter how sophisticated the AI might be, data fuels it.


