How Can Doctors Be Sure A Self-Taught Computer Is Making The Right Diagnosis?

4 min read
Curated from npr.org →

Some computer scientists are enthralled by programs that can teach themselves how to perform tasks, such as reading X-rays.

Many of these programs are called “black box” models because the scientists themselves don’t know how they make their decisions. Already these black boxes are moving from the lab toward doctors’ offices.

The technology has great allure, because computers could take over routine tasks and perform them as well as doctors do, possibly better. But as scientists work to develop these black boxes, they are also mindful of the pitfalls.

Pranav Rajpurkar, a computer science graduate student at Stanford University, got hooked on this idea after he discovered how easy it was to create these models.

The National Institutes of Health one weekend in 2017 made more than 100,000 chest X-rays publicly available, each tagged with the condition that the person had been diagnosed with. Rajpurkar texted a lab mate and suggested they should build a quick and dirty algorithm that could use the data to teach itself how to diagnose the conditions linked to the X-rays.

The algorithm had no guidance about what to look for. Its job was to teach itself by searching for patterns, using a technique called deep learning.

“We ran a model overnight and the next morning I woke up and found that the algorithm was already doing really well,” Rajpurkar says. “And that got me really excited about the opportunities, and the ease with which AI is able to do these tasks.”

Fast forward to February of this year, and he and his colleagues have already moved far beyond that point. He leads me to a sun-filled room in the William Gates (yes, that Bill Gates) Computer Science Building.

His colleagues are looking at a prototype of a new program to diagnose tuberculosis among HIV-positive patients in South Africa. The scientists hope this program will help fill an urgent medical need. TB is common in South Africa, and doctors are in short supply.

The scientists lean into the screen, which displays a chest X-ray and the patient’s basic lab results and highlights the part of the X-ray that the algorithm is focusing on.

Get the AI & data signal, daily.

335k+ subscribers read this every morning. One email, both newsletters. Unsubscribe anytime.

The scientists start scrolling through examples, making guesses of their own and seeing how well the algorithm is performing.

Stanford radiologist Matthew Lungren, who is the main medical adviser for this project, joins in. He readily admits he is not great at identifying TB on an X-ray. “We just don’t see any TB here” in the heart of Silicon Valley, he explains.

True to his warning, he misdiagnoses the first two cases he sees.

Rajpurkar says the algorithm itself is far from perfect, too. It gets the diagnosis right 75 percent of the time. But doctors in South Africa are correct 62 percent of the time, he says, so it’s an improvement. The usual benchmark for TB diagnosis is a sputum test, which is also prone to error.

“The ultimate thought from our group is that if we can combine the best of what humans offer in their diagnostic work and the best of what these models can offer, I think you’re going to have a better level of health care for everybody,” Lungren says.

But he is well aware that it’s easy to be fooled by a computer program, so he sees part of his job as a clinician to curb some of the engineering enthusiasm. “The Silicon Valley culture is great for innovation but it’s not got a great track record for safety,” he says. “And so our job as clinicians is to guard against the possibility of getting ahead of ourselves and allowing these things to be in a place where they could cause harm.”

For example, a program that has taught itself using data from one group of patients may give erroneous results if used on patients from another region — or even from another hospital.

One way the Stanford team is trying to avoid pitfalls like that is by sharing their data so other people can critique the work.

Continue Reading

Enjoyed this summary? Read the complete article at the source:

Continue at npr.org →

Yves Mulkers

Yves Mulkers is the founder of 7wData and a widely followed voice in the data and AI community. He curates the 7wData and AI Beat newsletters, reaching hundreds of thousands of data and AI professionals, and writes on data strategy, analytics, AI, and the evolving data ecosystem.