In 2016, the Nobel-winning Geoffrey Hinton predicted that radiologists — the physicians who read X-rays, ultrasounds and other images to help make diagnoses — would be replaced by computers within five years. A decade later the field is still here. In fact the number of practitioners is expected to expand by 26 percent or more over the next three decades.
But Hinton being wrong about labor market dynamics should not obscure the part he got right: physicians now have a silicon-based colleague in the room that matches or exceeds their performance.
Seeing the difference requires looking at the number: Hinton predicted the profession would disappear; what happened is that it grew while its content changed. That distinction is persistently conflated in debates about AI's effect on employment.
That is why radiology matters. It is medicine's hot spot for AI, and as such a bellwether for how expert decision-making systems get adopted across healthcare and perhaps other fields.
The scale
As of early 2026, about three-quarters of the 1,400 AI-enabled medical devices cleared in the US were for radiology. These tools do two different jobs:
- Efficiency — drafting reports or flagging the images that urgently need attention.
- Performance — identifying abnormalities that may not be visible to the human eye, interpreting images as well as, and sometimes better than, a trained radiologist.
An analysis of 43 clinical trials concluded that AI-assisted colonoscopies reveal more polyps than conventional ones.
Improving accuracy matters, because average human error rates on diagnostic images are estimated at 3 to 5 percent — millions of errors worldwide each year. Yet the solution is not as simple as replacing humans with machines. A decade after Hinton's prediction, the question is not whether humans or AI do statistically better at a given task, but how AI's technical precision can be combined with human experience and flexibility.
Two different kinds of intelligence
Curtis Langlotz, a radiologist who directs the Center for Artificial Intelligence in Medicine and Imaging at Stanford, explains with an example. Suppose an AI tool detects 95 percent of the lung nodules on a chest CT, while radiologists detect 90 percent.
"Some would say, 'Oh, the AI is better than the radiologists, therefore we should replace all the radiologists with AI,'" he says. "But I will tell you that the radiologist is going to detect some of that 5 percent that were missed by the machine. And that's because machine intelligence and human intelligence are different kinds of intelligence."
The difference lies here: AI can examine every pixel on an image without getting tired or distracted and compare it to every other image it has encountered. A radiologist's understanding of disease allows them to interpret images in ways AI cannot.
The new job: evaluating every decision
This changes the job description. The radiologist is now in the role of evaluating each AI decision — the vast majority of which will be correct — and identifying the rare instances where the algorithm got it wrong.
Paul Yi, a radiologist at St. Jude Children's Research Hospital, sums it up: "This requires a whole mental rewiring."
This role differs from the familiar pattern where a human decides and a machine assists. Here the machine decides first and the human audits it. The order of decision has reversed while the responsibility has not moved.
Physicians are used to overruling computers. Electronic medical record alerts triggered by rule-based algorithms have warned them for decades about things like dangerous drug interactions, and physicians override those warnings about half the time. But this moment is different.
The black box problem
The difference comes from the kind of tool. Rule-based algorithms give answers physicians can quickly interpret using their own medical knowledge. AI image analysis systems in radiology are usually built on neural networks that can identify tumor subtypes, outline lesion boundaries and perform diagnostic tasks with high accuracy. But they do not reveal how they reached a decision.
Charles Kahn, editor of Radiology: Artificial Intelligence, describes the difficulty: "Is that really an abnormality, or am I missing something that AI with its subtle mind or whatever is identifying? And that is really challenging for us."
A trap that runs both ways
Interpreting an image allows two kinds of error: a false negative misses disease, a false positive sees one that is not there. For uncommon diseases — and most diseases are relatively uncommon — any diagnostic tends to produce many false positives, simply because it encounters large numbers of people without the disease.
Two distinct biases follow:
- Automation bias — believing the machine's result too readily.
- Automation complacency — accepting something the machine missed, such as blood in the brain, without question.
Nina Kottler, chief medical AI officer at Mosaic Clinical Technologies, connects them: "These are both issues of letting the AI change the radiologist's level of suspicion without realizing it." Over time people may start believing AI's answers too much. One study found even experienced radiologists saw big drops in mammography accuracy when their decisions were made under the influence of incorrect AI predictions.
The opposite direction is more intuitive. Being suspicious of a new technology — especially one accused of taking your job — is natural until it proves itself. One detail feeds that: otherwise reliable AI systems make obvious mistakes unlike human errors. "So you see a silly mistake, a silly false positive, and you're like, 'Well, of course the AI is not smart,' and you could then dismiss it," Kottler says.
Rare disease, abundant false alarms
Understanding why false positives pile up is the key to understanding the radiologist's job. The issue is not that the tool is bad; it is that the disease is rare.
A screening tool encounters large numbers of people who do not have the disease. Each of them is another opportunity for the tool to raise a false alarm. The rarer the disease, the more false alarms accumulate per correct diagnosis — even if the tool's accuracy never changes.
The consequence is that a radiologist's job consists largely of saying no. They must overrule most of what the machine flags while not missing the rare correct diagnosis among them. That is exhausting mental work, demanding constant suspicion and constant attention at the same time.
How it gets managed
Kottler's proposal is measurement: monitor over time how much radiologists agree with their AI colleagues, and intervene if they drift too far either way. "If radiologists are accepting the AI result 99 out of a hundred times and we know it's only 95 percent accurate, we go talk to that rad."
But calling a tool 95 percent accurate is not enough. Training is essential so physicians know what to look for and when. They need to know, for instance, that a certain tool will be wrong on 30 percent of scans if the patient moved in the scanner.
There is a gap here. In a 2026 physician survey, more than a quarter of respondents said they had received no training about AI; only 11 percent said they had received a lot.
Kottler also advocates that tools report a confidence estimate for each evaluation instead of a simple yes/no. The reasoning is practical: no radiologist can be expected to know the intricacies of multiple AI systems, and those systems will only grow in number and complexity.
The number on the box is not the number in your hospital
One implication deserves emphasis. Most AI imaging tools are trained on data from particular patient populations, and a model's accuracy can drop on a population different from the one it learned from.
That sharpens the article's central finding: the accuracy printed on the box may not be the accuracy in your hospital. This is precisely why the measurement Kottler proposes — tracking how often radiologists agree with the tool — has to happen at the institutional level. A purchased tool that is not measured on site is running at an unknown accuracy.
Beyond radiology
Langlotz's summary captures the field's maturity: "We are just at the very beginning of understanding how to optimize the human/machine system."
This picture concerns more than radiologists. Every profession where expert judgment gets machine support will follow the same sequence: the tool arrives, then the job description changes, and only last comes the training and measurement to manage that change. It is always the last one that is missing.
The order matters too. When an institution buys the tool first and thinks about training later, staff use it on instinct in the interval — and that instinct drifts toward one of the two biases described above. Some over-trust the machine; some dismiss it entirely after one absurd error. Neither is visible unless it is measured.
What Langlotz said a decade ago, rebutting the godfather of AI's prediction, still holds: AI will not replace radiologists. Radiologists who use AI will replace radiologists who don't.