More and more often, radiology residents that work in our Breast Imaging Division at the IEO (European Institute of Oncology) ask me whether it still makes sense to subspecialize in breast imaging, given that within five years an algorithm will probably be reading the mammograms alongside her.
The stock reply is by now a slogan: AI won't replace radiologists, but radiologists who use AI will replace those who don't. It is repeated at conferences the way a mantra is repeated, and it has the same effect — it soothes without informing. It tells a trainee nothing about what to do on Monday morning. Worse, it implies that the required adaptation is a matter of adoption, of simply agreeing to use the tool. That is precisely the assumption the recent evidence should be dismantling.
Two studies, one week, opposite directions
Consider two findings published within days of each other this summer.
Filippo Pesapane
On July 14, a group led by Adnan Taib at the University of Nottingham published an eye-tracking study in Radiology. Readers evaluating mammograms with AI assistance saw their median sensitivity fall from 71% to 39% when the algorithm produced a false-negative suggestion. Eye-tracking data explained the mechanism: with a reassuring output on the screen, readers fixated less often and dwelled for less time on the very cases that contained cancer. The visual search itself contracted.
Seven days later another study published on the same journal, Radiology, came the opposite headline: adjunctive AI raised the cancer detection rate of general radiologists toward that of dedicated breast imaging specialists — a genuine argument for widening access to expertise in settings that have no subspecialist to spare.
One study shows it closing a gap; the other shows it closing our eyes. Both results are real, and the variable that reconciles them is not the algorithm. It is the human being reading beside it.
This is the part our training never covered. We were taught to evaluate an instrument by its performance characteristics, as though sensitivity and specificity were properties a device carries with it, but AI does not behave that way.
Its performance is contingent on the population it meets, the equipment that generated the images, the prevalence in the screening programme, and — as the eye-tracking data now shows — on what it does to the observer. The unit of evaluation must be the clinician-model pair, and almost none of our educational infrastructure is built to assess it.
The curriculum nobody wrote
So what should we actually teach? Not what a convolutional network is. Here is what I would put on the syllabus instead — nine competencies that exist only because the algorithm exists.
Form your reading before you open the AI output. The order of operations is a clinical decision, not an interface preference, and it is the cheapest available protection against anchoring on someone else's answer.
Interrogate the training data. Which population, which vendor's equipment, which years, which cancer prevalence. A model validated elsewhere is not automatically a model for your patients. If the manufacturer will not answer, that silence is itself informative.
Distinguish a score from a probability. An output of 87 is not an 87% chance of malignancy. Calibration and discrimination are different properties, and a great many disappointments in this field begin by conflating them.
Read a saliency map for what it is. It indicates where, not why. It can make an incorrect assessment feel more convincing rather than less. Explainability should function as a prompt to look again, never as reassurance.
Document your disagreements. One line, every time you overrule the model or it overrules you. After a year, that record describes where this tool fails in your department, on your patients — something no published validation study can tell you.
Learn to notice when you are reading faster. Fatigue announces itself; automation bias does not. It is invisible from the inside, which is exactly what makes it hazardous.
Accept that performance travels badly. Same model, different population, different numbers. And models drift: a tool that was accurate at installation can decay silently, without any of the signals we are trained to recognize as malfunction.
Practice public disagreement. Saying out loud in a multidisciplinary meeting that the AI says one thing, you say another, and here is your reasoning — that is now a core professional skill, not a display of stubbornness.
Learn to answer the patient. She will ask whether a computer read her mammogram. "Yes" requires a second sentence, and it cannot be a shrug. Our own survey work suggests women are considerably more interested in this conversation than the profession has assumed.
Not one of those nine items is about finding the cancer. Every one is about knowing when the answer in front of you can be trusted.
We have run this experiment before
There is a precedent that rarely appears on the programme of AI courses, and it should be on every one of them: computer-aided detection (CAD).
CAD was approved, adopted at scale, and reimbursed in the United States. Years later, studies involving hundreds of thousands of women failed to demonstrate the benefit that had been assumed. The technology did not fail in isolation. What failed was the way we evaluated it, deployed it, and — critically — the way we read alongside it.
The structural conditions for repeating that history are in place. There are now more than 1,100 FDA-cleared AI devices in radiology, representing roughly three-quarters of all authorized medical AI, with dozens more clearing each quarter. Regulatory approval is no longer the bottleneck; implementation is. And implementation is where the education deficit lives.
I would apply a simple test to any course, congress session, or curriculum claiming to teach AI in imaging. Does it answer this question: how would I notice that this tool is failing, in my department, on my patients, before anyone else does? If that question is not addressed anywhere in the programme, the programme is about something else.
A harder job, not a smaller one
The honest answer to my resident's question is not comforting, and does not need to be.
Yes — an algorithm will very likely read the mammogram. And someone will still have to decide, case by case and woman by woman, whether that reading holds. Doing so will require understanding the model's provenance, its calibration, its failure modes in the local population, and its effect on one's own perception. That is not a consolation prize handed to a displaced profession. It is a job description, and a considerably more demanding one than the job we had before.
Which suggests the goal of radiology education over the next decade is not to train physicians who can outperform machines at detection. That contest is uninteresting and, in narrow tasks, increasingly settled. The goal is to train physicians who are the reason the machines are safe.
Dr. Filippo Pesapane is a breast radiologist at the European Institute of Oncology (IEO) IRCCS in Milan, Italy, where his research focuses on the clinical validation, explainability, and regulatory framework of artificial intelligence in breast imaging.
The comments and observations expressed herein do not necessarily reflect the opinions of AuntMinnieEurope.com, nor should they be construed as an endorsement or admonishment of any particular vendor, analyst, industry consultant, or consulting group.













![A normal mammogram confirmed by three-year radiologic follow-up illustrates reader-marked regions of interest (ROIs) during (A) unaided (round 1) and (B) artificial intelligence (AI)–assisted (round 2) reading. Each colored dot represents an ROI for recall by a human reader. Readers could mark more than one ROI per case, represented by multiple dots of the same color. During AI-assisted reading, the AI system displayed three visible prompts: two with suspicion of malignancy scores of 35% (left mediolateral oblique [L MLO] and craniocaudal [L CC]) and one with a suspicion of malignancy score of 10% (right craniocaudal [R CC]), shown as polygonal overlays. Without AI, six of 10 readers (60%) marked a false-positive ROI. With AI assistance, this fell to two of 10 (20%). R MLO = right mediolateral oblique.](https://img.auntminnieeurope.com/mindful/smg/workspaces/default/uploads/2026/07/2026-07-14-radiology-mammogram-ai-auto-bias.H0bYO8QlWs.jpg?auto=format%2Ccompress&fit=crop&h=112&q=70&w=112)






