Skip to content
Miracle OlajuyigbePHYSICIAN · MEDICAL WRITER

Health-Tech

Why Healthcare AI Needs Doctors in the Loop

By Miracle Olajuyigbe 5 min read


14:32:06. A clinician approves an AI-generated recommendation.

14:32:19. She approves the next one.

Thirteen seconds apart. The audit log records two human reviews, and everything the compliance documentation promised is, technically, happening.

Thirteen seconds is not a review. It is a signature. And a signature is not oversight, it is the quiet transfer of liability from a company to a person who was never given the conditions to earn it.

“Doctor in the loop” has become the most reassuring phrase in healthcare AI. In a great many deployments, it describes almost nothing at all.

I do not think companies are lying about this. Something more awkward is going on. The clinician genuinely is there. She has an account, she sees the output, she clicks approve. Every box is ticked. The safeguard exists on paper and evaporates in practice, and almost nobody is measuring the difference.

What we meant by it

The intention behind human oversight is not complicated.

Software produces an output. A clinician weighs it against the patient in front of them, against the parts of the story that never made it into a structured field, and against their own accumulated sense of when something does not add up. Then they decide.

The clinician is supposed to be the error-correcting layer. The one who catches the cases where the software is confidently wrong, where the training data looked nothing like this patient, where the input was rubbish in a way the model could not detect.

That is a real job. It is also work, it takes time, and hardly anyone is costing it.

What it becomes

Picture a clinician with a full list and an inbox producing dozens of flags a day, most of which are not actionable. Now watch what sensible behaviour looks like.

She does not read every alert carefully. She cannot. So she develops a shortcut, usually some version of “the software is right about the obvious ones and wrong about the borderline ones, and the borderline ones are what I will be asked about.” Then the ward gets busier, and the shortcut degrades into recognising the shape of an alert rather than reading its contents. This is not carelessness. It is the correct response to a situation where most alerts are wrong and the available time is fixed. Ask anyone who has worked a night shift beside a noisy monitor how long it takes to stop hearing it.

What you end up with is a system where the human is functioning as an approval mechanism rather than a judgment mechanism.

And the difference between those two is invisible in every metric a company is likely to be watching, because the approval rate looks superb and the override rate looks reassuringly low. Low override rates get shown to boards as evidence the model is performing well. They are at least as consistent with nobody having time to disagree.

Why it happens

Three forces, and none of them is anybody’s bad intentions.

Throughput pressure never lets up. Every health system on earth is trying to see more patients with the same number of clinicians or fewer. When a tool arrives promising efficiency, that efficiency gets banked immediately, usually as more patients per session. The time oversight needs was meant to come out of the savings. It rarely survives the first budget cycle.

Alert volume destroys discrimination. This is the best-documented failure in clinical informatics, and we keep rebuilding it with newer technology. Studies of prescription safety alerts have found override rates ranging from around half to well over ninety percent. When most alerts are wrong, dismissing them is the rational default, and that default cannot tell the difference between the wrong alerts and the one that was not.

The liability structure rewards the wrong thing. This is the part I find genuinely uncomfortable.

The presence of a human reviewer is precisely what lets a company say the clinician made the decision. It is the mechanism by which responsibility moves outward, away from the vendor. Which means there is a strong incentive for the review step to exist, and very little commercial incentive for it to be rigorous, because rigour costs throughput and throughput is what the customer came to buy.

Nobody says this out loud in a meeting. The incentive does its work regardless.

What real oversight would need

If we are going to keep using the phrase, it should come with conditions. Four, at minimum.

Time that is protected rather than assumed. If reviewing an output properly takes ninety seconds, the workflow has to contain ninety seconds. Which means accepting a throughput cost, and somebody senior defending that decision when volume targets get set.

What happened13 secondsWhat a proper review takes90 secondsper output reviewed
Thirteen seconds against ninety. Both figures are from this piece.

Training on how it fails, not on how it works. Clinicians do not need the architecture. They need to know when this particular system is likely to be wrong, which patients it underperforms on, and what it reliably misses. That is the only curriculum that produces a reviewer capable of disagreeing intelligently.

An interface where disagreeing is as easy as agreeing. If approving takes one click and overriding requires typing a justification, you have put a thumb on the scale, and you will get exactly the approval rate you engineered. Make the override one click and capture the reason as a structured field.

That override log, incidentally, is the most valuable dataset your product will ever generate, because it is a record of your model’s real-world failures written by the people who caught them.

A culture where overriding is unremarkable. If clinicians learn that disagreeing with the algorithm invites scrutiny while agreeing never does, that behaviour resolves in one direction very quickly and very permanently.

The question worth asking about your own product

Go and pull one number.

The median time between an output appearing on a clinician’s screen and that clinician acting on it. Median rather than average, because the average will be rescued by the handful of complex cases somebody genuinely worked through. Break it down by individual user, over the last ninety days.

Then look at what you are putting on that screen and ask honestly whether a person could read it, weigh it against the patient, and reach an independent view in that amount of time.

If the answer is no, there is no human in the loop. There is a human on the receipt.

And then the harder follow-up. Look at your approval rate. If clinicians are agreeing with your model more than about ninety-five percent of the time, there are two possible explanations. Either your model is performing at a level your own validation data almost certainly does not support, or your review step is not functioning as a review step.

Most teams have never looked. The number is usually sitting there, unowned, absent from every dashboard that reaches leadership. Which is itself the finding.

Where I land

None of this is an argument against clinical AI. In a district hospital with no radiologist, the alternative to imperfect software is not a specialist, it is nobody, and I would deploy the imperfect software without much hesitation.

The argument is that “doctor in the loop” has been carrying enormous weight without ever being asked to prove anything. It shows up in pitch decks, regulatory submissions, and press releases as a settled safeguard, and it is almost never accompanied by a number.

So attach one. Publish your median review time. Publish your override rate, and what you do when it falls. Treat oversight as a feature you built and can evidence, rather than a category you get to belong to.

Until then, the phrase is not a safeguard. It is a sentence that makes everyone in the room feel better, including the people who should be asking the next question.