Skip to content
Miracle OlajuyigbePHYSICIAN · MEDICAL WRITER

AI and Medicine

What a Clinical Reviewer Looks For in AI-Generated Copy

By Miracle Olajuyigbe 5 min read


Two sentences. One word different.

Our algorithm flags suspicious lesions for radiologist review. Our algorithm detects suspicious lesions for radiologist review. The first describes a tool that marks something for a human to assess. The second describes a tool that finds something. In most markets, the first matches what imaging AI is actually cleared to do and the second makes a claim about intended use that your regulatory clearance probably does not support.

Your writer will not see a difference. Neither will your editor, your marketing lead, or the model that wrote it. Everyone in that chain is checking whether the sentence is fluent and roughly true, and by those standards both are fine.

This is the level at which AI-generated medical content actually fails. Not with hallucinated nonsense that anyone would catch. With one verb.

First, the thing I am not arguing

I use AI in my own process. It is genuinely good at structure, at first drafts, at reworking something into a different register, at the part of the job that is closer to typing than thinking.

The argument is narrower than “AI content is bad.” It is that fluency and accuracy are different properties, and these systems produce the first at a level that makes people stop checking for the second. A confidently written wrong sentence gets waved through in a way that a clumsy one does not.

In most categories that costs you some credibility. In healthcare it costs you more.

The verb ladder

This is where I start on any AI-assisted draft, because it is where the highest-consequence errors cluster. Each pair below moves from a defensible claim to one that is harder or impossible to support.

Flags → detects → identifies → diagnoses. A tool that marks a region for review is doing something quite different from a tool that determines what the region is. Diagnosis for most solid tumours happens on tissue, examined by a pathologist. Software that never touches tissue is not diagnosing anything.

Is indicated as an adjunct to → treats → cures. The first is what most labels say. The other two are what marketing copy drifts toward.

Reduces the risk of → prevents. Prevention is an absolute claim. Risk reduction is a statistical one. These are not synonyms, and the difference is the whole of the evidence.

Is associated with → causes. The most common single error in AI-generated health writing, and the model will make it while sounding entirely measured.

Supports → improves → optimises. Increasingly strong claims, each requiring more evidence than the last, and “optimises” usually requiring evidence nobody has.

May help with → helps with → is proven to. The last one needs a citation and often does not have one.

Run a find on the strong end of each of those. It takes ten minutes and it catches more real exposure than any other single check.

The other seven things I look for

Statistics with no denominator. “Studies show a 40 percent improvement.” In whom, compared to what, over what period, measured how. AI drops these qualifiers constantly, because the source sentence read more smoothly without them and smoothness is what it optimises for.

Relative risk dressed as absolute. “Doubles the risk” sounds enormous. If the underlying risk was two in ten thousand, it is not. Ask whether the number in front of you is a proportion or a change in a proportion, and whether a reader could tell.

Citations that are almost right. Real journal, real-sounding authors, wrong year, or a paper that exists but does not say what is claimed. This is the failure that most embarrasses companies, because it looks like fabrication rather than error. Check that every cited paper exists and says what the sentence claims. Both parts.

Research findings written as deployed capability. “AI can detect this years before symptoms appear” usually describes a retrospective study on curated data where the outcome was already known. That is a different activity from prospective detection in a clinic, and performance usually drops between them. Content has to say which one it is describing.

Hedges in the wrong place. This one is subtle and it matters. AI-generated health copy tends to state benefits flatly and hedge the warnings. “This treatment improves outcomes” followed by “some patients may experience side effects.” Flip the emphasis and look at what happens. Usually the hedging should be on the benefit and the warning should be the plain sentence.

Confidence on contested evidence. Where the literature genuinely disagrees, these models produce a clean, confident summary of one position, because the training data contained more of it. Remote monitoring in heart failure is a good example: some major trials positive, some negative. Any draft presenting that as settled has told you something about its source material.

Missing populations. A threshold, a normal range, or a screening cut-off presented as universal when it came from a specific study group. AI reproduces the default because the default dominates what it read.

The practical version

If you want something your team can run before anything reaches me:

  1. Search the draft for the strong verbs: detects, diagnoses, treats, prevents, cures, causes, proven, eliminates. Justify each one or downgrade it. 2. Every statistic gets a population and a comparator, or it comes out. 3. Every citation gets opened. Does the paper exist, and does it say this. 4. Every claim about capability gets labelled: study finding, or deployed product. 5. Read the benefit sentences and the risk sentences side by side. Check the hedging is on the right one. 6. Ask whether anything stated confidently is actually contested. That is maybe forty minutes on a 1,500-word page and it removes most of the risk.

What it will not remove is the last category, which is knowing what a claim commits you to under your specific clearance in your specific market. That part is not a checklist. It is the reason clinical review still has to happen.

Why this is worth building into the process

The usual pattern is that nobody looks until something forces them to. A claim in a two-year-old post surfaces during a funding round, a procurement review, or a conversation with a regulator, and then somebody audits eighty pages at once, under time pressure, with a lawyer on the call.

Catching it at the draft stage costs forty minutes. Catching it at the audit stage costs a fortnight and a great deal of goodwill.

The verbs are where I would start.