Skip to content
Miracle OlajuyigbePHYSICIAN · MEDICAL WRITER

Medical Review

Medical Content Audit

By Miracle Olajuyigbe 7 min read


Start with the headline, because the headline is where the problem starts:

How AI Diagnoses Cancer With 99% Accuracy: The Technology That’s Replacing Radiologists Three claims, packed into fourteen words. AI diagnoses cancer. It does so with 99 percent accuracy. It is replacing radiologists.

None of the three survives contact with what the technology actually does. Two of them carry regulatory exposure. And all three appear, in some form, across a great deal of published health-tech content, which is why this piece is worth walking through properly.

The article reviewed here is a composite, assembled from claims that recur across live company blogs. Everything flagged below is something I have found on a real published page.

Asset: blog post, roughly 1,400 words, top of funnel, on a diagnostic imaging company’s site Reviewed for: clinical accuracy, regulatory exposure, claim substantiation

Summary of findings

#IssueSeverityAction
1Headline says AI “diagnoses” cancer and is “replacing radiologists”CriticalRewrite title and thesis
2“99% accuracy” with no prevalence, sensitivity or specificityCriticalReplace with figures in context
3Research findings written up as current clinical capabilityHighReframe with stage and setting
4No mention that cleared tools are approved as assistive onlyHighAdd a paragraph
5Training population and external validation not statedModerateAdd two sentences
6No mention of overdiagnosis in a screening contextModerateAdd short section

I would not leave this page up as it stands.

Issue 1: The word “diagnoses” is doing something the product cannot do

Why it is clinically wrong. For most solid tumours, cancer is diagnosed on tissue. A sample is taken, processed, examined under a microscope, and classified by a pathologist. Imaging finds something suspicious. It does not establish that the something is cancer.

A radiologist reading a mammogram is not diagnosing cancer either. They are describing a finding and assigning a level of suspicion, which determines what happens next.

What these AI tools do is detection and characterisation. They mark areas worth a second look, assign risk scores, and push urgent studies up the reading queue. That is genuinely useful, and it is not diagnosis. The distinction is not pedantry. It describes exactly where clinical and legal responsibility sits.

Why it is a regulatory problem. Imaging AI authorised in major markets is overwhelmingly cleared as assistive, meaning the output goes to a qualified reader who makes the call. Publishing content claiming the technology diagnoses cancer and replaces radiologists is making a claim about what the product is for that its clearance does not support.

Marketing copy is discoverable. “Our blog post overstated it” is not a position anyone wants to be arguing from.

Why it also costs you commercially. Your buyer is a radiology department. Radiologists read your marketing. A page announcing their replacement is not a lead magnet for that audience, and it is a gift to any competitor whose content shows they understand the workflow they are selling into.

Recommended title: “How AI Supports Cancer Detection in Radiology: What the Technology Does, and What Still Needs a Radiologist”

Issue 2: “99% accuracy” is the most misleading number in health-tech writing

Accuracy, quoted as a single figure, is almost meaningless in cancer screening. It is also the statistic that appears most often, which tells you something.

Cancer is rare in a screening population. Take a group where five people per thousand have it. Software that called every single scan normal, detecting nothing whatsoever, would be right 99.5 percent of the time.

It would be “99 percent accurate” and completely worthless.

What a reader needs instead are three things. How often it catches real cancers. How often it wrongly flags healthy people. And, most importantly in screening, of everyone it flags, what proportion actually has cancer.

That last one is where low-prevalence maths bites. Even a very specific tool will produce mostly false alarms when the disease is rare, and every false alarm is a real person getting a recall letter, another scan, sometimes a biopsy, and several weeks of fear. Honest writing about screening has to carry that cost, because it is the cost patients actually experience. The article also gives its figure with no source, no population, and nothing to compare against. Three questions it should answer and does not: measured against what, in whom, and compared to which readers.

Issue 3: Studies described as if they were products

The article says AI can detect cancer “years before symptoms appear.”

Papers looking back at old scans of people later diagnosed with cancer do exist, and they are interesting. They are not a description of what a deployed product does today.

Finding something in a curated dataset when you already know the answer is a different exercise from finding it prospectively in routine practice. Performance typically drops, sometimes a lot, when a model meets a population it was not developed on. Content describing capability has to separate what a study showed from what is running in clinics, and say which is which.

Earlier detection also raises overdiagnosis, which the article never mentions. Finding more cancer is not the same as finding more cancer that would ever have harmed anyone. In screening, that is not a footnote. It is the substance.

Rewritten sections

Commentary is cheap. Here is the editing.

Before

AI is revolutionising cancer diagnosis. Deep learning algorithms can now analyse medical images with 99% accuracy, outperforming human radiologists and detecting tumours that doctors miss. This technology is transforming healthcare as we know it, and experts predict that AI will replace radiologists within the next decade.

After

AI is changing how cancer gets detected, though not in the way the headlines suggest. In large studies, deep learning models reading screening mammograms have matched or slightly bettered the detection rate of a single radiologist, and trials within European screening programmes have reported comparable results with a lighter reading workload. What these systems do is flag suspicious areas and score how concerning they look. A radiologist then decides what happens next, and if something needs to be confirmed as cancer, that call is made on tissue by a pathologist. The most credible near-term change is not replacement. It is sorting: software that orders a worklist so the scans most likely to matter get read first, and that offers a second opinion where no second reader exists.

Three changes worth pointing at. The unsourced accuracy figure is gone, replaced by what was actually measured and against whom. The claim is narrowed to one type of scan in one setting rather than “medical images” in general. And the argument moves from replacement to sorting, which is both accurate and the stronger commercial position, because sorting is what radiology departments are actively trying to buy.

Before

Because AI never gets tired and never has a bad day, it delivers consistent results every single time. Studies show AI can identify cancerous lesions years before they would be detected by conventional screening, potentially saving millions of lives.

After

Consistency is a real advantage, and it is worth being precise about why. A radiologist reading their two-hundredth scan of the day is working under a fatigue burden that is well documented. Software reads the last scan of a session exactly as it read the first. But consistency is not the same as accuracy. A model can be reliably wrong, and it fails in patterns rather than at random, which makes those failures harder to spot. Where a tired reader misses things unpredictably, software may miss one particular presentation, in one particular tissue type, in one particular group of patients, every single time, and generate a tidy audit trail while doing it. This is why a tool should be tested on your own patient population before it goes live, and why monitoring its performance afterwards is not optional.

The rewrite keeps the benefit the client is entitled to claim and adds the failure mode. It reads as more authoritative, not less, because it demonstrates knowledge of how these systems break.

Before

With AI-powered screening, cancer can be caught at stage 1 instead of stage 3, dramatically improving survival rates and reducing treatment costs.

After

Catching cancer earlier generally does improve outcomes, and it is the central argument for screening anything. It comes with a complication worth stating plainly. Some cancers found by screening would never have grown enough to cause symptoms or shorten a life. Treating those causes harm with no benefit attached. It also flatters the statistics, because a cancer found earlier appears to have been survived for longer even when the date of death has not moved at all. Screening programmes are designed and judged with these effects in mind, so a tool that raises detection rates has to be assessed on whether it finds more of the cancers that matter, not simply more cancer.

What I would do with the rest of the page

Add a short paragraph explaining that cleared tools are approved to assist a reader rather than to replace one. Name the population behind any performance figure quoted. Add two sentences on how performance is checked after deployment.

That last one is a sales asset in disguise. Every serious buyer asks about it, and almost no competitor content answers it.

Estimated effort for the full revision: half a day.

How this works as an ongoing arrangement

Most companies find this problem the expensive way. A claim in a two-year-old blog post surfaces during a funding round, a regulatory conversation, or a procurement review, and somebody spends a bad fortnight going through eighty pages at once.

The alternative is duller and much cheaper. New content gets a clinical read before it publishes, which catches most issues at the point they cost almost nothing to fix. The existing library gets worked through on a rolling schedule, highest-risk pages first. And the errors that keep recurring get written into a short claims guide your writers work from, so the same handful of mistakes stops being made.

The output looks like what you have just read. Specific findings, severity ratings, and rewritten copy you can actually publish, rather than notes telling your team to go and fix it themselves.