NHS AI scribes are getting drug names and diagnoses wrong, the watchdog says
Twenty-seven scribe systems are already in use across the health service. Documented errors include drug name confusion, missing prescription advice and a fabricated diagnosis.
AI scribes used by NHS doctors are producing wrong drug names, missing prescription advice and, in at least one case, a fabricated diagnosis, according to warnings reported by the Guardian drawing on Healthwatch England and clinical specialists. Twenty-seven such systems are already in use across the health service.
One documented error was a record of "null demyelination" — a diagnosis that does not exist, entered into a patient's notes as though it did.
Why the easy task is failing
Ambient scribing is supposed to be the safe application. The model listens to a consultation and drafts a note; the clinician reviews and signs it. There is a human in the loop by design, the task is transcription rather than judgement, and the productivity case is real — clinical documentation is a substantial and widely resented share of a doctor's day.
The failure mode is not that transcription is hard. It is that drug names are dense, similar-sounding and consequential, and that a model producing fluent clinical prose will produce a plausible term where it heard something unclear. "Null demyelination" is what a system does when it fills a gap rather than flagging one.
A transcription tool that says "inaudible" is useful. One that guesses is dangerous, precisely because the guess reads like the rest of the note.
The review that is supposed to catch it
The safeguard is clinician review before signing. The watchdog's warning raises the question of what that review actually consists of at the end of a full clinic list, presented with a fluent, correctly formatted, professional-sounding note.
That is the same problem the Dutch data protection authority valued at €825 million last month, when it fined Uber over automated driver deactivations conducted without meaningful human intervention. The regulator's finding was not that no human was present. It was that the human confirmed rather than decided.
A doctor approving an AI-drafted note under time pressure is in the same structural position, and the consequences are entered into a permanent medical record that other clinicians will later rely on.
The unresolved questions are administrative, not technical
The reporting identifies two that nobody has answered: who is responsible for clinical review, and how errors already written into patient records get corrected.
The second is harder than it sounds. A medical record is a legal document with an audit trail; it is not simply edited. An erroneous diagnosis propagates — into referrals, letters, insurance and other clinicians' reasoning — and the correction has to chase it.
With 27 systems in use, there is no single vendor to hold accountable, no common error-reporting mechanism, and no published rate at which any of them get things wrong.
The context this lands in
OpenAI connected ChatGPT Health to Epic's records for 325 million patients on 3 September, with read-only access to clinical notes, lab results and medications.
That product interprets records rather than writing them, which is safer in one direction and less safe in another: the reader is the patient, not a clinician who can recognise that a diagnosis does not exist.
The NHS experience is the closest available evidence about how these systems behave in real clinical use, and it says that the errors are subtle, fluent and hard to catch. Neither OpenAI nor Epic has published an evaluation of interpretation accuracy for the population they have just switched on.
Runs the newsroom. Rename this profile in the studio to your own byline.
Related
Every weekday, the AI stories that moved money or shipped code.
No cross-posting, unsubscribe anytime. See all newsletters