3 September 2026

Risk management for AI scribes and other generative AI medical devices

Risk management for AI scribes and other generative AI medical devices

Steven Byrne

Search…

Search…

In the previous post of this two-part guide, we looked at risk management for machine learning medical devices (MLMDs) based on traditional (non-generative) AI. But what about generative AI devices such as AI scribes? Below we set out how we think manufacturers should approach risk management for these devices.

What makes generative AI different?

Generative AI has some fundamentally different properties from traditional machine learning that shapes the risks it poses. 

  • The input and output spaces are effectively unbounded. Traditional machine learning maps an input to a point in a defined output space, such as a class, a measurement, or a mask. An AI scribe captures open-ended speech from an unpredictable clinical environment and produces new text sampled from a distribution.

  • Behaviour is probabilistic. With traditional machine learning, the same input data will always produce the same classification output. With an AI scribe, the same transcript will produce a different clinical note each time.

  • You (likely) did not train the model. Most generative AI devices are built on a third-party foundation model, trained on a vast array of internet data that the manufacturer neither owns nor can fully inspect. The training data, annotation controls and model validation activities that you control for an MLMD are largely unavailable for generative AI devices.

  • Failures are hard to see. A traditional (non-generative) MLMD fails by misclassifying within a known output space, and those failure modes can be enumerated and counted. A generative model fails by hallucinating or omitting details in prose that is fluent, plausible and indistinguishable from a correct answer. A scribe that fabricates a segment of patient history produces a note that reads exactly like a good note.

Start from the common ground

The newly published ISO/TS 24971-2 international guidance stops short of covering generative AI medical devices built with Large Language Models (LLMs). Despite this, much of the ISO/TS 24971-2 guidance is relevant to generative AI devices.

  • The requirement for competent personnel carries over directly. Your team needs people who understand the AI technology and the clinical workflow well enough to ensure a thorough risk management process.

  • Several risk categories transfer with little modification. Data shift, erroneous output generation, usability-related risks, and model deployment risks are all equally relevant for a generative device. Annexes B, C and D are worth working through as prompts for your own hazard identification, even though the examples are drawn from traditional (non-generative) models.

  • Bias is another risk addressed by  ISO/TS 24971-2. Annex A takes the position that every MLMD carries some degree of bias, so the task is to identify and manage it, rather than to claim its absence. That position holds for generative AI, with the added difficulty that you cannot characterise the training data. An AI scribe may handle certain accents, dialects or non-native speech markedly less well, and it may reproduce clinical language patterns that encode historical inequities. Identify these as hazards, test for them explicitly, and state the limitations you find.

  • The risk estimation guidance is also applicable. ISO/TS 24971-2 recognises that the probability of occurrence of harm cannot always be estimated for ML-related risks and that, where that is the case, risk should be estimated on the severity of possible harm alone. For generative AI, this will often be the case too.

  • Usability and human oversight guidance applies with little change. The hazards ISO/TS 24971-2 identifies for traditional MLMDs – automation fatigue, overtrust and the absence of a human intervention process – are also applicable to generative AI. What generative AI adds is a reason to expect overtrust rather than merely to guard against it. Fluent, well-structured output invites the reader to accept it, and the review burden of checking generated content against its source is greater than that of confirming or rejecting a flagged MLMD finding.

  • Post-market monitoring is also similar. Incident reports, performance logs and user feedback must be collected to evaluate the real-world performance of your device. For generative AI devices, signals of interest include: the edits clinicians make before accepting output, rates of rejection or regeneration, and flags raised by your automated checks. Change management is very important. Unlike MLMDs, where you set the cadence of model updates, for generative AI devices, you must monitor and react to the foundational model provider’s notices of model deprecation or change.

Further generative AI risks

In addition to the considerations above, we strongly recommend analysing the following additional risks in your generative AI risk file.


Risk control for generative AI

When planning risk controls for generative AI devices, ISO 14971’s required risk control priorities still apply: inherent safety design, then protective measures, then information for safety. 

A particular challenge of generative AI is implementation of ‘inherent safety design’ when you do not control the model. To address this, you should constrain the problem you give the model and build controls around it. Below are some common examples of risk controls for generative AI devices.


Verifying that your controls work

When verifying the effectiveness of your risk controls, you should consider the following kinds of testing:

  • Statistical evaluation: Because outputs are probabilistic, effectiveness should be demonstrated statistically, across an adequate evaluation set that represents your intended use, patient population, and operating conditions. This evaluation set should include many adversarial test scenarios that invoke the failure modes you wish to detect.

  • Usability evaluation: Similar to MLMDs, usability evaluation is key. Present suitable test participants with a series of example scenarios, both accurate and inaccurate, and observe their behaviour. It is the only way to establish whether clinicians will genuinely review generated content and spot mistakes, or defer to AI entirely.

Pulling it all together

When comparing MLMDs and generative AI devices, the overall risk management process remains unchanged, and there is substantial crossover in the risks you identify and the strategies you use to control them. What changes is how much of the system you control: the training data, the model design, the input and output spaces and the performance evaluation are looser or absent. That makes robust inherent safety design considerably harder to achieve and shifts weight onto protective measures, oversight, and vigilant post-market monitoring.

Manufacturers of generative AI medical devices should therefore:

  • Identify the risks common to all AI medical devices, using ISO/TS 24971-2 as a supporting resource.

  • Identify the risks specific to generative AI and to their own intended use.

  • Constrain the task and ground the outputs wherever possible.

  • Implement stringent protective measures around the model, and verify their effectiveness statistically and adversarially.

  • Design human oversight that clinicians can realistically exercise, and confirm through usability evaluation that they do.

  • Monitor performance in the field, treating clinician edits as a primary safety signal, and re-verify on every model or prompt change.

Need to certify an AI medical device?