20 August 2026
Steven Byrne
ISO/TS 24971-2 is a new international guidance for machine learning-enabled medical devices (MLMDs). In this two-part guide we’ll first outline its key considerations for MLMD risk management, and then look beyond it to consider how to apply risk management to LLM-enabled medical devices too.
Scarlet welcomes ISO/TS 24971-2 because it broadly aligns with the best practices we’ve seen in our assessments of MLMDs.
ISO/TS 24971-2 - What exactly is it?
ISO/TS 24971-2 is a technical specification that provides technical methods and examples and is aimed specifically at manufacturers of MLMDs. To date, manufacturers have built risk management documentation based on the guidance in ISO 14971, the risk management process standard, and ISO/TR 24971, a technical report that provides general guidance on applying ISO 14971. Both these existing publications are aimed at a broad array of medical devices and do not address unique considerations for MLMD development.
With the increasing development of AI-enabled medical devices, this has become a glaring gap between the published international guidance and the realities of actually developing innovative MLMDs. A gap that is now filled by ISO/TS 24971-2.
It is a valuable resource for MLMD manufacturers to incorporate into their state-of-the-art guidance for developing their products. It is worth noting that the specification does not apply to medical devices employing large language models or generative AI; it is specific to machine learning. We'll look at LLMs in the next blog though.
Which risks are specific to MLMD?
Your decision to use machine learning in your medical device introduces specific device characteristics and hazards that are not present in traditional medical devices. These must be thoroughly considered during your risk-analysis activities. To ensure that your risk analysis appropriately covers machine learning risks, you must have machine learning competencies within your team and consider the points in the machine learning lifecycle at which risks may arise.
The table below outlines some of the risks common to MLMDs.

ISO/TS 24971-2 Annexes B, C, & D provide questions, considerations, and examples that are worth consulting when you are analysing the risks specific to your MLMD.
ISO/TS 24971-2 devotes an entire annex (Annex A) to bias, on the basis that every MLMD carries some degree of it, so the task is to identify and manage bias rather than to claim its absence. Note also that bias is not always unwanted: where a device is intended for a specific patient population, training on that population is appropriate, provided users are told the device is not intended for others.
Estimating machine learning risks
Once identified, risks must be estimated and evaluated. When applying risk estimation scores to your machine learning-related risks, it is worth considering the following factors:
ISO/TS 24971-2 recognises that the probability of occurrence of harm cannot always be estimated for ML-related risks
Where that is the case, the risk should be estimated based solely on the severity of possible harm
This is consistent with IEC 62304, which assumes the probability of a software failure is 1 because it cannot be meaningfully estimated, so machine learning failure modes rooted in the software are scored solely on severity
Formative usability evaluations can be used to estimate the likelihood of usability-related risks
Assessment of data quality can inform the probability of risks arising from poor representation
Controlling machine learning risks
Under ISO 14971, manufacturers are required to consider risk-control options in the following order of priority: inherent safety design, protective measures, information for safety (including user training).
This holds true for MLMDs, where the best controls are high-quality training data, robust design of the machine learning algorithm/model, and validation of its performance. These are supported by protective measures, such as automated evaluation of model outputs and human oversight, and by information for safety, such as on-screen instructions, warnings, and alerts.
The table below provides some common examples of how to apply these risk-control options to MLMDs:

Where you rely on human oversight as a risk control, consider whether the level of autonomy leaves the user enough time to intervene, and whether the device makes clear how to do so. ISO/TS 24971-2 Annex D sets out levels of autonomy from manual support through to full autonomy; the higher the level, the harder both of these become.
Verifying the effectiveness of risk-control options is a key part of the risk management process, and this is particularly important for MLMDs, where inaccurate outputs and automation bias are a real risk. Your verification and validation activities should include test scenarios that simulate hazardous situations arising from machine learning failures, and evidence that these risks are detected and mitigated by your risk controls. For example, synthetic data may be used to simulate an error case and verify that invalid outputs are automatically evaluated and flagged.
Similarly, usability evaluation is a vital tool for verifying that your information for safety controls is effective. It can demonstrate that the software explains its outputs adequately, that users understand the on-screen instructions, warnings, and alerts, and that users can effectively override bad outputs.
When evaluating the overall residual risk, give particular thought to failures that occur silently, without any signal to the user. A device that fails without announcing it undermines the protective measures and safety information that the rest of your risk-control strategy relies on.
Monitoring risk post-market
Machine learning models are not static. Once deployed, they may be proactively updated with new real-world training data to improve performance or expand the model’s scope to include new use cases. They may also be updated reactively in response to incident reports or drift.
Drift is the gradual, unintended degradation in performance that occurs as the data seen in use diverges from the training data. Post-market monitoring of the model’s performance is therefore a critical part of the risk management process for MLMDs.
Where performance monitoring is necessary for safety, ISO/TS 24971-2 requires the risk management plan to define how that monitoring will be done and which data will be collected. The plan should also set out what triggers an update or retraining, and the provisions for rolling back to a previous version. The information collected should include:
Incident reports relating to model failures, errors, and misuse
Performance logs capturing the model’s performance over time and detecting any degradation or bias
User feedback relating to the model’s adoption, usability, and effectiveness
This information should be reviewed and analysed to identify trends and patterns that may indicate a risk. For example:
Consistently incorrect outputs requiring human intervention may point to a problem with the model’s design or training data
Performance that degrades on real-world data may indicate drift or unwanted bias
Incident reports showing that faulty outputs are not being overridden may suggest that usability-related risk controls are ineffective
Conclusion
Machine learning introduces unique challenges to medical device risk management. It requires a competent team that can comprehensively identify risks associated with their MLMD, can effectively control these risks using machine learning-specific risk-control methods, and can monitor and maintain the performance of the ML model post-market.
ISO/TS 24971-2 provides the foundational understanding needed for MLMD risk management with worked examples that were previously missing from the published guidance.
Next time, we will look beyond the scope of ISO/TS 24971-2 and consider how to apply risk management to LLM-enabled medical devices.