Medical Technology and Regulation · us
FDA Maps a Risk Framework for Generative AI Medical Devices, Proposes Incorporating Clinical Review into Oversight
From premarket benchmark testing to continuous postmarket monitoring, the U.S. regulator is developing tiered rules for medical systems that generate content and may even perform tasks autonomously; the public consultation will continue through October 19.
As generative artificial intelligence enters healthcare settings, the question is no longer simply whether its answers are accurate. When a system can draft diagnostic summaries, recommend courses of action, or even connect with other tools to perform tasks, a single plausible-looking erroneous output may be amplified as it moves through the clinical workflow. The U.S. Food and Drug Administration (FDA) has therefore launched a public consultation in an effort to translate the risks of such medical devices into regulatory requirements that can be reviewed and tracked.
The FDA has proposed a two-axis risk framework that stratifies products according to their clinical impact and the risks introduced by their generative functions. The specific assessment approach remains to be refined through consultation feedback, but the central direction is to ensure that evidence requirements are proportionate to the potential harm, rather than placing every product that uses a large model into the same category.
At the premarket stage, developers may be required to use predefined benchmark tests to evaluate system performance in representative scenarios. Such validation cannot focus solely on average accuracy; it must also address the variability specific to generative models: the same input may produce different answers, while rare cases and different populations may expose gaps not covered by the training data. The FDA announcement did not provide trial data for specific products or a uniform threshold, so questions remain about how benchmarks should be designed and which clinical populations the data should cover.
For higher-risk uses, the FDA has also identified “clinical confirmation” as a possible safeguard. If AI output will affect diagnosis or treatment, requiring a qualified healthcare professional to review it before action is taken could reduce the chance that an error reaches the patient directly. However, whether this safeguard is effective will still depend on interface design, users’ ability to identify errors, and whether busy workflows cause the review to become a mere formality.
Regulatory attention also extends beyond a product’s market launch. The FDA is considering risk-proportionate monitoring to track model updates, changes in input data, and performance drift in real-world clinical environments. This is especially important for generative systems: the underlying model may be supplied by a third party and updated continuously, potentially causing product behavior to change quietly without any obvious change in appearance.
The consultation also specifically addresses foundation models and agentic systems. Foundation models can be shared across multiple medical products, but may make training data, version control, and accountability less transparent. Agentic systems not only generate text or images, but may also plan steps, call tools, and execute tasks, allowing errors to expand from a single answer into a sequence of actions. Whether regulatory responsibility should rest with the model supplier, the medical device developer, or the deploying institution will be critical to the framework’s design.
The public comment period will close on October 19. What has been released at this stage is a regulatory direction, not finalized review standards. Without independent data sources on the same matter and complete technical documentation, it is not yet possible to determine the final evidence threshold or which products will be subject to the strictest requirements. However, the FDA has clearly communicated one principle: generative AI medical devices cannot simply pass a single test before market launch; their safety must be demonstrated continuously throughout the entire product lifecycle.