
FDA weighs competency-based oversight for GenAI medical devices, signaling regulatory shift for AI builders
Published by AINave Editorial • Reviewed by Ramit
The FDA is exploring a competency-based evaluation framework for generative AI medical devices, a shift that could fundamentally change how AI tools in clinical care are regulated. The agency published a discussion paper, first shared with Axios, outlining risk-based methods for preapproval and postmarket oversight of GenAI devices. For builders shipping AI into healthcare, this signals that the regulatory playbook is being rewritten.
FDA proposes competency-based evaluation for GenAI medical devices
The FDA's discussion paper raises considerations for both preapproval and postmarket regulation of AI-enabled devices and solicits stakeholder feedback. The agency notes that generative AI devices are substantially different from other regulated medical devices, particularly because of the variation in their outputs and because they evolve over time. The paper begins by outlining a method for evaluating the risk of a GenAI device, measuring the type of activity performed by a device against the severity of the consequences from incorrect outputs. One big question the paper poses is what standard GenAI devices should be compared against.
Why GenAI devices break the traditional regulatory mold
Traditional medical devices have fixed outputs and don't change after approval. GenAI devices, by contrast, produce variable outputs and can evolve in real time as models are updated or fine-tuned. This makes it hard to apply the same preapproval testing and postmarket surveillance used for static devices. The FDA's competency-based approach, similar to how doctors are evaluated, would assess the device's performance against the potential consequences of its mistakes, rather than against a fixed specification. For builders, this means that the regulatory burden may depend on the clinical risk of the use case, not just the model's benchmark scores.
How the risk-based framework would work
The FDA proposes mapping device activity to the severity of potential outcomes. For example, a GenAI tool that suggests treatment plans would face higher scrutiny than one that summarizes patient notes, because the consequences of an error are more severe. This risk-evaluation approach aims to adapt regulatory scrutiny to the dynamic nature of AI in clinical care. Builders should expect that the level of evidence required for approval will scale with the potential harm of incorrect outputs. The paper also asks what standard GenAI devices should be benchmarked against, leaving open whether comparisons will be to human clinicians, traditional devices, or some other baseline.
What's still unclear for developers
The discussion paper is just that: a discussion. No firm implementation timeline or specific requirements have been proposed. The FDA is soliciting stakeholder feedback, meaning the final framework could change significantly. Additionally, the evidence for this story comes from a single Axios report about the paper; the actual paper may contain more detail. Builders should monitor the FDA's public docket and consider submitting comments. The regulatory landscape for AI in medicine is evolving, and early engagement could shape the rules.





















