Business, energy, technology, markets and global industry news from Business News Today
Features & Analysis

How should FDA regulate a medical AI system that can generate a different answer every time?

Traditional medical-device validation assumes a relatively stable product. A blood analyzer receives a sample and performs a defined measurement, while conventional software executes a known algorithm against defined inputs. Generative artificial intelligence creates a more difficult regulatory object because its response can vary according to wording, context, model updates and prior interactions, and future agentic systems may perform sequences of actions rather than produce one isolated prediction. In August 2026, the US Food and Drug Administration opened a public discussion on whether those characteristics require a regulatory framework materially different from the one developed for conventional device software.

FDA’s discussion paper explores a two-axis risk framework, competency-based premarket evaluation, risk-proportionate postmarket monitoring and specific considerations for foundation models and agentic AI. The agency emphasizes that the document is not draft or final guidance and does not establish new regulatory requirements; it is a request for stakeholder input intended to inform future policy. Comments are being accepted through October 19, 2026.

Why does generative AI fit poorly into the traditional medical-device testing model?

A conventional diagnostic algorithm can often be evaluated against a defined dataset using sensitivity, specificity or another fixed performance measure. If the same input repeatedly generates the same output, regulators and manufacturers can characterize failure modes relatively directly.

Large language models and other generative systems are probabilistic. Small differences in instructions or contextual information may change the response, and the same underlying foundation model can perform several clinical tasks ranging from summarizing a record to recommending a diagnosis or orchestrating actions across connected tools. That makes the intended use, user interaction and potential consequence of an incorrect output central to determining risk.

A system that drafts a discharge letter for clinician review does not create the same risk as an autonomous agent that interprets symptoms, orders tests and recommends medication without meaningful human verification. FDA is therefore exploring whether regulation should evaluate not only model architecture but the task being performed and the severity of harm that could follow from failure.

What does FDA mean by competency-based evaluation?

The agency’s discussion paper considers an approach inspired at a high level by how physicians demonstrate competence. Instead of relying exclusively on one aggregate benchmark score, a medical GenAI system might have to demonstrate acceptable performance across representative tasks, clinical situations and failure conditions relevant to its intended use, followed by clinical confirmation that it performs adequately in the real environment.

This approach recognizes that a clinical assistant may need multiple capabilities simultaneously. A system used in cardiology might need to extract information accurately from records, reason across contradictory test results, identify emergencies, recognize when information is insufficient and communicate uncertainty appropriately. Excellent performance on routine cases could conceal dangerous weaknesses in rare but consequential scenarios.

Competency assessment could therefore emphasize breadth, stress testing and boundary recognition rather than asking whether an AI achieved one attractive average accuracy number. The unresolved issue is how regulators define the test. Physicians undergo years of standardized education and supervised clinical exposure; a foundation model can change dramatically between software versions and may possess capabilities its manufacturer did not explicitly program.

FDA is exploring competency assessment, risk-based monitoring and new lifecycle controls for generative AI medical devices, addressing a regulatory challenge created by outputs that can change with prompts, clinical context and evolving software. Representative image.
FDA is exploring competency assessment, risk-based monitoring and new lifecycle controls for generative AI medical devices, addressing a regulatory challenge created by outputs that can change with prompts, clinical context and evolving software. Representative image.

Why is postmarket monitoring more important for GenAI than for ordinary software?

Generative systems can encounter real-world inputs that are impossible to reproduce comprehensively during premarket testing. Patient language, incomplete records, local medical terminology and unusual combinations of disease can expose behaviors that were rare or absent in development datasets.

Performance can also shift as the surrounding healthcare environment changes. New medical knowledge appears, clinical guidelines evolve and the populations using a product may differ from the population in which it was initially tested. If a manufacturer updates the underlying model, new capabilities may emerge while old ones degrade.

This is why regulatory researchers have argued for a total-product-lifecycle approach in which postmarket monitoring is not treated as an afterthought but as an integral component of GenAI device governance. Potential mechanisms include ongoing performance surveillance, structured adverse-event reporting, version tracking and predefined thresholds for investigating deterioration.

Can a manufacturer update a medical AI model without returning to FDA every time?

FDA has already been developing predetermined change control plans for AI-enabled device software. The principle is that a manufacturer can prospectively describe certain modifications, how those changes will be developed and validated and what controls will ensure safety, potentially allowing specified updates without requiring an entirely new marketing submission for every improvement. FDA finalized guidance on such predetermined plans in 2025.

Generative AI makes that framework harder because foundation-model updates can produce broad emergent changes rather than altering one isolated performance parameter. Updating the model to improve language reasoning might unexpectedly affect clinical calculation, summarization or hallucination behavior.

A regulator therefore has to decide how wide a pre-authorized change can be before the device effectively becomes a different medical product. Too little flexibility could freeze medical AI into obsolete versions; too much could allow materially altered software to reach patients without sufficient independent review.

What additional problem do foundation models create?

Many medical products may eventually be built on foundation models supplied by another company. A medical-device manufacturer could design a clinical interface and workflow while depending on an underlying model it did not train and cannot fully inspect.

If the foundation-model provider updates that model, downstream regulated products could change even when the medical-device company has not modified its own application code. This creates questions about supplier controls, version locking, contractual responsibilities and how manufacturers establish that a third-party model remains fit for the medical claim built on top of it.

Foundation models can also perform tasks outside their authorized medical use. A clinician might discover that a regulated documentation assistant can also answer treatment questions, creating what amounts to off-label software behavior without a clear physical barrier preventing use.

Why is agentic AI a larger regulatory leap than a chatbot?

A chatbot primarily provides information. An agent can potentially plan a task, call external software, retrieve records, order services or execute a sequence of steps without requiring a human to approve every intermediate action.

In medicine, that difference is profound. An AI that incorrectly summarizes one laboratory result can be caught by a clinician; an agent that interprets that result, changes a workflow and triggers another action can propagate one error across multiple systems before a human notices.

FDA’s discussion paper therefore explicitly raises agentic AI as an area requiring special consideration. Risk may depend not only on whether an answer is correct but on what authority the system has to act on that answer.

Could stricter GenAI regulation slow useful medical innovation?

It could if regulatory expectations become so expensive or rigid that only the largest technology companies can comply. Healthcare systems already face clinician shortages and administrative workload, and generative systems could potentially automate documentation, improve access to medical knowledge and help clinicians manage increasingly complex records.

The opposite risk is more serious: weak oversight could allow confident but incorrect systems to scale errors much faster than individual humans can produce them. Medical-device regulation therefore has to permit continuous software improvement while establishing evidence standards appropriate to products operating in safety-critical environments.

The hardest problem is that static validation and dynamic intelligence are inherently uncomfortable partners. A regulator wants to know exactly what product has been tested, while a developer wants the model to keep learning and improving.

FDA’s 2026 discussion does not resolve that contradiction. It does something arguably more important at this stage by identifying the regulatory unit that may need to change. Future medical AI may not be evaluated solely as a piece of software frozen at submission; it may have to be evaluated as a continuing capability whose competence, update process and real-world behavior all form part of the medical device.

Leave a Reply

Your email address will not be published. Required fields are marked *