The question that decides everything else
Most conversations about regulating clinical AI start in the wrong place. They start with which pathway a product will take, how much evidence a submission needs, and whether the model is explainable enough to survive review. All of that is downstream of a question almost nobody asks first, and the answer to it is not a property of the model.
The question is whether the software is a medical device at all. The 21st Century Cures Act carved five categories of software function out of the device definition in 2016, and one of them covers clinical decision support. A function inside that carve-out is not regulated by the FDA, needs no submission, and has no pathway to choose. A function outside it needs all three.
What decides which side you are on is intended use. Section 201(h) defines a device by what it is intended for, 21 CFR 801.4 reads that intent out of the labelling, the advertising and the circumstances of sale, and every one of the four criteria below is phrased as something the software is intended to do. Not the architecture, not the accuracy, not whether anyone can explain it. Two teams can build the same model and land on opposite sides of the line, and the thing that moved them is written in the labelling rather than in the code.
And the FDA is one of five authorities, not the whole of the question. The other four attach to the organisation running the software rather than to the one that wrote it, which is why they are so often discovered by a customer.
The whole territory, six domains deep, is below. Pick the one you arrived for; the rest of this piece walks them in order.
THE STACK · SIX DOMAINS
Pick the one you arrived for. Every figure is from the statute, the regulation, the guidance, the docket or the published study, and dates are as at October 2026.
What takes clinical software outside the device definition?
5
Cures Act exclusions
4
criteria for the CDS one
6 Jan 2026
guidance replaced
Criterion 1 is a test on the input
A medical image, a signal from an in vitro diagnostic, or a pattern from a signal acquisition system. A pattern is multiple, sequential or repeated measurements, so hourly observations out of a chart qualify and vitals from separate encounters generally do not.
The exclusion stops at the clinician
Criteria 3 and 4 are written about a health care professional. Patient-facing software is outside the exclusion entirely, however transparent it is.
New: one clinically appropriate output
Where only one option is clinically appropriate and every other criterion is met, FDA intends to exercise enforcement discretion. That is a device with a policy, not an exclusion.
Once it is a device, which way does it reach the market?
96.2%
510(k) · 1,553 entries
2.5%
De Novo · 40
1.3%
PMA · 21
Substantial equivalence is a comparison
The 510(k) asks whether a device resembles one already legally marketed, not whether the model works on your patients. The useful question is never whether something is cleared, but cleared against what, and on whose data.
De Novo creates codified special controls
Which is why two district courts have now found state product liability claims against De Novo-authorised Class II devices expressly preempted. One device family, no appellate decision, parallel claims unaffected.
Class does not follow risk
An interoperable insulin controller is Class II through De Novo; an integrated closed-loop system from another manufacturer is Class III through PMA. Same autonomy, different architecture, different class.
How is a model that retrains handled after authorisation?
515C
the statute behind it
3
parts of a plan
34 of 1,080
radiology devices carrying one
Description, protocol, impact assessment
A predetermined change control plan pre-authorises a bounded, enumerated list of modifications. A change outside the envelope still needs a new submission, so the boundary is the whole value of the plan, and the sponsor writes it.
The guidance about the model is still a draft
The January 2025 lifecycle guidance covers data management, model development, validation and performance monitoring. Twenty-one months on it is unfinished, and CDRH’s fiscal 2026 agenda places it on the B-list.
A trigger you cannot observe is not a trigger
A performance threshold is only detectable if the fall exceeds the uncertainty of the measurement at the volume being monitored. Decide the fall that matters clinically, then check the volume can see it.
What does the system around the software have to look like?
2 Feb 2026
QMSR in force
2 of 15
old subparts kept
29 Mar 2023
SBOM duty began
ISO 13485 by reference
Part 820 now incorporates the standard, keeping supplemental provisions for records and for labelling and packaging. Design control is where a model’s lineage has to live: an artefact with no traceable link to its training data is a design output with no design history.
An inadequate SBOM stops the submission
Section 524B makes a software bill of materials, vulnerability disclosure and patching a condition of a cyber device submission. Failure is grounds for a refuse-to-accept decision rather than a review question.
Documentation level is assessed before risk controls
Enhanced documentation turns on what the software could do if it failed, considered prior to mitigations. The categorical triggers for Class III and combination products are recommendations a sponsor may rebut with a rationale.
What can the postmarket system actually see?
11 Mar 2026
AEMS replaced MAUDE
0.63
validated AUC against 0.76–0.83 claimed
18%
of admissions alerted on
AEMS is a database, not telemetry
It consolidates MAUDE and six other legacy systems into one reporting platform. It does not ingest model performance streams from health systems and it does not compute population-level variance.
Event reporting cannot see drift
A model wrong 10% of the time is working as specified on any individual case, so no case is a malfunction, and gradual degradation across a population produces no report. Changing the database does not change that.
The obligations that bind are elsewhere
45 CFR 92.210 has bound providers since May 2025, and the predictive decision support criterion binds certified health IT. Both ask for subgroup performance information only the vendor has.
What happens to a model that writes, and who pays for any of it?
18 Aug 2026
discussion paper
23 Dec 2025
first patient-facing LLM cleared
~$41
CPT 92229 national rate, 2024
Two axes, four named steps
Independence runs from non-directive information through action-directing information and supervised action to fully autonomous action. Consequence runs limited to severe. Evidence scales with position, not with the presence of a language model.
The clearance came first
K253281 cleared device software with a patient-facing large language model by substantial equivalence to a 2018 insulin titration system, eight months before the paper asking how to regulate the category. The generative layer is confined; the dosing logic is deterministic.
Payment classifies by output, not technique
CPT Appendix S sorts AI services as assistive, augmentative or autonomous. For hospital outpatient work CMS has proposed a Software as a Medical Service category; that rule is not final.
What FDA rewrote in January
This got harder to keep straight because the rules changed recently and most of what is written about them predates the change. On 6 January 2026, re-issued on 29 January, FDA published a new final Clinical Decision Support Software guidance superseding the September 2022 one. A new General Wellness guidance landed the same day, replacing the 2019 version.
The four statutory criteria did not move. FDA’s reading of three of them did, and the changes are not cosmetic.
| Criterion | What the statute says | What is now different |
|---|---|---|
| 1 · Inputs | Not intended to acquire, process or analyze a medical image, a signal from an in vitro diagnostic, or a pattern or signal from a signal acquisition system | A pattern is defined as multiple, sequential or repeated measurements. Discrete, episodic point-in-time measurements, such as vitals at separate encounters, generally are not one |
| 2 · Medical information | Displaying, analyzing or printing medical information about a patient or other medical information | Read widely: demographics, symptoms, test results, discharge summaries, guidelines, studies, labeling. The test is whether each input’s relevance comes from accepted sources |
| 3 · Recommendations | Supporting or providing recommendations to a health care professional | The time-critical clause was removed from this criterion. And a single output no longer always fails it |
| 4 · Independent review | Enabling the professional to independently review the basis, so as not to rely primarily on the recommendation | Now a specification: intended use and user, required inputs, a plain language account of development and validation including the data and the clinical results, and the knowns and unknowns in the output |
Criterion 1 is the one that catches people, and it catches them on data they already have. FDA’s worked examples include software that analyses hourly pulse oximetry and heart rate pulled from the chart, and software that reads glucose from a continuous monitor every thirty minutes. Both are devices because both analyse a pattern. The data never touched a sensor the software controls. It came out of a record, and it is still a pattern, because a pattern is defined by the repetition rather than by the source.
Criterion 4 is where machine learning teams should spend their time. FDA now asks for a summary of the general approach, naming AI and ML techniques as the example, a description of the data relied on, and the results of clinical validation studies, and states that this applies regardless of the complexity of the software and whether or not it is proprietary. Commercial confidentiality is not an answer to criterion 4. It is a reason the function is a device.
How little it takes to move the line is best shown by FDA’s own triad. One cardiovascular risk function, three variants, and the model is identical in all of them.
Check your own product against it
The panel below walks the four criteria with a gate in front of them, and every option in it is drawn from the guidance or from one of its worked examples. It answers one question and no others: is this software a device. Nothing you select leaves the page.
IS IT A DEVICE? · FDA FINAL GUIDANCE, JANUARY 2026
Pick what fits your software. Every option is drawn from the guidance or one of its worked examples. Nothing you answer leaves this page.
Non-device CDS
Outside the device definition
Not a device
All four criteria are met, so the function is outside the device definition. It is not outside everything else: a provider using it still has to find and mitigate discrimination risk under 45 CFR 92.210, and inside certified health IT the predictive decision support criterion still asks for the source attributes.
The exception nobody expected
The most consequential change in January is the easiest to miss. FDA restates that a function giving a specific preventive, diagnostic or treatment output fails criterion 3, and then adds that where only one option is clinically appropriate, and the function otherwise meets every criterion, it intends to exercise enforcement discretion.
That dissolves a drafting game that had become standard practice. Teams were padding outputs with alternatives no clinician would pick, so that a single recommendation could be presented as a list. FDA’s own examples now bless the honest version: an antibiotic selector naming a specific agent, a treatment plan a clinician reviews and finalises, a differential that collapses to one diagnosis when the alternatives are improbable.
It is a third answer, not a second one. Enforcement discretion is not an exclusion. The software is a device in law and FDA has published an intention not to enforce. There is no letter to show a buyer, the policy can be withdrawn, and there is no device-specific federal requirement for a state law claim to conflict with. Treating a discretion policy as equivalent to a clearance is a mistake with three different ways of surfacing.
The limits are the useful part. Each of FDA’s discretion examples is paired with a variant that stays regulated, and the variants break on inputs or on urgency, never on the single output. Add PET images to the cognitive impairment planner and criterion 1 fails. Move the back pain pathway classifier from chronic pain to acute trauma and criterion 4 fails. The discretion is for the shape of the output. It is never for what goes in, and it is never for a decision that has to be made now.
The wall at the clinician
Criterion 3 says recommendations to a health care professional, and criterion 4 is written about the same person. Software labelled for a patient, a carer or any unlicensed user cannot satisfy either, so the clinical decision support exclusion is unavailable to it. No amount of transparency changes that.
FDA’s example is blunt: software helping a person with diabetes calculate a bolus insulin dose is a device, failing criterion 3 because it is not intended for a professional and criterion 4 because the decision is time-critical. The practical consequence is that a clinician-facing function and a patient-facing version of the same function are different regulatory objects. Opening the clinician view to patients is not a feature flag. It is a new intended use and most likely a submission.
Almost everything goes through a predicate
Once software is a device, the route is overwhelmingly one route. FDA publishes an AI-Enabled Medical Device List and cautions that it is not comprehensive, which is worth repeating whenever a number from it is used as a denominator. At 5 September 2026 it held 1,614 entries.
Only one of those three questions is about the device in front of the reviewer, and it is the one asked twenty-one times. Substantial equivalence is a comparison, and a comparison does not require prospective evidence. For most software it means retrospective testing on held-out data and a bench comparison against the predicate’s claims. That is a defensible standard for a device that genuinely resembles its predicate, and it is why the useful procurement question is never whether something is cleared but cleared against what, and on whose data.
The De Novo route has acquired a second consequence since 2024, in litigation rather than in regulation. Because a De Novo grant creates special controls codified in 21 CFR, two district courts have now held that state product liability claims against a Class II device authorised that way are expressly preempted, which 510(k)-cleared devices have never enjoyed. That is two district courts about one device family, with no appellate decision and with parallel claims surviving on their own terms. It is not settled law. It is also no longer irrelevant to choosing a pathway.
A worked example of how little the class tells you. An algorithm that predicts glucose and commands an insulin pump is about as consequential as clinical software gets. Medtronic’s MiniMed 780G reached the market as a Class III device under a PMA in April 2023. Tandem’s Control-IQ reached it as a Class II device with special controls, through De Novo, in December 2019, because FDA had split the closed loop into three separately classified interoperable components. Same clinical function, same autonomy, different architecture, different class.
Three classes, and two categories that are not classes
Class follows the risk the device poses and the controls needed to manage it, and for software it is less informative than people expect.
| Class | Controls | Usual route | Representative AI software |
|---|---|---|---|
| I · low | General controls | Mostly 510(k) exempt | Software that organises or displays health data without analysis that affects a clinical decision |
| II · moderate | General and special controls | 510(k), or De Novo where no predicate exists | Nearly all of it: imaging detection and triage, retinopathy screening, deterioration and sepsis prediction, the interoperable insulin controller |
| III · high | General controls and premarket approval | PMA | Integrated automated insulin delivery, algorithms embedded in life-sustaining implants |
Class I is not a synonym for exempt. Most Class I devices are exempt from premarket notification, but the exemption is listed per classification regulation and some Class I devices still need a 510(k). The answer for a product code is in the regulation, not in the class.
And the IMDRF categories are not FDA classes. The international framework in document N12 sorts software on two axes: the significance of the information it provides, from informing clinical management through driving it to treating or diagnosing; and the state of the healthcare situation, from non-serious through serious to critical. That grid gives categories I to IV. It is a genuinely useful way to think about risk and it does not map onto 21 CFR classes. A Category IV piece of software is not thereby a Class III device, and writing "Category III" for a reviewer who is thinking about Class III starts a conversation nobody wants.
How much documentation, and the test that comes before your mitigations
The June 2023 final guidance on the content of premarket submissions replaced the old three-step Level of Concern with two levels, and the way the line is drawn catches teams out.
Enhanced documentation is recommended wherever a failure or flaw of any device software function could present a hazardous situation with a probable risk of death or serious injury, to a patient, to a user, or to others in the environment of use. Basic covers everything else. Three features of that test matter more than the definition.
- The risk is assessed before risk controls are applied. A team that has mitigated a hazard to an acceptable level has not moved itself into basic documentation. The question is what the software could do if it failed, considered prior to the mitigations.
- Foreseeable misuse and cybersecurity are inside the assessment. The guidance names the likelihood that functionality is compromised by inadequate device cybersecurity as part of the determination, which ties the security work to the submission rather than leaving it a parallel workstream.
- The categorical triggers are recommendations, not rules. FDA generally recommends enhanced for Class III devices and for the device constituent part of a combination product, and for several blood-related categories. For the first two a sponsor may determine that basic applies and supply a detailed rationale. Published summaries calling enhanced documentation strictly required for Class III are overstating it.
At basic level the system-level test protocols and reports may be summarised. At enhanced level they are expected in full, with unit and integration testing, architecture and design documentation, and a record of unresolved anomalies assessed for clinical impact. For an AI device that last item is the uncomfortable one, because an honest list includes known failure modes of the model that no amount of engineering will remove.
The guidance about the model is still a draft
On 7 January 2025 FDA published a draft guidance on AI-enabled device software functions, covering data management, model description and development, validation, performance monitoring, transparency and bias across the lifecycle. It is the most complete statement the agency has made about what it wants from a medical AI programme.
Twenty-one months later it is still a draft, and CDRH’s fiscal 2026 guidance agenda puts finalising it on the B-list, to be done as resources permit, below eight higher-priority documents. Meanwhile the quality system regulation was replaced on 2 February 2026, the adverse event system was replaced on 11 March 2026, the cybersecurity requirements were finalised in June 2025, and the predetermined change control plan guidance was finalised in December 2024. Everything around the model is settled. The document about the model is the one that is not finished.
Build to the draft anyway, because it is the clearest available statement of what a reviewer expects and the alternative is guessing. Just do not cite it as authority, and treat anything expensive and specific in it as a bet that can still move.
Its most portable idea is worth taking regardless. The draft recommends conveying performance through something like a model card, with metrics reported separately across demographic subgroups. The ASTP/ONC predictive decision support criterion asks a certified EHR for substantially the same information from the other direction, and a hospital’s own non-discrimination obligation needs it too. One artefact satisfies three demands, which makes it the document in an AI regulatory file that repays the most effort.
Ten principles, and three parts of a plan
Two instruments sit where the draft guidance will eventually go. Neither binds, and both are what a reviewer has in mind.
Good Machine Learning Practice is ten principles published jointly by FDA, Health Canada and the UK MHRA in October 2021, endorsed by the IMDRF. They were followed by five principles for change control plans in October 2023 and a set on transparency in June 2024, which extends the audience beyond the clinician to patients and payers.
The one most often read as a platitude is the seventh. Where a device is used with a clinician, the evidence should be about the clinician’s performance when assisted, not the model’s alone. A model more accurate than the clinician it assists can still make the pair worse through automation bias, and a study design that measures only the model cannot see that happening.
The change control plan is the one mechanism built for a model that moves. Section 3308 of the 2022 omnibus added section 515C to the Act, giving FDA express authority to clear or approve a plan; a change made within an authorised plan needs no new submission. The AI-specific guidance was finalised on 4 December 2024, and it renamed machine learning device software functions to AI-enabled device software functions, widening its own reach.
| Part | What it has to contain | How it fails |
|---|---|---|
| Description of modifications | A specific, bounded, enumerated list of intended changes with the limit of each stated | Written as a category. "Periodic retraining as data becomes available" is not a modification description |
| Modification protocol | The data management, retraining and evaluation procedures, and the verification and validation for each enumerated change | No acceptance criterion stated in advance, so the protocol cannot fail; or no stated way to detect, halt and revert a change that degrades performance |
| Impact assessment | How the planned changes alter benefit and risk, and the risk controls for each, inside the device’s risk management process | Treated as a summary of the other two rather than a risk analysis of the act of changing |
Two things about plans in practice. The envelope’s boundary is the whole value of the plan and the sponsor writes it, so a change outside it still needs a submission. And the mechanism is barely used: of 1,080 radiology AI submissions listed by FDA between 2015 and 2025, 34 carried an AI-specific plan, and three of those documented a postmarket surveillance plan.
There is an arithmetic trap in the monitoring half, and it has cost real programmes real time. A trigger expressed as a fall in a performance measure is only observable if the fall exceeds the uncertainty of the measurement at the volume being monitored. A concordance rate computed on a few hundred cases a month carries an interval of a couple of points either side, so a trigger written at two points fires on the calendar rather than on the model. Decide the fall that would matter clinically, work out the volume needed to see it, and only then write a number into a document a regulator will hold you to.
The system around the software
A model does not ship on its own. It ships inside a quality system, a risk file, a usability study and a security posture, and those are the parts of this that were settled long before anybody asked about AI. The useful question is not which of them applies, because they all do, but which of them binds.
The quality system changed in February. Part 820 is now the Quality Management System Regulation, incorporating ISO 13485:2016 by reference. Two of the old fifteen subparts survive, with supplemental provisions for records and for labelling and packaging controls. For a company already certified to the standard this removes duplicated work rather than granting an exemption, and medical device reporting, corrections and removals and device tracking were never part of it and did not move. The clause that costs an AI developer most is design control, because that is where a model’s lineage has to live: an artefact with no traceable link to the dataset that produced it is a design output with no design history.
Risk management needs a vocabulary conventional hazard analysis lacks. ISO 14971 works from a fault to a hazard, and a model does not fault; it is wrong at a rate, in a distribution, and the distribution shifts. FDA recognises AAMI CR34971, which applies the standard to AI and machine learning and was the first AI-focused document on its consensus standards list. Its hardest point is the one teams resist: a warning in the instructions for use is a weak control. The standard already ranks inherent safety by design above protective measures above information for safety, and a model’s misclassification is where that ranking bites, because the information for safety is exactly what a user cannot act on in the moment.
And usability is where the model meets the person. IEC 62366-1 asks for hazard-related use scenarios, formative evaluation during development, and a summative evaluation with representative users in a realistic simulated environment before submission. The distinction that decides responsibility is between a use error, which the manufacturer is expected to engineer out, and abnormal use, which is a deliberate violation. Misreading a probability because the interface presented it badly is a use error. So, in practice, is a false positive rate that is tolerable in a paper and intolerable at a hospital’s volume, because the quantity that matters to a nurse is alerts per shift and not specificity.
Security, and the record of what the model did
Section 3305 of the 2022 omnibus added section 524B to the Act, and its obligations took effect on 29 March 2023. A sponsor must submit a plan to monitor and address postmarket vulnerabilities including coordinated disclosure, maintain processes giving reasonable assurance the device is cybersecure and make patches available, and provide a software bill of materials covering commercial, open-source and off-the-shelf components. FDA’s final guidance followed on 27 June 2025, and an inadequate bill of materials is grounds for a refuse-to-accept decision. That makes this the one requirement in this piece whose failure mode is the submission never reaching review.
Three supply chain realities are specific to a model and generic security work tends to miss them. Open-source machine learning frameworks and their transitive dependencies belong in the bill of materials, and the dependency tree of a training stack is large. Model weights obtained from a third party are a component whose provenance is part of the analysis. And adversarial manipulation of inputs is a foreseeable attack rather than a research topic, which puts it in the security risk analysis and, through it, in the safety risk analysis.
Then the record. 21 CFR Part 11 governs electronic records and signatures where a predicate rule requires the record, which for a manufacturer means quality system records, complaint files and reporting records kept electronically. The 2003 scope guidance narrowed enforcement considerably; what remains is system validation, secure computer-generated time-stamped audit trails, and attributability of a record to a person. Attributability is where autonomous software raises a real design question: if a model’s output and a clinician’s confirmation are written into one entry under the clinician’s identity, the record no longer distinguishes what the software proposed from what the person decided.
The reasonable architecture logs the inference as its own event, with model version and timestamp, and the human action as a separate signed event referring to it. That is a conclusion drawn from the attributability requirement rather than a quotation. FDA has issued no rule saying an AI inference and a human signature must be separate log entries, and secondary writing in this area states it as though the agency had. It is good practice, and it should be presented to an auditor as the manufacturer’s reasoning.
Privacy sits beside all of it and shares nothing with it. FDA governs whether the device is safe and effective; HIPAA governs the protected health information used to train and run it. A model hosted outside the covered entity makes its operator a business associate. Using a hospital’s records to improve a product is not treatment, payment or operations for the vendor, so it needs an authorisation, a limited data set, de-identification or an institutional review board waiver — and a business associate agreement permitting the vendor’s own product development is the clause most often assumed and least often present. A model trained on identifiable records can also leak them, which is what membership inference and model inversion name, so de-identifying a training corpus does not de-identify the artefact trained on it. The NIST AI Risk Management Framework, built on govern, map, measure and manage, is the structure most organisations use to keep track of this. It is voluntary. The rewrite of the HIPAA Security Rule that would specify the risk analysis in detail was proposed in January 2025 and is not final.
The rules that bind after you ship
A premarket submission is finite work with a decision at the end of it. The obligations that attach to a model once it is running have no end state, and most of them are not the FDA’s. They attach to the organisation using the software rather than the one that wrote it, which is why they are usually discovered by a customer.
| Obligation | Who it binds | Since | What it asks for |
|---|---|---|---|
| 45 CFR 92.210, Section 1557 | Covered entities — the hospital, not the vendor | 1 May 2025 | Reasonable efforts to identify patient care decision support tools using inputs that measure race, colour, national origin, sex, age or disability, and to mitigate the discrimination risk. The definition covers any tool, automated or not, device or not |
| ASTP/ONC predictive DSI, 45 CFR 170.315(b)(11) | Certified health IT developers | 1 January 2025 | Thirty-one source attributes exposed to users, covering development data, validation, performance, maintenance and funding. Thirteen for an evidence-based intervention |
| State disclosure and review statutes | Clinicians, facilities and insurers | Varies | Disclosure that AI was used in diagnosis or treatment, in Texas. Disclosure of AI-generated patient communications and a physician decision on medical necessity, in California |
The first row is the one most often missed. 45 CFR 92.210 has been in force since May 2025, its definition of a patient care decision support tool is deliberately broad enough to include a score sheet, and a vendor’s clearance letter does nothing to discharge it. Satisfying it requires an inventory of every model in clinical use and subgroup performance information that only the vendor has. Neither party usually prices that gap before signing.
The second row has a jurisdictional limit worth knowing, and it mirrors the FDA’s. The predictive decision support criterion binds certified health IT. A model inside a certified EHR is covered. The same model in a PACS, an imaging orchestrator or a standalone web application is not, however identical its clinical effect.
What the postmarket system can and cannot see
FDA replaced MAUDE and six other legacy databases with the Adverse Event Monitoring System on 11 March 2026. The consolidation is real and anything written about searching MAUDE is now a historical note. It is worth being precise about what AEMS is, because secondary writing has got ahead of it: it is a reporting and surveillance database. It does not ingest model performance streams from health systems and it does not compute population-level performance variance.
That matters because the structural problem is unchanged. Device reporting is built on events, and AI failure is usually not an event. A model wrong 10% of the time is working as specified on any individual case, so no individual case is a malfunction, and gradual degradation across a population produces no report at all. The reporting system cannot see the characteristic failure mode of the technology, and changing the database does not change that.
The external validation of the Epic Sepsis Model published in JAMA Internal Medicine in 2021 is still the clearest illustration, and it is worth quoting accurately because it is usually quoted loosely.
The shallow reading is that the model was bad. The better one is that it was validated on one population and deployed on many. The reading worth carrying into a design review is in the bottom row: an alert on 18% of all admissions at a positive predictive value of 12% is a workflow problem before it is a statistical one, and nothing in the premarket framework would have caught it, because the model was not a device and no submission existed.
The generative case was decided before the question was asked
On 18 August 2026 FDA published a discussion paper on generative AI-enabled medical devices, proposing a two-axis risk framework and a competency-based premarket evaluation modelled on how clinicians are assessed. One axis scores how independently a function acts, in four named steps from non-directive information through action-directing information and supervised action to fully autonomous action. The other scores the consequence of relying on an incorrect output. Evidence expectations scale with position on the grid rather than with the presence of a language model. Comments closed on 19 October 2026.
While that paper was in preparation, FDA cleared one. On 23 December 2025, 510(k) K253281 cleared device software for adults with type 2 diabetes that recommends the next insulin dose through a conversational layer built on a large language model. It is described as the first software as a medical device with a patient-facing LLM. It is Class II, under the product code for a drug dose calculator, cleared by substantial equivalence to a 2018 insulin titration system, and it carries a change control plan whose permitted modifications are bounded so as to preserve the deterministic dosing logic.
Three things follow, and each contradicts something commonly said about generative AI in medicine. A patient-facing generative device does not require a De Novo or a PMA, because one has been cleared without either. The pathway is not the interesting variable; the confinement of the generative component is, and here the language model routes structured information into a clinician-configured algorithm rather than deciding the dose. And a change control plan is the natural instrument for a model like this, because the plan is where a sponsor writes down what the generative part is not allowed to become.
Who pays, and why that decides the design
Reimbursement is not a safety authority and it is the strongest single force on how clinical AI gets built, because it decides which products have a business model at all. The vocabulary is CPT Appendix S, effective from 1 January 2022 and revised since, which classifies AI services by their output rather than by their technique.
| Category | What the machine does | Who interprets and reports |
|---|---|---|
| Assistive | Detects clinically relevant data | A physician or other qualified professional |
| Augmentative | Analyses or quantifies the data in a clinically meaningful way | A physician or other qualified professional |
| Autonomous | Analyses, interprets and produces clinically relevant conclusions | Nobody. That is the definition |
Retinal screening is the one place where the financial consequence of autonomy is visible, because there is a code at each level of human involvement. CPT 92227 covers imaging with remote clinical staff review; 92228 covers remote physician interpretation; and 92229 covers point-of-care automated analysis with a report and no human over-read. The autonomous code has a national rate, about $41 in 2024 and about $44 proposed for 2025, with receipts varying by locality. Use went from 143 services in 2021 to 1,427 in 2023.
One number this piece will not quote. Secondary sources commonly put autonomous retinal screening at $50 to $100 an examination and describe the clinic as retaining the full fee. The published national rate is in the low forties, and the economics turn on not having to pay a reader rather than on a larger fee. A business case built on a doubled rate will not survive its first finance review, and the honest version of the argument does not need it.
What is still unsettled is the hospital outpatient side. In the CY 2027 outpatient rule published on 7 July 2026, CMS proposed a payment category called Software as a Medical Service with a new status indicator making designated codes separately payable. Comments closed on 31 August 2026 and the final rule has not issued. Inpatient, the New Technology Add-on Payment remains a two to three year bridge rather than a rate. A Category III code has no national price at all, because it is priced by the local contractor, which is why none appears here.
What to do about it
Six things, in the order they pay off. The first three are free and the other three are not.
- Write the intended use in one sentence and list every input. If any input is an image, a signal from an in vitro diagnostic, or a repeated measurement of the same parameter, the exclusion is already gone and the rest of the non-device analysis is wasted effort.
- Name the labelled user before anything else. If it includes a patient or an unlicensed carer, there is no clinical decision support exclusion to argue about and the only remaining question is device or general wellness.
- Try to write the criterion 4 package and see whether you can. Intended use and user, required inputs, a plain language account of development and validation, the data behind it, clinical validation results. A team that cannot produce this cannot claim the exclusion, and will struggle with a submission either way.
- Build the subgroup performance artefact once. It answers the draft guidance, the predictive decision support source attributes, and your customer’s obligation under 45 CFR 92.210. Producing it three times in three formats is the most common avoidable cost in this area.
- Make the model traceable to its training data before design transfer. Design control needs it, a change control plan needs it, and retrofitting the link after the fact is the most expensive remediation on this list.
- If the model will change, write the envelope and check you can see the trigger. A performance trigger is only observable if the fall exceeds the uncertainty of the measurement at the volume you actually monitor. A threshold set below that detection floor commits you to noticing something the monitoring cannot show.
None of this is a reason to be pessimistic about shipping clinical AI in the United States. The framework is navigable and a record number of products went through it last year. It is a reason to stop treating the FDA question as the whole of the regulatory question, because for most deployed models it is the smaller half.
The whole of it, on one screen
Everything above, as one thing to keep. Nothing in it is new: every figure appears earlier in the piece with the statute, the regulation, the docket or the study it came from.
If you want the whole of it, Clinical AI Regulation is the paper behind this piece. Thirty-three pages on the five authorities and what each one actually regulates, the four criteria worked through with FDA’s own examples, the pathways and what each proves, documentation level, change control plans, the quality and risk standards, the privacy layer, postmarket reality, generative AI, who pays, and twenty questions for a submission or a tender. It is free at /resources/us-medical-ai-regulation.

