The decision that is made before the model
An agent that reads a patient record is usually discussed as a model problem. Which model, what context window, how the retrieval is chunked, whether the summary can be trusted. All of that is downstream of a decision taken earlier and somewhere else, normally by whoever filled in a client registration form, and it is the decision that determines what the agent can reach, what the hospital can prove afterwards, and whether anybody can answer the question an auditor will eventually ask.
The decision is which kind of thing the agent is. OAuth has a vocabulary for this and healthcare has a profile of it, and neither has a word for an agent. An agent is a client, not a user, and the record it is talking to was built to tell those two apart.
Three subjects, and an agent is none of them
SMART App Launch defines the authorization layer almost every FHIR server in the United States exposes, because the certification criterion requires it. Its scope grammar begins with one of exactly three prefixes, and they are not three levels of the same thing. They are three different claims about who is at the other end of the connection.
patient/ is a single patient in context, chosen at launch. user/ is, in the specification’s own words, data available to the signed-in user rather than data about them, which means the ceiling is whatever that person’s own account already reaches. system/ is for the case the spec describes as having no user in the loop: a client authenticates with a signed assertion, gets a token, and the only consent that ever happened was an administrator registering it.
Two sentences in that specification do more work than the grammar does. The first: scopes allow a client to request the delegation of a specific set of access rights, and such rights are always limited by underlying system policies and permissions. The second, with the exclamation mark the authors put on it: the scopes ultimately granted by the authorization server may differ from the scopes requested by the client. A scope is a ceiling, not a grant. Asking for more does not produce more, and a client that assumes it got what it asked for is a client that will fail in production on a patient whose record is restricted.
An autonomous agent fits none of the three cleanly. It is software, so it is not a user. It acts because a person asked, so it is not unattended. It persists across sessions and acts when nobody is watching, so it is not an app in a clinician’s hands. Every deployment resolves this by borrowing one of the three, and the choice is made early, quietly, and almost always for operational reasons rather than for the reason that will matter later.
What a scope can actually narrow
The second half of a scope is the part people think they control. It is narrower than it looks.
The letters after the dot are a subset of cruds, in that order, and each maps to specific FHIR interactions rather than to a general permission. c is type-level create. r is instance-level read, vread and history. u is instance-level update and patch. d is instance-level delete. s is type-level and system-level search and history.
The consequence that catches people is that r on its own is close to useless. An instance read needs an id, and a client that has not searched does not have one. Search is what finds things, so the real minimum for an agent that explores a chart is .rs, and a scope review that waves through read-only access has approved the ability to enumerate.
SMART 2.2 added a query-string filter, so patient/Observation.rs?category=laboratory is a legal scope that reaches lab results and not the social-history answers filed in the same resource type. This is the only general mechanism for sub-resource granularity, and its coverage is thin on purpose: US Core requires a certifying server to support three category filters on Condition and five on Observation, and names them as the ones the HTI-1 final rule requires. Nothing else has a required filter. patient/DocumentReference.rs is every note, including the psychiatry note and the social work note, and there is no standard scope that is less than that.
Build one and read what it reaches. The warnings underneath are the specification’s, not ours.
The bridge, and the token that cannot cross it
Most agents do not speak FHIR. They call tools, and something behind the tool speaks FHIR. In 2026 that something is increasingly an MCP server, and the Model Context Protocol has an authorization specification with a position on exactly this arrangement.
The specification makes an MCP server an OAuth 2.1 resource server. Clients must implement RFC 8707 resource indicators and send a resource parameter naming the server the token is for. Servers must validate that a token was issued for them as the intended audience, must only accept tokens valid for their own resources, and must not accept or transit any other token. Then the sentence that decides the architecture: if the MCP server calls an upstream API, the token used there is a separate token issued by the upstream authorization server, and the MCP server must not pass through the token it received.
That rule is right and the reason for it is good: a token replayed at a service it was not issued for is the confused deputy problem, and audience binding is the mitigation. But follow it through in a hospital. The clinician authorises the agent. The agent gets a token for the MCP server. The MCP server, forbidden from forwarding it, authenticates to the FHIR server as itself, which in practice means SMART Backend Services: a signed JWT assertion, the client credentials grant, system/ scopes, and a token the profile says should expire within five minutes.
The record now sees one client. Not the agent, and not the clinician behind the agent. It sees a service account holding the union of everything an administrator pre-authorised, doing a thing, at a time. Break-the-glass, consent directives, the narrower permissions on a trainee’s account and every sensitive-category rule are all evaluated against an identity that is not the one that asked.
Four topologies, and the same five questions
There are four shapes this takes in practice. They are not a maturity ladder, and the first is the right answer whenever it fits.
Delegation, if anyone implemented it
There is a standard for the thing the third topology is missing, and it is twenty-five years of OAuth thinking old. RFC 8693 defines token exchange: a client presents a subject_token representing the party on whose behalf the request is made, optionally an actor_token representing the acting party, and receives a token that carries both. The act claim expresses that delegation has occurred and identifies the actor.
The RFC also draws the distinction this problem needs. Under impersonation, the agent is given all the rights of the user and is indistinguishable from them in that context. Under delegation, the agent keeps its own identity and the user has delegated some rights to it. Only the second is auditable, because only the second leaves two names in the token. There is even a may_act claim, which states in advance that one party is authorised to become the actor for another.
Almost nobody in healthcare implements it. SMART does not profile token exchange, so there is no conformance statement to point a vendor at, no row in a certification test, and no way to specify it in a procurement document except by writing the specification yourself. The work most likely to close this is elsewhere: SMART Permission Tickets, discussed in the Argonaut and US Core session at FHIR DevDays in 2026, aim at portable pre-verified permission grants a client can present at any participating endpoint. That is a different problem with an overlapping shape, and it is the thing to watch.
In the meantime the gap is a procurement question rather than an engineering one. Nothing will require a vendor to offer delegated tokens, so the only mechanism that produces them is customers asking.
Read is certified. Write is not.
The read path an agent uses is a standard. §170.315(g)(10) requires FHIR R4.0.1, the US Core profiles matched to the USCDI version in force, SMART App Launch, token introspection, refresh tokens valid for at least three months for confidential apps, and Bulk Data with the group export operation. A health system can hold its vendor to all of it.
The write path is not a standard, and the criterion says so in as many words: these services specifically exclude write capabilities. Nothing an agent does that changes the record is inside the certified surface. Every create and every update is a bilateral arrangement, enabled per customer, scoped to the resource types that vendor chose to expose, on terms that are commercial rather than regulatory.
| Read | Write | |
|---|---|---|
| Covered by §170.315(g)(10) | Yes. Single patient and population, to the USCDI in force. | No. Explicitly excluded from the criterion. |
| Scope vocabulary | Required scopes listed in US Core, including the eight category filters. | No required scopes. Whether a write scope exists is a vendor decision. |
| What you can hold a vendor to | Certification, real-world testing, the information blocking rules. | A contract, and whatever the app review programme decided. |
| What this does to an agent | A portable read path that works the same way across estates. | A per-vendor integration, re-done for each one, with its own review. |
This is the single most load-bearing fact for anyone planning an agent that does something rather than says something. The demo where it drafts a note and files it is the demo. The production version of that is a negotiation, per vendor, per customer, and it is on the critical path.
Naming the actor in the record
Suppose the write happens. Who does the record say did it?
FHIR R4 is mostly ready for this. AuditEvent.agent.who, Provenance.agent.who and DocumentReference.author may all reference a Device, so an AI can be named as the author of a note and as a participant in an event. The gap is in one place and it is a conspicuous one: Observation.performer is defined as who was responsible for asserting the observed value as true, which is precisely the question, and its reference list has no Device in it.
HL7 is working the problem. AI Transparency on FHIR, sponsored by the EHR work group and balloted in January 2026, defines an AIProvenance profile requiring at least one agent whose who resolves to an AIDevice, allowing human agents beside it, fixing the reason to the AIAST security label for artificial-intelligence-asserted content, and slicing the entity list so the model card and the input prompt are both carried as references. There is an AIconfidence extension for the probability, usable at resource or element level.
It is worth knowing two things about it. The guide is a continuous build and not an authorised publication, so it is something to design toward rather than conform to. And its own first example runs straight into the performer gap: the source carries a note saying the element cannot hold a Device, and falls back to the alternate-reference extension. The data model is not quite ready for an author that is not a person, and the people fixing it have found that out in the first file.
The audit log is where the action survives
Six months after an agent reads a chart, the only account of what happened is in an audit record. The question will not be a general one. It will be whether this specific patient’s record was accessed on this specific day, by whom, and why, and it will arrive from a privacy officer with a complaint in front of them.
IHE Basic Audit Log Patterns profiles FHIR AuditEvent for exactly this, with patterns for each RESTful interaction and separate patterns for the SAML and OAuth security token cases and for a consent decision. The important structural property is that AuditEvent.agent repeats, and each agent carries a requestor flag saying whether that party initiated the event. One record can name the clinician, the application, and the organisation, and say which of them started it.
Which means the data model has been ready for agent attribution for years, and the thing that defeats it is upstream: if the token that reached the FHIR server carried only a client identity, there is no second agent to write. The audit record is complete, correct, and silent on the only question anyone will ask. The answer then lives in the agent platform’s own logs, which is a different system, with a different retention period, a different access control list and a different auditor.
If you are stuck in that position, BALP has a slot for the join and it is worth using rather than inventing a header. The query pattern defines an entity slice typed XrequestId, for a cross-system request identifier, so one logical transaction can be followed across logs belonging to different systems. Write the same value beside the user in the agent platform’s log and the reconstruction becomes a documented two-step procedure. Then have somebody who was not on the project actually perform it, on a real day and a real patient, because a correlation id that has never answered a question is a design intention rather than a control.
The data is also the instruction channel
One more thing changes when the client is a model rather than a program. A program that fetches a note and a model that reads one are doing different things, because the note can address the model.
The best evidence available is narrower than the headline it usually gets, and the narrowness matters. A JAMA Network Open study published in December 2025 ran three lightweight commercial models across twelve clinical scenarios and found injected instructions followed in 94.4% of evaluations at turn four, 102 of 108, with two of the three models at 100%. The injections were inserted into the user queries, not into retrieved chart content, so this is a measurement of whether these models follow instructions that reach them rather than a measurement of indirect injection through a record. It is the right number to quote and the wrong number to round up.
The structural point stands either way. An agent with read access to free text has an input channel that it cannot distinguish from its own instructions, and a patient-portal message, a scanned referral and an outside note all reach it through that channel. The defence is not better instructions. It is that the scope on the token is the one boundary the model has no ability to argue with: a token that cannot write cannot be talked into writing.
This is also why a tool definition is not a control. MCP lets a server annotate a tool with behavioural hints, and the specification is blunt about their standing: clients must consider tool annotations untrusted unless they come from trusted servers. A tool marked read-only is a label. A .rs scope is an enforcement point, in a different process, that the model never sees.
What to settle before anything connects
Six questions. If the answers are not written down somewhere, the integration has made the choices anyway.
- Which subject, and why not the one above it. If the answer to "why not a user token" is that the service account was easier to register, the attribution gap was bought for convenience and nobody priced it.
- What the scope string is, in full. Written out, with the filters.
user/*.crudsis shorter to type than the right answer and the difference is invisible until someone reads a log. - Whether a human identity crosses the bridge, and how. Token exchange, a signed assertion of your own, or a correlation id logged at both ends. Those are the three answers. "It is in the application logs" is the fourth and it is the one that fails an audit.
- What the audit record will say. Write the AuditEvent you expect, before building, and check it answers the privacy officer’s question rather than the engineer’s.
- Whether anything writes, and under whose name. If it writes, the attribution question is now a clinical safety question and the certified surface does not cover it.
- What the token can still reach when the model is wrong. Assume the instructions came from the chart. The scope is the answer, and it is the only one that holds.
None of this is about the model, and none of it gets easier by choosing a better one. The hard part is in the same place it has always been: between the systems, in the part where a field means two different things and nobody owns the join.

