What is per-call authorization for AI agents, and why does it beat per-session?
Per-call authorization means the agent gets a fresh token for each tool call. Least-privilege, short-lived, scoped to one resource and one action. Not a single long-lived key for the whole session. A per-session setup works the other way. It hands out a broad key at session start. It reuses that key for each call that follows.
The gap between the two shows up the moment something goes wrong. An agent is a program that reads unsafe text. Then it decides what to do with its rights. Plant commands in that text: a poisoned GitHub issue, a lead form, a support ticket. A per-session agent already holds broad standing power. It will point that power wherever the injection tells it to. A per-call agent holds nothing broad to begin with. To cause real damage it would need a token scoped for that damage. No such token was ever minted.
The practical result is a much smaller blast radius. Instead of all the user can touch, the exposure is one call. One resource. For the next few seconds. You also gain a clean place to enforce policy, revoke access mid-session, and log activity. This is least-privilege applied at the level of the action. Not the login. It is the non-human identity idea made concrete. The agent is its own principal, with its own scoped keys. Not a process quietly borrowing a human's broad session.
How per-session and per-call actually differ, mechanically
Per-session is the default most agent frameworks ship with. It is the easy path. Authenticate once, stash a token, attach it to each outbound request. It is the same instinct that gave us long-lived API keys. It fails for the same reason. The key's power has no relationship to what any given call needs.
Per-call changes the flow. Before each tool call, the agent's runtime trades its identity for a narrow token. One that describes just this call: this resource, this action, this audience, expiring soon. A read of invoices/2024 gets a token good for reading invoices/2024 and nothing else. A new call gets a new token. Freshly minted and freshly allowed.
What this buys you is non-transferability. With a per-session key, call A and call B carry the same power. Byte for byte, no difference. So a hacker who redirects the agent can reuse A's key for B. It costs them nothing. With per-call tokens, A's token is useless for B. The key also becomes far less worth stealing. It expires quickly and opens just one door. Steal it. You get a few seconds of read access to one invoice.
The obvious objection is cost. It deserves a straight answer. Minting a token before each call means an extra hop to the authorization server. Some issuance latency. That's load on the server you did not carry before. Not when one cached session key covered it all. Real, but tractable. RFC 8693 token exchange tokens carry their own life span. So you cache a narrow token for its short TTL. A burst of reads against invoices/2024 can reuse it. This skips the mint cost on each call. You tune the window. Make it long enough to amortize the exchange. Make it short enough that a leak dies fast. Per-call is about scope. Not a religious one-round-trip-per-call rule.
One caveat to keep honest. A per-call token is still a bearer token. A bearer token is worth just its possession. Grab it inside that short window, and you can replay it. Holding it is the whole proof. A short life shrinks the window. It does not close it. To close it, bind the token to its sender. Then a stolen copy is inert in anyone else's hands. That is what DPoP (RFC 9449) does with a client-held key. And what mTLS-bound tokens (RFC 8705) do with the client certificate. Pair short-lived per-call scope with sender constraint. A token lifted mid-flight is dead on arrival.
What the MCP authorization spec requires (revision 2025-06-18)
The Model Context Protocol authorization spec, revision 2025-06-18, never uses the phrase per-call authorization. It bakes the primitives in as hard rules instead. The ones that matter here:
- Authorization is per request, not per session. The
Authorizationheader must ride on each HTTP request from client to server. Even when those requests are part of the same logical session. There is no log in once and coast mode. - Audience validation is mandatory. An MCP server acts as an OAuth 2.1 resource server. It must verify the access token was issued specifically for it. It must reject tokens that do not name it in the audience. A token minted for service X is dead on arrival at service Y.
- No token passthrough. Say an MCP server calls an upstream API. It must get a separate token from the upstream authorization server. It must not forward the token it received from the client. That single rule closes the confused-deputy problem. Where a downstream API wrongly trusts a forwarded key.
- Clients must send resource indicators. MCP clients implement RFC 8707. They pass a
resourceparameter, so tokens carry an audience, and no one can replay them across services. One precise caveat: the spec makes the client-to-authorization-server binding conditional. It applies when the authorization server supports the feature. The server-side audience validation above is unconditional. So a server always rejects a token that does not name it. Even where a given authorization server cannot yet bind the audience at issue time. - The spec also calls for short-lived tokens. And for rotating refresh tokens issued to public clients. This limits the fallout from any single leak.
Combine those, and you have described per-call, resource-scoped, audience-bound authorization. Without ever naming the pattern. Building on MCP? The companion piece on MCP server security picks up where this leaves off.
The standards under the hood: RFC 8693 and RFC 8707
Two IETF documents supply the machinery. Neither is new or exotic.
RFC 8693, OAuth 2.0 Token Exchange, a Proposed Standard from January 2020, defines a Security Token Service. You hand in one token and get back another. Its important contribution: it formally separates delegation from impersonation. Under delegation, principal A keeps its own identity while acting on behalf of B. The act (actor) claim, plus an optional may_act claim, expresses this. Under impersonation, A simply becomes B, and it loses the trail. Delegation is the one you want. It preserves the whole Human to Agent to Service chain inside the token. Token exchange is the primitive behind this. It lets an agent get a fresh, least-privilege, audience-scoped token for each downstream call. Without breaking that chain.
RFC 8707, Resource Indicators for OAuth 2.0, published the same year, defines the resource request parameter. It binds an issued token to a specific target audience. It is the mechanism MCP relies on. Each per-call token works at its intended server, and nowhere else.
A caveat is worth stating plainly. A vendor may eventually pitch this to you. Per-call authorization names a design pattern. It is not a normative term in any single spec. No RFC carries that title. No RFC mandates the pattern by name. What the standards give you are the parts: token exchange, resource indicators, per-request audience validation. You build those into the pattern yourself. Also note: RFC 8693 and RFC 8707 are Proposed Standards. They offer building blocks, not a finished, interoperable per-call profile. If someone tells you a standard requires per-call by name, they are overselling it. The standards give you the pieces. Not the built pattern.
Why prompt injection makes external authorization non-negotiable
The root cause is uncomfortable. LLMs mix trusted commands and unsafe data in the same context window. They cannot reliably tell which is which. Simon Willison's long-running analogy is SQL injection. The model concatenates your commands with the hacker's text and runs the blend. This is not a bug you patch once. It is a structural property of how these models read input.
That's the reason authorization has to live outside the model. You cannot prompt-harden or fine-tune your way to a model that reliably refuses injected commands. Just as you cannot comment your way to SQL that ignores malicious input. The enforcement point has to be a part the model cannot talk around. You must check it at each call.
The standards bodies see the same shape. In the OWASP Top 10 for LLM Applications 2025, published 18 November 2024, Prompt Injection is LLM01. The number one risk for the second edition running. Right next to it sits LLM06 Excessive Agency. The category for agents handed too much autonomy and too many permissions. OWASP's own mitigation for Excessive Agency is least-privilege plus human authorization on sensitive actions. Per-call authorization is a way to operationalize both at once.
Worth keeping in perspective: this is an engineering threat model for a single control. Not a complete security program. It is one strong beam, not the whole house.
Real incidents: what over-permissioned agents actually do
Two well-documented cases show the failure mode clearly. Both are proofs-of-concept, shared responsibly, not breaches seen in the wild. Worth being precise about when you discuss risk internally.
The first is the GitHub MCP private-repo exfiltration. Shared by Invariant Labs on 26 May 2025. A hostile public GitHub issue carried an indirect prompt injection. A user's agent, running the official GitHub MCP server, read the issue. It followed the planted commands. It pulled data from the user's private repositories. Then it leaked that data through an auto-created public pull request. Invariant reads it this way: a server code bug did not cause it. The cause was architectural. The agent held broad cross-repo session power. So one poisoned input could steer it anywhere that power reached. Their recommended fix: restrict the agent to one repository per session. Also issue least-privilege tokens. That is per-call scoping under another name.
The second is ForcedLeak in Salesforce Agentforce (Einstein AI). Shared by Noma Security at CVSS 9.4. Noma says the team found the issue on 28 July 2025. Salesforce remediated it on 8 September 2025. The team shared it in public on 25 September 2025. Noma's account describes hackers stuffing commands into a Web-to-Lead description field. That gave them a large block of text to work with. An employee later queried lead data. Agentforce then ran the injected commands. It exfiltrated data to an expired domain still on the CSP allowlist. A domain the researchers say they acquired cheaply. As in the GitHub case, the agent acted on broad standing rights. Not a power scoped to the task at hand. These figures come from a single vendor's disclosure. Treat them as Noma's reported account, not an on its own confirmed detail.
Both cases rhyme with Willison's lethal trifecta, described on 16 June 2025. Access to private data. Risk from unsafe content. And a way to communicate outside. Line up all three, and one poisoned input can exfiltrate data. With no traditional vulnerability in the picture. Per-call authorization attacks the trifecta. It makes sure the agent never holds enough standing power to complete the chain.
Per-call authorization best practices: how to prevent agent overreach
Here is the checklist to put this into practice. It's what the specs and the incident responders point to.
- Mint a fresh token per tool call via RFC 8693 token exchange. Scope it to that call's resource and action, with the shortest life you can tolerate. One call, one token, one purpose.
- Bind each token to its audience with RFC 8707 resource indicators. Have each server validate the audience claim, and refuse anything not addressed to it. A stolen token should open just one door.
- Sender-constrain the token, so it is not pure bearer. A short life limits a leak. But DPoP (RFC 9449) or mTLS (RFC 8705) changes that. A token grabbed inside its window is useless without the client's key. Short-lived and sender-bound combined is what neutralizes token theft.
- Never pass tokens through. When your agent's server calls an upstream API, exchange for a new upstream-scoped token. Do not forward the client's. That is the confused-deputy fix. It is a hard rule in MCP 2025-06-18.
- Use PKCE and exact redirect-URI matching. This blocks code interception during the auth flow.
- Scope the runtime, not just the token. Invariant's one-repository-per-session guidance generalizes. Give each session the minimum surface it needs. No adjacent resources.
- Grant access just in time. As agents move from reading data to acting on it, scope access per action. Not through a persistent session. Red Hat's analysis of MCP security makes the same case. Enforce least privilege at the point of each tool call.
- Rotate refresh tokens. Keep access tokens short-lived, so any single leak has a short, narrow window.
- Keep unsafe input away from actions that matter. Say a call can move money, delete data, or publish externally. Gate it behind explicit human authorization. Do not let injected text reach it unmediated. Do not build the lethal trifecta in the first place.
- Assume the model is compromised. Place the enforcement point somewhere it cannot argue with. The broader OWASP MCP Top 10 catalogs the MCP-specific risks these checks are meant to contain.
Govern the data plane, not only the actions
Per-call authorization governs what an agent is allowed to do. There is a second question that often gets skipped. What an agent can reach at all. Both matter. The strongest posture answers both.
That is the angle we build DataShield around. Tokenize sensitive fields at ingest. A hijacked agent that slips its leash finds tokens where the PII used to be. Not raw data it can exfiltrate. Authorize each tool call on its own. Keep the ability to revoke power mid-session, the instant something looks wrong. Do not wait for a long-lived session to time out. And seal each call into a tamper-evident audit chain you can verify afterward. Which call touched what, and under whose power: that question gets a cryptographic answer. Not a shrug.
The underlying point is this. Do not rely only on constraining what the agent will do based on its reasoning. Prompt injection can corrupt that reasoning directly. Constrain what it can reach, and what each call is allowed to touch. Outside the model. At the level of each single call. A per-session key cannot give you that level of control. Per-call can. And the building blocks for it are not new. Token exchange and resource indicators have sat in the RFC series for years.
Watch: related explainers
Three short watches on agent identity, MCP, and the injection that makes authorization matter.
Frequently asked questions
What is per-call authorization for AI agents?
It is a design pattern. The agent gets a fresh, least-privilege, short-lived token for each tool call. Scoped to that specific resource and action, instead of reusing one broad key all session. It is built from OAuth 2.0 Token Exchange (RFC 8693, 2020) and Resource Indicators (RFC 8707, 2020). MCP's 2025-06-18 authorization spec mandates the underlying pieces: per-request authorization, audience validation, and no token passthrough.
Why does per-call authorization beat per-session authorization?
Because it makes power non-transferable across calls. A per-session agent holds one broad key. So a hacker who hijacks it via prompt injection can reuse that key. For any call the agent can make. A per-call agent only ever holds a token scoped to the one call in progress. It expires in seconds. So a redirected agent cannot recycle it into another, more damaging action. The blast radius drops from all the user can touch to one resource for a few seconds.
Does the MCP specification require per-call authorization?
Not under that name. The MCP authorization spec (revision 2025-06-18) requires the primitives that make it up. The Authorization header must be sent on each request, even within one logical session. Servers must validate token audience per RFC 8707. Servers must not pass the client's token through to upstream APIs. And clients must use resource indicators. No spec is titled per-call authorization. The standards supply the parts you build into the pattern.
Do short-lived per-call tokens stop a stolen token?
Only partly. A per-call token is still a bearer token. So anyone who grabs it inside its short window can replay it. Possession is the whole proof. A short life shrinks that window but does not close it. To close it, sender-constrain the token with DPoP (RFC 9449) or mTLS-bound tokens (RFC 8705). Then a stolen copy is useless without the client's key. Short-lived plus sender-bound is the combination that neutralizes token theft.
Can prompt injection be fixed by better prompts or model training instead?
No, and that is the crux. LLMs mix trusted commands and unsafe data in one context. They cannot reliably tell them apart. That makes prompt injection structural, not a patchable bug. In OWASP's LLM01:2025 ranking it is the number one LLM risk. Authorization has to be enforced outside the model. At each call, by a part the model cannot talk its way past. Prompt hardening lowers frequency but never makes refusal reliable.
Have per-session agents actually been exploited in the wild?
The best-documented cases are two. The GitHub MCP private-repo exfiltration (Invariant Labs, 26 May 2025). And ForcedLeak in Salesforce Agentforce (Noma Security, CVSS 9.4, shared 25 September 2025). Both are proofs-of-concept, shared responsibly, not confirmed in-the-wild breaches. Both worked for the same reason. The agent held broad, standing session rights. Indirect prompt injection could redirect those rights. The lesson holds either way, weaponized or not. Over-permissioned agents fail this way.