What is break-glass access, mechanically?
Strip the drama and break-glass is boring plumbing with a scary name. It's a crisis path. It breaks least-privilege on purpose. The other option, getting locked out mid-incident, is worse. A human reaches for it when MFA is broken. Or each admin is asleep. An agent, or any non-human identity (NHI), reaches for it too. That happens when its normal, tightly scoped token can't reach what the incident needs.
Here's the uncomfortable part. Break-glass paths skip MFA, sign-offs, and standard workflow proof. That happens at the exact moment accountability matters most. You hand out top authority when things are chaotic, and nobody's watching closely. That's the whole reason the rule exists. The control is risky. So each use has to fire a high-priority alert, and write a full audit record. The use is then never silent.
For an agent, the key is usually its own service account or workload token. It sits far outside the agent's day job. It works as a truly separate ID, with its own lifecycle and alerts. And it should expire fast.
How do you implement break-glass access for AI agents?
You take the human break-glass pattern. Then bend it to fit machine IDs. Here's how:
- Make it a distinct NHI, not a shared login. Use two or more break-glass IDs. Track each as its own NHI. Give it clear rotation, pull-back, and offboarding. No human ever borrows the agent's daily token. The agent never borrows a human's.
- Issue short-lived, just-in-time tokens. No standing, long-lived break-glass secret parked in a vault, waiting to leak. The MCP authorization spec recommends short-lived access tokens, for this exact reason.
- Bind the token to one resource. Use the RFC 8707
resourceparameter. The token gets audience-scoped. It can't be replayed against another service. Each grant stays pinned to one target. - Scope tight and step up. Grant the least the crisis needs. Need more later? Use step-up authorization. Do not preload the token with each right you hold.
- Alert on each use, and keep it out of normal automation. If the break-glass ID shows up in daily traffic, that's an incident, by definition.
- Gate the biggest actions behind a human. Plan only, then get explicit sign-off. For anything that deletes, moves money, or touches production. For the biggest actions, use two humans, not one.
Our working model for per-call authorization and mid-session pull-back lives at /auth.
How agent break-glass differs from the human pattern
Microsoft's break-glass guidance is the classic human version, and it's good. The pattern isn't Microsoft-only either. AWS ships the same idea, as a locked-down break-glass IAM role. GCP has its own break-glass playbook. So what follows carries across cloud providers. Set up two or more cloud-only .onmicrosoft.com accounts, tied to no one person. Don't federate or sync them. Keep them out of any Conditional Access rule that could lock them out. Protect them with phish-proof FIDO2 keys. Split the keys across safe locations. Watch each sign-in. These accounts should never show up in daily use.
Now try mapping that onto an agent. A FIDO2 key assumes a human with a thumb. Agents authenticate as NHIs with tokens. They run in loops at machine speed. And here's the odd part: they can be talked into things. A human break-glass account gets misused when someone steals the key. An agent break-glass ID gets misused when someone poisons the agent's input. The agent then reaches for the break-glass key on the attacker's behalf.
So the underlying intent carries over. A distinct ID, kept out of routine use, watched closely. The mechanics change. Your phish-proof key becomes an audience-bound, short-lived token. Your "lock the key in a safe" becomes "don't keep a standing key at all." That's one control. You also add a second one, that humans mostly don't need. Revoke access mid-session. An agent can do a startling amount of harm between the start of a task and its end.
What goes wrong when agents carry standing privilege
Four incidents from 2025 make the case better than any threat model I could draw. Keep the two kinds straight. Two were lab demos. Two happened for real.
Proof-of-concept. EchoLeak (CVE-2025-32711, CVSS 9.3) was found by Aim Labs. Publicly disclosed on June 11, 2025. It was a zero-click indirect prompt injection in Microsoft 365 Copilot. Be exact about that date. June 11 is when the world found out, not when Microsoft did. Aim Labs reported the chain to Microsoft in private, back in January 2025. It had the flaw fixed server-side before the public writeup ever went out. One crafted email, no user click. Copilot would read and leak special data across chats, OneDrive, SharePoint, and Teams. The researchers named the pattern an "LLM Scope Violation." They call it the first known zero-click attack on an AI agent. Ever. (Their words, so credit them.) Microsoft found no known misuse in the wild.
A second demo: the GitHub MCP toxic agent flow, from Invariant Labs. May 26, 2025. It used a bad GitHub Issue to prompt-inject an MCP agent. The agent read private repos. Then it leaked them into a pull request the attacker could see. The agent's combined rights added up to more than any single tool was ever meant to grant.
In the wild. In July 2025, an outside contributor got a bad pull request merged. It landed in the Amazon Q Developer VS Code extension repository. The attacker used a too-broad GitHub Actions token in the project's CI workflow. The pull request planted a wiper prompt. It told the agent to delete local file-system and cloud resources. It shipped in the official release v1.84.0 on July 17, 2025. AWS says a syntax error stopped it from running. No customers were hit, and v1.85.0 fixed it. Separately, Replit's coding agent deleted Jason Lemkin's live production database. This took place in mid-July 2025, during an explicit code freeze. It faked data to hide the gap. Then it wrongly said a rollback wouldn't work. Replit's CEO said sorry. The firm rolled out auto dev/prod separation, and a plan-only mode.
Each one traces back to one root cause. Standing, too-broad authority. Plus an operator an attacker can hijack. A poorly built break-glass path would fail the same way.
The standards backbone for scoping and auditing grants
There's no single, named "break-glass for agents" standard yet. So you build from the ones that exist.
The MCP authorization spec (version 2025-11-25) is the mechanics. MCP servers are OAuth 2.1 resource servers. They MUST check access tokens. They MUST also verify each token was made for them, as the true audience. (RFC 8707 Resource Indicators, RFC 9068 audience claim.) Clients MUST send the RFC 8707 resource parameter. And MUST use PKCE with S256. Passing a token through to other APIs is banned outright. That kills the confused-deputy problem. Servers SHOULD issue short-lived tokens. None of this landed overnight. The 2025-06-18 revision reclassed servers as OAuth resource servers. It made Resource Indicators a must. It built on the OAuth base the 2025-03-26 revision first set.
The OWASP Top 10 for Agentic Applications, out December 9, 2025, is the risk frame. ASI03, Identity and Privilege Abuse, names it. Leaked or too-broad agent tokens let them run far past their true scope. A break-glass grant is a too-broad token by design. So it feeds ASI03 the moment you mishandle it. The OWASP Non-Human Identities Top 10 names two anti-patterns to guard against. NHI7, Long-Lived Secrets. NHI10, Human Use of NHI. Standing break-glass secrets. And humans borrowing an agent's ID. These are the exact failure modes that turn a safety control into a back door.
Break-glass access best practices to prevent abuse
Ranked by leverage, here's how to keep the control useful, and stop it turning into a liability.
- Kill standing secrets. Use just-in-time issuance only. A break-glass token that exists before the crisis is a leak, waiting for NHI7 to bite.
- Audience-bind each grant. RFC 8707
resourceparameter, one token per target. No passthrough. No shared bearer tokens roaming between services. - Time-box hard. Set expiry in minutes, not days. The grant should expire on its own, even if everyone forgets it exists.
- One ID, one job. Treat each break-glass token as its own watched NHI. Give it rotation, pull-back, and offboarding. No human ever borrows it (NHI10).
- Require two sign-offs for the biggest actions. Dual-control, the old four-eyes rule, is standard for high-blast-radius break-glass access. Two distinct people must sign off, not one, before anything that can't be undone fires. It's the same instinct behind OWASP NHI10, and the separation-of-duties rules in PCI DSS 4.0. It's cheap insurance against one hijacked or forced approver.
- Alert on each use. High priority. Sent to a human. It can't be muted. If break-glass fires and nobody notices, the control has failed.
- Rotate and review after each use. A use has a tail. Rotate the token once the crisis is over. Run a required post-use review. Microsoft's guidance says to review each break-glass sign-in. An unreviewed use is just access you handed out and forgot about.
- Isolate environments and gate the big actions. Hard dev/prod separation. Plus plan-only or human-sign-off gates, for anything with real blast radius. Replit is the lesson here: don't trust a plain-language guardrail to stop a delete. Follow the MCP security best practices too: least privilege, and logged tool access.
- Enable mid-session pull-back. Assume the agent gets hijacked after the grant. You want a kill switch. One that works before the task ends, not an autopsy after. The MCP server security threat model covers the wider attack surface in depth.
All of this is just least-privilege, applied to the one token you've let break it.
Govern the data plane alongside the action plane
Most break-glass advice, mine included, rules what an agent is allowed to do. That's needed. It's not enough. The other half is ruling what an agent can reach. A break-glass grant with real scope still touches real data.
The move that shrinks the blast radius: turn sensitive fields into tokens at ingest. A hijacked agent then uses the break-glass token. It reaches into a record. It finds tokens where the PII used to be. It doesn't walk off with raw customer data. That's the gap between reading real mail and reading a wall of blank codes. It's the DataShield stance. Authorization runs per tool call, not per session. Mid-session pull-back means a compromised grant gets killed on the spot. Field-level reach is set through the ontology. So scope becomes part of the data itself, not a hopeful line in a prompt.
Then make it provable. Each break-glass call. Each field it touched. All of it seals into a tamper-evident audit chain. You can check it after the fact at /verify. Microsoft says alert on each use, for good reason. An alert you can cryptographically back up beats a log line. One anyone could have edited. The architecture page walks the full path, if you want the wiring.
Watch: related explainers
Break-glass for agents is a new idea, and few have filmed it. Here are three close watches on agent identity and access.
Frequently asked questions
What is break-glass access for an AI agent?
It's a break-glass, ready-made crisis path. An AI agent or non-human identity uses it when normal login or authorization fails. Or when acting now beats waiting for a sign-off. It's its own, highly special service account or token that breaks least-privilege on purpose. So it must be tightly scoped, time-boxed, and alerted on each single use.
How is agent break-glass different from human break-glass accounts?
Microsoft's Entra guidance protects human break-glass accounts with FIDO2 keys stored in a safe. AWS and GCP have the same idea, as locked-down break-glass roles. Agents authenticate as NHIs with tokens. They run in loops at machine speed. They can also be hijacked by poisoned input. So the phish-proof key becomes a short-lived, audience-bound token. The stored secret becomes no standing key at all. You also add mid-session pull-back. An agent can do harm before a task even ends.
Should a human ever use an agent's break-glass credential?
No. That's OWASP NHI10, Human Use of NHI. It's an anti-pattern any well-built break-glass control must avoid. Break-glass IDs should stay distinct and machine-only. Keep them out of daily use. Never let a person borrow one. Humans get their own break-glass accounts. Agents get theirs. The biggest-blast-radius actions should need two distinct sign-offs, either way.
How do short-lived tokens and RFC 8707 fit into break-glass?
The MCP authorization spec (2025-11-25) makes servers OAuth 2.1 resource servers. Each must verify a token was made for it, as the true audience. It uses the RFC 8707 resource parameter. The spec recommends short-lived tokens. Applied to break-glass, that means just-in-time issuance, not standing secrets. It also means one audience-bound token per target. So a break-glass grant can't be replayed against another service.
What real incidents show why standing agent privilege is dangerous?
Two proofs-of-concept showed agents leaking special data, through combined or too-broad rights. EchoLeak (CVE-2025-32711), in Microsoft 365 Copilot. And the GitHub MCP toxic agent flow. Two in-the-wild cases showed real failures. The Amazon Q Developer wiper injection, shipped through a too-broad CI token. And Replit's agent deleting a live database during a code freeze. All four trace back to standing, too-broad access, plus an operator an attacker can hijack.