Est.

MCP Authorization Layer Design for Production Deployments

Protocol overhaul closes security gaps exposed by three production incidents.

Staff Writer · · 11 min read
Cover illustration for “MCP Authorization Layer Design for Production Deployments”
Access Control for Agents · October 5, 2026 · 11 min read · 2,441 words

MCP launched in November 2024 without a settled authentication model, and the developer community building on it didn't wait around for one. By mid-2026, tens of thousands of MCP servers were running in production, with SDKs pulled down tens of millions of times a month, all according to the protocol's own Wikipedia record. Adoption ran so far ahead of the security groundwork that the NSA put it in writing: a May 2026 report found that MCP's "rapid adoption has significantly outpaced the development of its security model." That's a government agency stating, on the record, that the industry shipped the car before it built the brakes.

Three incidents show what that gap actually cost. In April 2025, Invariant Labs demonstrated tool poisoning: an MCP server's tool descriptions get loaded straight into the agent's context window as trusted text, so anyone who controls a description can bury instructions inside it that the model will read and obey. Different surfaces, same failure. In each case the protocol assumed that whatever reached the model had already been cleared for use, and nothing in the system actually checked that assumption at the moment it mattered.

The July 2026 spec revision reads like a direct answer to that record. It doesn't patch around the incidents one at a time. It rebuilds the parts of the protocol that let them happen.

What the July 2026 spec revision changed about authorization

The specification dated July 28, 2026, which Anthropic technical staff member David Soria Parra called the most substantial change to the protocol since authorization was added at all, rebuilt the transport layer and the authorization layer in the same release. That pairing isn't a coincidence of scheduling. The two problems were the same problem looked at from different angles.

On the transport side, MCP dropped its session model. The initialize/initialized handshake is gone (SEP-2575), and so is the Mcp-Session-Id header (SEP-2567). A request now carries everything it needs to be understood and judged on its own. It can land on any server instance sitting behind an ordinary round-robin load balancer, no sticky sessions required. The old stateful design made mid-session authorization failures awkward to handle, since the client was pinned to one instance and the session was already open when something went wrong. Once every request stands alone, an authorization check can run, pass or fail, on every single one of them, which is what the new OAuth work actually needs underneath it. The spec also now requires Mcp-Method and Mcp-Name HTTP headers (SEP-2243), so a gateway, a rate limiter, or a WAF can make routing and authorization decisions by reading headers instead of parsing the JSON body of every request, a change that saves real CPU cycles for anyone enforcing policy at that layer.

On the authorization side, six SEPs do the hardening work. RFC 9207 issuer validation (SEP-2468) requires authorization servers to return an iss parameter and clients to check it before redeeming a code, which closes off a class of mix-up attacks between authorization servers. Client credentials are now locked to whichever issuer minted them, with no reuse across servers (SEP-2352). Dynamic Client Registration, long the default way clients registered themselves with a server, is formally deprecated in favor of Client ID Metadata Documents; DCR keeps working for now, but its removal from a future version of the spec is already on the calendar. Multi Round-Trip Requests (SEP-2322) let a server ask the client a mid-call question, through an InputRequiredResult pattern, without needing to hold a connection open, solving a real problem for confirmation and elicitation flows under the new stateless model. And server-initiated requests, which used to be merely recommended to run only while a client request was in flight, are now required to behave that way (SEP-2260).

Taken together, these changes draw the skeleton: a stateless transport that can carry authorization decisions request by request, and an OAuth layer hardened enough to make those decisions trustworthy. What the spec does not do is tell an implementer how to put flesh on that skeleton. That's the work the rest of this piece covers.

Resource server and authorization server must stay separate

The July 2026 authorization spec draws one line before anything else: an MCP server is an OAuth 2.1 resource server, full stop, and it validates tokens. Issuing those tokens, running login, and collecting consent is someone else's job, handled by a separate authorization server. Building a token issuer into the MCP server itself collapses the whole model, because nothing downstream can trust that the thing checking permissions isn't also the thing handing them out.

Which authorization path applies depends entirely on how the server is reached. Servers running over HTTP are expected to follow the OAuth 2.1 spec. Servers running over STDIO, which by definition live on the same machine as the client, are explicitly told not to bother with that flow and instead pull credentials straight from the environment. That's a meaningful fork in deployment strategy: a local STDIO server and a remote HTTP server are not the same security problem wearing different clothes, and treating them as interchangeable is how teams end up building OAuth machinery nobody needed.

Discovery is where the spec gets genuinely prescriptive. Every MCP server must implement RFC 9728 Protected Resource Metadata. When a client hits a server without a valid token, the server returns a 401 along with a WWW-Authenticate header, and that header is how the client finds out which authorization server it needs to talk to, via either RFC 8414 Authorization Server Metadata or OpenID Connect Discovery 1.0. Clients are required to support both discovery methods, not just one, because the spec doesn't get to assume which one any given authorization server will use.

Strung together, the full flow runs in six steps: the client hits the server and gets a 401 pointing at a Protected Resource Metadata document; the client fetches that document and finds the list of authorization servers and supported scopes; the client fetches the authorization server's own metadata to find its endpoints; the client registers itself, through CIMD, through pre-registration, or through DCR as a fallback; the user runs through a standard OAuth 2.1 authorization code flow with PKCE and the client walks away with an access token; and from that point forward, every MCP request carries that token as a Bearer credential. None of those six steps is optional, and skipping one doesn't simplify the system, it just moves the failure somewhere harder to find.

RFC 8707 resource indicators tie a token to the specific server it was issued for, and any organization running more than one MCP server behind a shared authorization server needs them, because without that binding a token minted for Server A can be replayed against Server B. That's the default behavior of OAuth tokens unless someone explicitly scopes them to a resource, and it's also the first crack in the one-server-at-a-time mental model, widening once an organization runs a fleet of these things rather than one.

Diagram: The Six-Step OAuth Flow Every MCP Server Now Requires. Visualizes: Visualize the mandatory six-step authorization sequence that the July 2026 MCP spec requires for HTTP servers.

Scope design as the practical expression of least privilege

The spec is specific about how a client should decide what scope to request, and the order matters. If the 401's WWW-Authenticate challenge names a scope, that scope is authoritative and the client uses it. If it doesn't, the client falls back to the full list in scopes_supported from the Protected Resource Metadata document. If that list doesn't exist either, the client omits the scope parameter. The spec is blunt about which of these takes priority when they conflict: the challenged scope wins.

The instinct to request one broad scope and call it done is exactly backward. Access should be broken apart per tool or per capability wherever that's practical, with the resource server checking the required scope on every route or every tool call, not once at the connection level. That distinction matters because MCP tool calls don't map one-to-one onto REST-style API endpoints. A single MCP connection might expose a dozen tools doing very different things, so validating scope once, at connection time, tells you nothing about what any individual tool call is actually allowed to do. The spec's own term for the balance it's after is "Agent Experience," or AX: cut unnecessary friction for the agent without loosening the boundaries that keep it contained. Each call should carry exactly the permission it needs and nothing more.

When a client needs to step up its access, re-authorization is supposed to bundle the new scope in alongside the scopes already granted, rather than restarting the whole consent process from zero. That's a deliberate design choice, closer to a step-up authentication flow than a full re-consent cycle, and it matters because without it, every scope expansion risks silently dropping permissions a working integration already depended on.

None of this is free. Splitting scopes per tool multiplies the number of things that need separate configuration: every MCP server acting as its own resource server needs its own Protected Resource Metadata, its own token audience, and its own RFC 8707 resource indicator setup. For a five-person startup running one internal tool, that's an afternoon of config. For a platform team running thirty MCP servers across a dozen teams, that's an ongoing maintenance line item with its own on-call rotation.

That granularity has a second, quieter consequence: it turns scope design into tool design. A tool built to accept a raw SQL string or a raw shell command can't be scoped meaningfully, because there's no way to look at the argument and know what it's actually going to do, which also means there's no way to review it in a confirmation dialog before it runs. A tool built to accept a customer_id can be scoped tightly and reviewed at a glance, because the argument itself tells you the blast radius. Scope enforcement and human review are solving the same problem from two different directions, and both depend on the same upstream decision: how narrowly the tool's interface was designed.

Per-invocation authorization and human-in-the-loop gates as production requirements

Checking permissions once, at connection time, isn't enough for a system where an agent can fire off dozens of tool calls a second. Each individual invocation needs to be checked against whatever the requesting agent is actually allowed to do right now, which can shift based on context, user identity, or what's happened earlier in the session. The Cloud Security Alliance's guidance on agentic MCP security recommends just-in-time permission escalation for high-privilege operations specifically because session-level permission, granted once at connection time, doesn't shrink back down when the operation calls for less trust.

The reason this matters more for agents than it ever did for humans comes down to speed. An agent inherits access scoped for a human user and then burns through it at machine speed, with none of the natural pauses a human operator builds in just by being a person who occasionally hesitates. Most of the access control models currently in production were designed around the assumption that a human would notice something was wrong and stop. Agents don't stop on their own.

The MCPoison vulnerability, tracked as CVE-2025-54136, is what that failure looks like in practice. Cursor approved an MCP configuration once and never checked it again, so a configuration that was legitimate at the moment of approval turned into a standing attack vector the instant someone swapped out the payload behind it. Per-invocation validation would have caught the swap on the very next call. A one-time approval checked nothing about what happened after.

A sensible default for most enterprises, drawn from current enterprise MCP security guidance, splits the world into three buckets: auto-approve reads, require explicit approval for writes, and never auto-approve anything that can't be undone. That policy only works, though, if the tool arguments are actually readable by the human being asked to approve them, which loops straight back to the tool design point from the previous section. A reviewer staring at a raw SQL string has no real basis for a yes or no. A reviewer looking at a customer_id and an action name does.

The new server-initiated request rules in the spec give this kind of gate somewhere to run. A server can only issue a request back to the client while it's actively processing one of that client's requests (SEP-2260), and the MRTR pattern makes a confirmation step possible without holding a connection open the old stateful way did. The policy itself, deciding what gets auto-approved, what gets a human in the loop, and what never runs without one, is not something the spec can hand down from above. It has to be built and run as infrastructure that sits on top of the protocol, which is exactly the gap Enterprise Managed Authorization is built to close from the identity side.

Enterprise Managed Authorization

Enterprise Managed Authorization, promoted to stable on June 18, 2026, hands the authorization decision over to the organization's identity provider rather than leaving it scattered across individual OAuth grants. It sits on top of the OAuth 2.1 foundation the spec already requires. It doesn't replace any part of that foundation.

The mechanism runs through something called an Identity Assertion JWT Authorization Grant, or ID-JAG. During single sign-on, the client pulls an ID-JAG from the identity provider and trades it for an access token issued by the MCP server's own authorization server. The user never sees a separate per-server consent screen, because the identity provider already vouched for them once.

That one change solves two problems that individual OAuth grants handled badly. On the provisioning side, an admin approves a server a single time and every employee who's supposed to have access gets it automatically, no per-user OAuth dance required for each new hire. On the deprovisioning side, pulling someone's access in Okta cuts them off from every EMA-connected MCP server at once. Under the old model, an offboarded employee's access had to be revoked server by server, by hand, a process that's easy to forget and nearly impossible to audit after the fact.

EMA's current limits are the following. At launch, Okta is the only identity provider it works with. An organization running a different IdP has nothing to plug EMA into yet, which is a real constraint for a lot of enterprises and not a small print footnote. Client support at launch covers Claude, including Claude Code and Cowork, along with VS Code. Server support at launch covers Asana, Atlassian, Canva, Figma, Granola, Linear, and Supabase. That's a credible starting roster, but it's a starting roster, and any enterprise evaluating EMA today is evaluating what it covers right now, not what the architecture might eventually grow into.

Sources

  1. The 2026-07-28 MCP Specification Release Candidate
  2. Authorization - Model Context Protocol
  3. Understanding Authorization in MCP - Model Context Protocol
  4. The 2026-07-28 Specification
  5. Key Changes - Model Context Protocol
  6. Agentic MCP Security Best Practices Guide
  7. Evolving OAuth Client Registration in the Model Context Protocol
  8. Enterprise-Managed Authorization: Zero-touch OAuth for MCP

More in Access Control for Agents