MCP Server Security Hardening for Agent Authorization
Hardening MCP servers requires OAuth, token exchange, and runtime checks beyond the spec.

MCP was built to solve a wiring problem: how to connect a language model to tools without writing a custom integration for every pair. Developers read "optional" as "skip it," and the result is that a large share of remote MCP servers still require no authentication at all, a majority carry command-injection vulnerabilities, and most handle credentials in plaintext.
Every hardening control described in this piece exists to answer one structural design fact: an MCP server is an untrusted third party until proven otherwise, regardless of how politely it was onboarded.
The six threat classes that define the MCP attack surface
The MCP attack surface splits into two families that chain together: classic application-security bugs sitting in server code, and model-layer manipulation delivered through content the model reads as instruction. Missing or weak authentication is the foundational failure in the first family. An unauthenticated server has no principal to authorize, no identity to log, and no scope to limit, so the entire authorization model collapses before enforcement can even begin.
Tool poisoning sits in the second family. Invariant Labs demonstrated this against Cursor in April 2025: hidden instructions buried in a tool description silently read and exfiltrated SSH keys and config files through what looked like an innocuous tool parameter. The MCP specification's notifications/tools/list_changed mechanism lets a server push updated tool definitions after a client has already connected, and because the base protocol sets no re-approval trigger, version pin, or content hash on tool definitions, a server can serve a clean description at approval time and swap in a malicious one later.
Confused deputy and token passthrough belong to the application-security family. Session IDs must be secure and non-deterministic, generated from a cryptographically secure random source rather than anything sequential or guessable, and the spec is explicit that MCP servers must not use sessions for authentication.
These threats chain. Three incidents confirm this is not a theoretical exercise. The Azure DevOps MCP authentication bypass, CVE-2026-32211 with a CVSS score of 9.1, surfaced in April 2026 and made API keys and tokens accessible with no credentials required. Microsoft Incident Response published a walkthrough on June 30, 2026, tracing a tool-poisoning-to-exfiltration attack through four stages: a silently modified tool description, dynamic re-trust without re-approval, agent execution, and exfiltration carried out through a call the system had already approved.
What the July 2026 spec revision standardized
The 2025-06-18 revision standardized the bottom half of enterprise authorization, including the RFC 9728 requirement, and the 2026-07-28 revision, the largest update since MCP's launch, pushed that hardening further while leaving the top half for engineering teams to build themselves.
| The spec now requires | The spec leaves to the implementer | |---|---| | OAuth 2.0 Protected Resource Metadata (RFC 9728) for automatic discovery of the correct authorization server | Tool-level RBAC: which identities may call which tools, per server | | Resource Indicators (RFC 8707) binding a token to its intended target server | Approval and human-in-the-loop queues for mutating calls like deploy, rollback, or delete | | A progressive, least-privilege scope model that starts minimal and elevates through WWW-Authenticate challenges | Structured audit-record schemas, currently a roadmap extension rather than core spec | | server/discover as a mandatory RPC servers MUST implement for advertising protocol versions and capabilities | Credential isolation, so server-side secrets never reach the agent or client | | | Multi-tenant isolation for MCP servers serving more than one organization |
The NSA's May 2026 Cybersecurity Information Sheet on MCP made a similar point from the regulatory side: traditional cybersecurity principles remain necessary, but agentic systems introduce risks, such as dynamic tool invocation, implicit trust relationships, and context sharing, that those principles alone were never built to catch. Everything from here forward in this piece is the construction work that sits on top of that floor.
Layer 1, Identity (implementing OAuth with PKCE and audience-bound tokens)
Every other control in this piece assumes you already have a validated identity. Without one, authorization, logging, and runtime enforcement are all theater. The 2026-07-28 spec now mandates a cold-start flow, and it runs in a fixed sequence. The server answers with a 401 Unauthorized response and a WWW-Authenticate header pointing to its Protected Resource Metadata document under RFC 9728. The resource server publishes its trust anchors at a well-known, machine-readable location so the agent never has to infer anything.
Every protected request then needs the same validation sequence applied in order: check the issuer, verify the signature, confirm expiration, match the audience, and check the granted scope. RFC 8707 Resource Indicators say the token must name this MCP server as its intended resource, and RFC 9207 says you need to verify the iss claim against the expected authorization server, so a token from the wrong identity provider cannot slip through as an IdP mix-up.
Public clients, desktop assistants and CLI agents among them, need PKCE on every authorization request, with the code verifier kept out of reach of any untrusted content the agent might process. The MCP client is the OAuth client, the MCP server is the OAuth resource server, and the authorization server is a separate entity that may be a completely different identity provider. You capture consent once at grant time, but the agent then acts across potentially hundreds of calls afterward, so keeping that consent fresh is a job for the layers above the protocol, not something OAuth itself was built to track.
Layer 2, Token handling (preventing passthrough, confused deputy, and scope creep)
A correctly issued, correctly validated token can still cause damage if it gets forwarded somewhere it was never scoped for. Forwarding a client's token unchanged to a downstream API lets a credential meant for one audience reach a system it has no business touching, and the MCP security guidance forbids the practice. So you exchange or mint a fresh credential for the downstream resource using OAuth 2.0 token exchange under RFC 8693, scoped narrowly to that one downstream service, and this preserves a clean audit trail.
Confused deputy attacks follow a specific mechanical path: a proxy using a static third-party client ID, allowing dynamic clients, and skipping explicit per-client consent gives a malicious server the opening to trick that proxy into leaking an authorization code. Automated integration tests should deliberately exercise the failure cases: a changed redirect URI, a missing state value, a reused authorization code, a token issued for the wrong audience, a client lacking explicit consent. Every one of those should fail closed, not silently pass.
The Five Eyes "Careful Adoption" guidance requires short-lived credentials for every agent identity, because long-lived tokens simply widen the blast radius of any single compromise, and that applies as much to stdio deployments using dedicated, narrowly scoped process identities as it does to anything running over HTTP.
Layer 3, Tool integrity (pinning manifests and detecting rug pulls before the model executes)
Authentication answers who connected. It says nothing about whether the tool the agent is about to invoke is the same tool a human reviewed and approved, and that gap is what tool poisoning and rug pulls exploit. Don't ask the model itself to judge whether a tool is safe based on the tool's own description, because the model reads that description as instruction, not as metadata to be skeptical of.
Hashing tool descriptions and schemas at approval time, then re-prompting for human re-approval on any hash mismatch, is the direct countermeasure to the rug pull. A server showing a clean description at approval time and a malicious one at execution time is invisible without this pinning step in place.
Tool outputs need the same suspicion as tool descriptions: arguments, returned data, and error strings are all potential injection channels, so server-side validation should check ranges and allowed values, not just types, and high-impact calls should pass through an out-of-band check before the model acts on whatever came back. Invariant Labs' open-source scanner, mcp-scan, flags tool poisoning, prompt injection, rug pulls, and cross-origin escalation attacks in MCP configuration files, though scanning at registration time is not sufficient on its own: the rug-pull window opens after approval, so re-scanning on every session start matters just as much. Red Hat's RHEL MCP servers show what this looks like organizationally. They ship read-only by default, with an optional guarded command-execution mode where a gatekeeper component reviews proposed scripts, a human approves changes through an MCP Apps compatible client, and systemd-run constrains permissions, a pattern Red Hat presented at its 2026 Summit.
Layer 4, Execution isolation (confining what a tool process can reach)
Tool integrity controls verify intent. Tool processes should run as non-root, with narrow filesystem mounts, explicit resource limits, and separate process accounts for each tool. systemd-run or an equivalent sandbox primitive constrains what a process can execute, the same pattern Red Hat's guarded command-execution mode uses in production.
SSRF defense requires validating every tool-supplied URL before the server fetches it, blocking private IP ranges, loopback addresses, and cloud metadata endpoints like 169.254.169.254. An OX Security disclosure in April 2026 showed the STDIO transport executing OS commands without sanitization; if you want to fix it, cut shell-out patterns where you can, use parameterized execution APIs, and validate every argument strictly before any exec call runs. An MCP endpoint remains an HTTP or process boundary reachable by any caller, not only the model sitting in the loop, so a direct attacker with network access to the port is a threat in its own right.
AWS Bedrock AgentCore Gateway and Google Cloud's managed MCP servers both offer gateway layers with IAM controls and network isolation in front of tool servers, giving self-hosters a working pattern to copy.
Layer 5 — Runtime enforcement: RBAC, human-in-the-loop gates, and policy that fires before execution
Spec compliance and tool integrity establish what an agent is allowed to do in principle. Runtime enforcement stops a technically compliant agent from doing something its token permits but its policy prohibits, and it has to do that at machine speed, before the action completes. The 2026-07-28 spec defines no RBAC model, so you have to put tool-level access control, which decides which identities may call which tools on which servers, in the gateway or the server's own policy layer. So you map a validated token to an agent identity, the agent identity to a human principal, and that principal to a permitted tool set, and you map IdP groups to tool-level permissions and set session lifetimes for delegated authority explicitly.
Human-in-the-loop gates are where the spec's caution and operational reality diverge. The NSA CSI and the Five Eyes "Careful Adoption" guidance treat human checkpoints as mandatory for destructive or irreversible operations: deploys, rollbacks, deletes, any write to production. The MCP spec uses "SHOULD" rather than "MUST" for human oversight, and a year of incidents, the Claude Code RCE chief among them, shows that "SHOULD" gets read as "skip it for the demo" often enough to treat it as a defect. Treat human approval as mandatory for anything mutating, full stop, regardless of what the spec's language technically permits. Here is what a structured audit record for an agent-triggered deploy needs to look like.
The gateway is where this policy actually fires. The counter-argument is that MCP relies on persistent, stateful, bidirectional JSON-RPC streams over STDIO or WebSockets, which a generic reverse proxy may not handle correctly. Snowflake's acquisition of Natoma for enterprise MCP governance, agent identities, permissions, and audit trails is a sign this layer has become a purchased capability, not a research curiosity, for organizations without a platform team to build it from scratch.
Layer 6, Audit telemetry (covering what to log and in what structure)
Compliance pressure on MCP is arriving from outside the spec process. SOC 2 auditors are starting to ask about MCP server inventories, ISO 27001 reviews want to see tool allowlists, and PCI auditors want a straight answer on whether agents have any path to cardholder data. None of that pressure comes from the protocol itself: it says nothing about audit-record schemas and leaves them as a roadmap extension, not a core requirement.
Every invocation needs to capture event type, tool name, server identity, agent identity, the human principal the agent acted on behalf of, a hash of the arguments rather than the arguments themselves in plaintext, the result, a trace ID, and a timestamp. The authorization decision itself, meaning whether the token validated, the scope checked out, and the RBAC match succeeded, should be logged as its own event, separate from the tool's result, since the two carry distinct forensic value during an incident review.
If a record is not created at the moment of invocation, it cannot be reconstructed afterward. The audit trail needs to stream into existing SIEM tooling, and revocation events (a token revoked, an agent terminated, a tool pulled from the allowlist) deserve first-class entries of their own, because nobody should have to account later for silent state changes.
Putting the layers together: a reference architecture for a hardened MCP deployment
The six layers compound. If your deployment has working RBAC but no structured audit telemetry, it cannot prove anything happened the way it was supposed to once an incident review starts.
If you want to hold all six together, you put a single gateway in front of every tool server. It terminates OAuth and validates identity (Layer 1), checks token audience and blocks passthrough (Layer 2), maintains the approved tool registry and checks hashes on every session start (Layer 3), routes calls into isolated execution environments (Layer 4), enforces RBAC and sends mutating calls to the human approval queue (Layer 5), and writes structured records out to the SIEM pipeline (Layer 6). That gateway has to handle persistent, stateful, bidirectional JSON-RPC streams correctly, not just ordinary HTTP request-response pairs, to work across MCP's STDIO and WebSocket transports without breaking the streaming semantics underneath. The 2026-07-28 spec's stateless core makes header-based routing possible, so that policy enforcement can happen without requiring changes on every individual tool server behind it.
A server inventory sits underneath this as the foundation, and none of it works without that in place. Since server/discover is still a roadmap item rather than a universal registry, the allowlist has to be maintained as a first-party artifact, not pulled from any public source, and the network should be scanned regularly for MCP servers that were never approved and never made it onto that list, exactly as the NSA and CISA guidance recommends. Identity binds the whole structure together: each agent carries a verified, cryptographically anchored identity with short-lived credentials, as the Five Eyes guidance requires, and the gateway performs the identity join the spec itself never defined, from token to agent identity to


