What an MCP Gateway Actually Does and Why Security Matters

The Model Context Protocol (MCP) defines a JSON-RPC interface that lets large language model agents call external tools, fetch resources, and subscribe to prompts. Because every server an agent can reach is a potential code-execution surface, most enterprises now sit an MCP gateway in front of that fan-out, much like an API gateway sits in front of microservices. Cloudflare's reference architecture describes the gateway as the policy enforcement point for tool discovery, request validation, and audit logging, while AWS positions Amazon Bedrock AgentCore Gateway as a managed equivalent that collapses dozens of individual MCP servers into a single, observable endpoint. Wiz's research briefing frames the gateway as the single chokepoint where identity, transport encryption, and content inspection can be enforced consistently, because a scattered fleet of MCP servers is operationally untenable.

Also worth reading: How Do You Troubleshoot LLM Access Denied Errors in Enterprise AI Deployments? · How Should Organizations Structure an Enterprise AI Evaluation Checklist for 2026 Deployments? · How do you architect an enterprise agentic ai policy engine design for autonomous multi-model deployments?

The risk profile is unusually wide. A compromised MCP server can issue tool calls that exfiltrate data, modify systems, or simply waste tokens. SOC Prime's taxonomy of MCP risks separates prompt-injection carried through tool outputs, confused-deputy actions against the underlying API, and supply-chain attacks against the MCP server binary itself. Treating the gateway as a thin proxy will not contain any of these. The gateway must own authentication, authorization, schema validation, and rate limiting, and it must do so in a way that survives both legitimate and malicious traffic.

Authentication, Identity, and Mutual TLS

The first best practice is to never expose an MCP server without strong, scoped identity. OAuth 2.0 with PKCE is the de facto standard for the client-to-gateway leg, because it allows the agent runtime to obtain a short-lived token tied to the end user, the agent identity, and the requested tool scopes. The gateway then brokers a separate identity to the upstream MCP server, which is commonly a service account or workload identity from a cloud IAM system. The two-leg model keeps the end user's identity out of the server's trust store and prevents a leaked server credential from impersonating a human.

Mutual TLS between the gateway and the upstream MCP server is the second layer. mTLS pins the server's certificate to a known CA, which neutralizes many on-path attacks and DNS hijacks. InfoQ's write-up of a least-privilege AI agent gateway shows the pattern working at scale: every MCP server is fronted by an OPA policy decision point, and every request is mTLS-authenticated with short-lived certificates minted by an internal CA. The result is that a stolen cookie or API key cannot be replayed against the gateway from outside the service mesh.

Rotation discipline matters as much as the handshake. Tokens should live for minutes, not hours, and server certificates should rotate on the order of days. Any token that survives longer than an agent run is a liability, because the agent may have already been decommissioned by the time it expires.

Authorization, Scoping, and Least Privilege

Authentication proves who is calling. Authorization decides what they may do, and in an MCP context that is the harder problem. Each MCP tool should carry a JSON Schema describing its inputs, and the gateway should bind every tool invocation to a scope that names the resource, the verb, and the maximum data volume. OPA, Cedar, and Rego are common policy engines, and InfoQ's reference design shows OPA evaluating a policy on every request in under five milliseconds, which is fast enough to keep in the hot path.

Least privilege is not a slogan; it is a measurable property. A reasonable target is that no agent identity holds scopes for more than five tools, that no tool scope allows both read and write on the same resource, and that destructive tools (delete, drop, revoke) require a separate, human-approved break-glass flow. SOC Prime's risk catalog recommends a deny-by-default posture with an explicit allow list, because an allow-by-default posture in MCP has historically led to agents discovering and calling tools the operator did not know were exposed. A weekly review of granted scopes, ideally automated, catches drift before it becomes an incident.

Input Validation, Output Sanitization, and Prompt Injection

A surprising share of MCP incidents in 2025 originated not from a malicious server but from a malicious response. Tool outputs are fed back into the model as context, and a tool that returns a page fetched from the public web can carry an indirect prompt injection payload. The gateway is the natural place to inspect tool outputs for known injection patterns, redact obvious payloads, and quarantine responses that exceed a token or byte budget. Wiz's briefing recommends a two-pass scanner: a fast regex pass on every response, and a slower model-based classifier on a sampled subset.

Input validation on the way in is equally important. MCP servers commonly expect JSON, but the schema is enforced by the server, not the transport. The gateway should validate the JSON-RPC envelope, confirm the tool name is on the allow list, check the arguments against a registered schema, and reject anything that does not conform before a single byte reaches the server. AWS AgentCore Gateway implements this as a contract-first model: tools are registered with a schema, and calls that violate it are rejected at the edge with a structured error. The result is a much smaller blast radius when a tool is replaced or its API changes.

Observability, Audit Logging, and Cost Controls

A gateway is only as good as its logs. Every request should produce a structured record with a correlation ID, the agent identity, the tool name, a hash of the arguments, a hash of the response, the policy decision, and the latency. Logs that are not queryable are not logs; they are storage. Most teams ship these to a SIEM within seconds, with alerts on unusual tool sequences (for example, a read followed by a delete on the same resource within ten seconds) and on sudden spikes in error rates that may indicate an agent loop.

Cost is a security control in this domain. An agent that calls an expensive tool in a tight loop can exhaust a budget in minutes. The gateway should enforce per-agent and per-user rate limits, both in requests per minute and in estimated dollar cost, and it should emit a metric that the FinOps team can chart alongside the security alerts. A common pattern is a soft limit at 80 percent of the daily budget, a hard cap at 100 percent, and a human-in-the-loop override that requires MFA. This converts a runaway agent from a security incident into a budget conversation, which is a much cheaper problem to have.

Practical Rollout: A 90-Day Plan

The fastest credible path to a secure MCP deployment fits in a quarter. Days 1 through 15 should focus on discovery: inventory every existing MCP server, classify each by data sensitivity, and pick a gateway product. The middle of the quarter is for hardening: deploy the gateway in front of the highest-value servers, wire it to the existing identity provider, write the first OPA policies, and turn on structured logging. Days 60 through 75 should bring the rest of the fleet behind the gateway, with parallel-run comparisons to catch regressions. The last two weeks are for tabletop exercises: simulate a prompt-injection response, a stolen token, and a misconfigured scope, and confirm the alerts, the runbooks, and the on-call rotation all line up.

A common mistake is to ship the gateway as a pure observability layer and defer policy enforcement. That is a reasonable first step for buy-in, but it must be followed by enforcement within a single quarter, or the gateway becomes shelfware. A second common mistake is to centralize every MCP server behind a single gateway instance, which creates both a performance bottleneck and a high-value target. Cloudflare's architecture shows the opposite: multiple regional gateway clusters, each enforcing the same policy bundle, fronting a smaller set of upstream servers.

Comparison: Gateway Approaches

FeatureSelf-Hosted Gateway (e.g., OPA + Envoy)Managed Gateway (AWS AgentCore)Sidecar per MCP Server
Operational overheadHigh (you run the control plane)Low (AWS owns uptime)Moderate (one per server)
Policy expressivenessVery high (full Rego)Medium (managed policy DSL)High (local OPA)
Latency overhead1-5 ms per request3-10 ms per requestUnder 1 ms per request
Best fitRegulated industries, custom policyTeams already on AWS, fast pilotSmall fleets (<10 servers)
Cost modelEngineering time + infraPer-request or per-toolEngineering time
Failure blast radiusCluster-wide if sharedCluster-wide if sharedIsolated to one server
Each row reflects a real trade-off, not a preference. A regulated bank will usually pick the self-hosted column because the policy engine is auditable down to the line. A startup running three MCP servers will pick the sidecar column because the operational tax of a central control plane exceeds the benefit. The managed column is the path of least resistance for teams that already trust their cloud provider with the data in question.

Common Mistakes and When to Act

The mistake list is short and well-rehearsed. First, treating MCP servers as internal APIs and skipping threat modeling; MCP is closer to running untrusted code than to calling a microservice. Second, allowing agents to register new tools at runtime without a human review; this is convenient and dangerous. Third, logging only the request and not the response, which makes post-incident forensics guesswork. Fourth, ignoring egress control: an MCP tool that fetches a URL can exfiltrate data just as easily as it can fetch it, so the gateway should restrict outbound destinations to a known list. Fifth, leaving default credentials on MCP servers; this is still the most common finding in published audits.

The right time to act is before the second MCP server ships, not after the tenth. A gateway inserted at pilot stage costs roughly a week of engineering time; a gateway inserted after a fleet of fifty servers has been deployed costs a quarter of migration effort and several awkward compliance conversations. As of mid-2026, every major cloud provider ships a managed MCP gateway option, and the open-source projects (Envoy-based gateways, the Cloudflare reference, the MCP Inspector) are mature enough that the only remaining barrier is internal alignment. Pilot-stage governance is the cheapest, most defensible posture, and it is what enterprise evaluation platforms like enterpriseailabs.io are designed to support with shared policy templates, evaluation harnesses, and audit-ready logging.