Large Language ModelsGenerate imagesGenerate videos
MCP Gateway AWS vs Azure: Options and Setup Compared
An MCP gateway sits between your AI agents and your tools, so choosing where it runs matters. This comparison puts Amazon Bedrock AgentCore Gateway, Azure API Management, and Microsoft's open source MCP Gateway side by side on authentication, routing, cost, and setup effort.
Your agent runs fine with three tools. Then the team adds a billing API, a ticketing system, a search index and a few internal services, and every MCP client ends up carrying its own URLs, its own credentials and its own idea of a safe request. An MCP gateway fixes that by putting one controlled endpoint in front of every tool. The hard part is deciding where it lives. AWS offers a managed gateway built for agents, Azure extends an API management product many teams already run, and both clouds let you self-host a proxy when you want full control.
This comparison puts the real options side by side: what each one does, how authentication works, how billing is structured, and the exact steps to stand one up. Product names, tiers and limits move fast, so the details below follow the vendor documentation as of October 2026. Recheck that documentation before you commit to anything.
What an MCP Gateway Does
A gateway is a reverse proxy that speaks the Model Context Protocol. Clients send JSON-RPC messages such as tools/list and tools/call over Streamable HTTP, and the gateway decides who may call what, forwards each call to the right backend, and records what happened. Your AI agents never need to know whether a tool lives in a Lambda function, a REST API or a container in another region.
One Front Door for Tools
Without a gateway, each MCP client points straight at each MCP server. Ten agents and eight servers means eighty connections to configure, secure and monitor. With a gateway, ten clients talk to one endpoint and the gateway fans calls out to the servers behind it. Swapping a backend becomes a gateway change instead of a client rollout.
Where a Gateway Pays Off
Central authentication: validate tokens once instead of inside every server.
Rate limits and quotas: stop a looping agent from hammering your billing API.
Audit trail: one place to see which agent called which tool, and when.
Tool curation: expose the 12 operations agents need instead of all 200 your API has.
Mixed backends: Lambda functions, REST APIs and remote MCP servers behind a single tools/list.
💡 Tip: Ask who calls each tool and how often. If the honest answer is "one developer, locally," a gateway is premature. Once a second team or an outside customer shows up, it stops being optional.
The AWS Options
AgentCore Gateway Basics
Amazon Bedrock AgentCore reached general availability in October 2025, and AgentCore Gateway is its entry point for agent traffic. A gateway holds one or more targets in three categories:
MCP targets: the gateway acts as an MCP server and aggregates every MCP target into one virtual server, so clients see a single consolidated tools/list.
HTTP targets: traffic goes straight to the backend, such as another agent, with path-based routing and no aggregation or protocol translation.
Inference targets: LLM requests reach model providers through one endpoint, and the model field in the request picks the destination.
Every gateway needs an inbound authorizer. The options are OAuth with JWT, IAM with AWS Signature Version 4, an "authenticate only" mode that validates tokens and leaves authorization to the target, and no authorization at all for development. Production should use one of the first two.
Targets You Can Plug In
For MCP targets, AWS lists several source types:
Lambda functions: the gateway invokes your function and turns the response into MCP format.
API Gateway REST APIs you already run.
OpenAPI specifications: an existing REST API becomes a set of MCP tools.
Smithy models: useful for AWS services and custom APIs.
Remote MCP servers: the gateway can pass along prompts and resources as well as tools.
Outbound calls to OpenAPI and MCP targets go through credential providers, which can store API credentials or OAuth settings, sign with SigV4, or skip authentication for public endpoints. Lambda and Smithy targets use the execution role you attach. Two extras matter at scale: semantic tool search, which helps an agent find the right tool among hundreds, and capability synchronization for MCP targets.
Self-Run on ECS or EKS
Some teams skip the managed option and run an open source MCP proxy on ECS or EKS behind an Application Load Balancer. You own scaling, patching, token validation and logging, but you also get custom routing, one codebase you can reuse on any cloud, and no dependence on a single vendor's feature list. It suits platform teams with Kubernetes experience and becomes an expensive habit for everyone else.
The Azure Options
API Management as MCP Gateway
Azure API Management (APIM) offers two built-in ways to publish MCP servers:
REST API as MCP server: any REST API managed in APIM can be exposed, and its operations become MCP tools.
Existing MCP server: put APIM in front of a server built with LangChain, LangServe, Azure Logic Apps or Azure Functions.
Governance runs through the policy engine. Policies currently apply to all operations exposed as tools in a given MCP server, and Microsoft lists rate limiting and quotas, JWT validation with Microsoft Entra ID or another identity provider, IP filtering and response caching. Traffic shows up in Azure Monitor and Application Insights, and Azure API Center can act as a private registry so teams can find the servers that exist. Everything can also be managed as code with Bicep, Terraform, the Azure CLI or ARM templates.
Availability is broad: the classic Developer, Basic, Standard and Premium tiers, the v2 tiers (Basic v2, Standard v2, Premium v2), and the self-hosted gateway for your own infrastructure. Two limits matter. APIM supports MCP tools only, with no resources or prompts, and the MCP features are not available in workspaces.
Microsoft's Open Source MCP Gateway
Separate from APIM, the microsoft/mcp-gateway repository on GitHub is a reverse proxy and management layer for MCP servers on Kubernetes. Its data plane uses session-aware stateful routing, so requests that share a session ID reach the same server instance, and its control plane deploys, updates and deletes server adapters. Authorization relies on Entra ID roles, and the repo ships PowerShell and Bicep deployment paths onto AKS with Azure Container Registry and Cosmos DB.
💡 Spec change: The 2026-07-28 MCP specification revision removed protocol sessions, dropping the initialize handshake and the Mcp-Session-Id header. Session affinity still matters for servers on older revisions or servers that hold state in memory, but new servers no longer need it.
Backends You Can Put Behind It
Microsoft's own examples of existing servers behind APIM are LangChain, LangServe, Logic Apps and Functions. Any MCP-compatible server that APIM can reach over HTTP is a candidate, which makes the gateway a handy wrapper around a mixed estate of older and newer servers.
Side-by-Side Comparison
Feature
AgentCore Gateway
API Management
Microsoft MCP Gateway
Model
Managed AWS service
Managed Azure service, or self-hosted gateway
Open source, you run it on Kubernetes
Tool sources
Lambda, API Gateway REST, OpenAPI, Smithy, MCP servers
REST APIs in APIM, existing MCP servers
MCP servers you deploy to the cluster
MCP features
Tools, prompts, resources
Tools only
Depends on your servers
Inbound auth
OAuth (JWT), IAM SigV4, authenticate only
JWT from Entra ID or other providers, plus other methods
Entra ID role authorization
Billing
Per MCP operation, search query and indexed tool
By tier and scale units
AKS, registry and Cosmos DB resources you run
Ops workload
Low
Low to medium
High
Best fit
AWS-first teams, many mixed tool sources
Teams already using APIM for REST
Platform teams on Kubernetes
Identity. Both managed options validate JWTs, and APIM explicitly accepts tokens from Entra ID or other identity providers. AgentCore adds SigV4 for callers that already hold AWS credentials, which is handy for service to service traffic. Whatever you choose, test how the gateway answers an unauthenticated call. The MCP authorization model is built on OAuth, and a clean 401 that points clients to the right authorization server is what lets them sign in without hand-written config.
Billing. AgentCore Gateway charges for the calls your agents make through it, counted as MCP operations such as ListTools, CallTool and Ping, plus search queries and tools indexed for semantic search. APIM is priced by tier and scale units, so your bill follows capacity instead of calls. The Microsoft gateway is open source, but you pay for the cluster, the registry, the database and the people who patch them. Spiky, low volume traffic tends to favor per-operation pricing, and steady heavy traffic tends to favor flat capacity. Check each vendor's pricing page for current numbers.
💡 Tip: Because ListTools and Ping count as billable operations on AgentCore, chatty clients that refresh the tool list on every turn raise the bill. Cache the list on the client for a few minutes.
Setup on AWS and Azure
Both paths are quick for a first tool if your backend already exists. The steps below stay at the level the documentation guarantees, so check console labels and CLI flags against the current docs.
AgentCore Gateway Setup Steps
Choose the inbound authorizer. Point an OAuth (JWT) authorizer at your identity provider's OpenID configuration URL and allowed audiences, or pick IAM for AWS-native callers.
Create the gateway with that authorizer.
Add a target. Attach a Lambda function with its tool schema, upload an OpenAPI spec, or enter the URL of a remote MCP server.
Attach outbound credentials through a credential provider when the backend needs an API credential or OAuth token.
Copy the gateway's MCP endpoint into your client configuration.
Call tools/list and confirm every tool you expect appears, and nothing else does.
API Management Setup Steps
Confirm the tier. Use a classic or v2 tier that supports MCP, or the self-hosted gateway.
Create the MCP server from a managed REST API and pick which operations become tools, or point APIM at an existing MCP server.
Add inbound policies. Validate the token first, then limit the call rate, because a limiter placed before authentication lets anonymous traffic burn through the quota.
Wire up monitoring with Application Insights and add a correlation ID header.
Register the server in API Center so other teams can find it.
Servers on older spec revisions may expect an initialize call first and an Mcp-Session-Id header afterward, while a server on the 2026-07-28 revision does not. For anything beyond a smoke test, point the MCP Inspector at the gateway and run a real tool call, because a one-shot request will not show how streamed responses behave.
Mistakes That Break Rollouts
Exposing every operation as a tool. Models tend to pick worse tools from a list of 200 than from a list of 12. Curate on APIM, and lean on semantic tool search on AgentCore.
Rate limiting before authenticating. Anonymous callers should be rejected first, not counted.
Forwarding the client's token to backends. Let the gateway hold its own outbound credentials, so a backend never accepts a token that was issued for a different audience.
Assuming sessions exist. Design tools to be stateless. If a tool needs state, return an ID from a creation call and have the model pass it back as an ordinary argument.
Heavy policy logic on the MCP path. Anything that reads or buffers a streamed response can delay events. Test with a real client.
Ignoring the tools-only limit. If your servers rely on prompts or resources, APIM will not carry them today.
Whatever you pick, route the gateway's telemetry into the dashboards your team already watches before the first agent goes live. Tool-call volume, error rates and latency per tool tell you what agents actually use, and which of your 200 operations nobody ever calls.
Which One Should You Pick
Pick AgentCore When You Live on AWS
Choose it when your tools are Lambda functions, API Gateway REST APIs or OpenAPI specs, when you want one managed endpoint that aggregates them, and when IAM or OAuth already governs your callers. It also suits teams that want semantic tool search without building it, and servers that rely on prompts and resources.
Pick API Management If You're Microsoft-First
Choose it when your REST APIs already sit in APIM, your identity is Entra ID, and your platform team knows the policy engine. You get rate limits, quotas, caching and monitoring you already trust. Pick the Microsoft open source gateway only when you want Kubernetes-level control and can afford to operate it.
Running on both clouds? AgentCore can aggregate remote MCP servers, and APIM can front servers it reaches over HTTP, so a gateway in one cloud can sit in front of tools in the other. That works, but it adds latency, egress charges and a second identity system to reconcile. In practice, one gateway per cloud with a shared identity provider is usually simpler than one gateway stretched across both.
💡 Quick check: Count your tool sources, then your identity provider, then your traffic shape. Mostly Lambda and OpenAPI points to AgentCore. Mostly REST already in APIM points to API Management. A Kubernetes platform team with strict control needs points to the self-run option.
Make Your Own Images on Picasso IA
An MCP gateway is plumbing, and good plumbing is the part nobody sees. The same applies to the services your agents call. PicassoIA, for example, exposes a developer API at api.picassoia.com/v1 with bearer authentication, asynchronous jobs you create and then poll, and a cap of 5 concurrent predictions per account that API and MCP connections share. That kind of ceiling is exactly what a gateway should enforce for you: queue or throttle on your side instead of letting agents collect rejections.
Language models help with the setup work too. You can draft Bicep, Terraform or an OpenAPI spec for your gateway with Claude Sonnet 5, review a policy with GPT 5.6 Sol, or summarize a long vendor changelog quickly with Gemini 3.5 Flash. Treat their output as a first draft, and run it through the same review you would give a colleague's pull request.
Then there are the visuals. Architecture write-ups, internal wikis and launch posts all read better with a strong hero photo. Open Picasso IA, pick a text-to-image model such as Seedream 4.5, describe the scene you want in a few sentences, and generate several variations. Change the lens, the light or the setting, and compare. Your first image will not be your best one, and the fifth usually is. Browse every model at picassoia.com/en/all-models, and create your own images today.