Large Language ModelsGenerate imagesGenerate videos
Deploy an MCP Server to AWS Lambda, Azure and Cloud Run Side by Side
One stateless TypeScript MCP server, three hosts. See the exact Lambda Web Adapter setup, the Azure Functions host.json for self-hosted servers and the Cloud Run deploy command, plus the auth options, timeouts and cold start trade-offs that decide which platform fits your project.
Your MCP server runs fine on your laptop over stdio. Then a teammate asks for a URL they can paste into a client, and the real work begins. A remote server needs HTTPS, authentication, a transport that survives load balancers, and a host that does not bill you while nobody is calling it. This article takes one small TypeScript server and puts it on three platforms: AWS Lambda, Azure Functions and Google Cloud Run. You get the config that matters on each one, the access control that keeps strangers out, and a plain comparison so you can pick a host in ten minutes instead of a week.
💡 Scope: every snippet below assumes the Streamable HTTP transport. Stdio is for local child processes, so a server that only speaks stdio needs an HTTP front end before any of these hosts can run it.
Pick the Transport Before the Cloud
Why Stateless Wins on Serverless
Serverless platforms start and stop instances whenever they like. Request one lands on instance A, request two on instance B, and request three triggers a cold start on instance C. If your server keeps a session in memory, that sequence breaks it.
The fix is a stateless Streamable HTTP server: one /mcp endpoint that accepts a POST, answers, and forgets. All three platforms are built around that shape. Azure's self-hosted preview only accepts stateless servers on the streamable-http transport. Cloud Run documents SSE and Streamable HTTP as its two remote options, with built-in HTTP response streaming. Lambda behaves the same way once you put a web adapter in front of it.
What the July 2026 Spec Changed
The 2026-07-28 revision of the MCP specification pushed in the same direction:
No protocol-level sessions. The Mcp-Session-Id header is gone from Streamable HTTP.
No handshake. The initialize exchange was removed, and every request now carries its protocol version and client capabilities in _meta.
State through handles. A server that needs memory between calls mints an explicit handle and passes it around as an ordinary tool argument.
No stream resumability. A broken response stream loses the in-flight request, and the client must send it again with a new request ID.
HTTP+SSE is deprecated. New work should use Streamable HTTP.
In practice you can write the server as plain request and response code and let the platform run as many copies as it wants. One caution: SDK releases trail spec revisions, so pin your SDK version and test against the clients you care about before you trust a deploy.
One Server, Three Targets
The Handler Every Host Shares
The demo server exposes two tools that front the PicassoIA developer API: one starts an image job, the other checks it. The API is Replicate-style and asynchronous, with a base URL of https://api.picassoia.com/v1, Bearer authentication, POST /models/{owner}/{name}/predictions to start a job and GET /predictions/{id} to read it. Splitting the work into start and check keeps every request short, which suits platforms that bill by the millisecond. The job below targets PicassoIA Image through its picassoia/picassoia-image slug.
A fresh server and transport per request is the stateless pattern from the SDK examples, and it costs almost nothing because registering two tools is cheap. Method names shift between SDK releases, so match the snippet to the version you install, and check the request and response fields against the PicassoIA API docs.
Secrets Stay Out of the Image
Bake nothing sensitive into the container. Read PICASSOIA_API_TOKEN from the platform's secret store: AWS Secrets Manager or SSM Parameter Store on Lambda, an app setting that points at a managed vault on Azure, and Google Secret Manager on Cloud Run.
💡 PicassoIA accounts allow 5 concurrent predictions, shared across tokens and MCP connections. Cap your host's fan-out with Lambda reserved concurrency, Cloud Run --max-instances and --concurrency, or the Azure maximum instance count, instead of finding the ceiling in production.
Deploy to AWS Lambda
Lambda Web Adapter Setup
The least invasive route runs your Express app unchanged through the Lambda Web Adapter. For a container image it is one extra line:
FROM public.ecr.aws/docker/library/node:22-slim
COPY --from=public.ecr.aws/awsguru/aws-lambda-adapter:1.1.0 /lambda-adapter /opt/extensions/lambda-adapter
ENV PORT=8080 AWS_LWA_INVOKE_MODE=response_stream AWS_LWA_READINESS_CHECK_PATH=/health
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY dist ./dist
CMD ["node", "dist/server.js"]
Prefer zip packages? Attach the adapter layer, set AWS_LAMBDA_EXEC_WRAPPER to /opt/bootstrap, and point the handler at a startup script. The adapter reads the port from AWS_LWA_PORT (it falls back to PORT, default 8080) and probes the readiness path before it forwards traffic.
Function URL and Response Streaming
Put a Function URL in front and set its invoke mode to RESPONSE_STREAM, matching the adapter variable above. The default buffered mode holds the whole response until the tool finishes, which defeats streaming. Lambda gives you up to 15 minutes per invocation and up to 10 GB of memory, far more than a start-and-check tool needs.
Two access paths exist:
AWS_IAM Function URL. Callers sign requests with SigV4. Good for service-to-service traffic, awkward for desktop MCP clients.
NONE plus your own check. Run OAuth inside the server or place an authorizer in front. A Cognito or Lambda authorizer through API Gateway is the usual choice, but its default integration timeout sits near 30 seconds, so long tool calls favor the Function URL.
The Serverless Framework v4 can wire all of this from a few lines of YAML:
mcp:
servers:
images:
server: index.ts
💡 That post flags two catches: interactive OAuth login needs a custom domain at the root rather than the default execute-api URL, and Cognito has no dynamic client registration.
Deploy to Azure Functions
The host.json That Matters
Azure runs SDK-built servers as custom handlers: the Functions host receives the request and proxies it to your process. Microsoft's self-hosted MCP documentation gives this minimal file for a TypeScript server, and the Node quickstart shows it in a working project:
The mcp-custom-handler profile turns on HTTP proxying, routes every path ({*route}) to your server and clears the route prefix, so /mcp arrives untouched. Make the port value match the port your server listens on. Test locally with func start, since the F5 debugger is not supported yet, then publish with func azure functionapp publish <APP_NAME>.
Preview Limits and Entra Sign-In
Read the fine print before you commit: this feature is in public preview. It supports stateless streamable-http servers only, written with the Python, TypeScript, C# or Java SDKs, and the app must run on the Flex Consumption plan. If you need state, Microsoft points you to the Functions MCP extension instead. Flex Consumption can keep always-ready instances to trim cold starts, at the cost of paying for idle capacity.
Authentication is where Azure shines. The platform's built-in server authentication implements the MCP authorization requirements for you: it issues the 401 challenge, publishes the Protected Resource Metadata document, and sends clients to Microsoft Entra ID to sign in. The docs' expanded host.json sets defaultAuthorizationLevel to anonymous and leaves sign-in to that platform layer, so switch it on before the URL goes anywhere public.
Deploy to Cloud Run
One Command From Source
Cloud Run needs the least ceremony. With a Dockerfile or a Node project in the folder:
Already have an image? gcloud run deploy --image IMAGE_URL --port PORT does the job. Cloud Run injects PORT, and the server must bind to 0.0.0.0, which the shared handler already does. The adapter line in the Dockerfile from the Lambda section is just an inert file here, so one image can serve both platforms.
Private by Default
A new Cloud Run URL requires the Cloud Run Invoker (roles/run.invoker) IAM role on every request. For a local client, Google's docs recommend a proxy that injects your identity:
gcloud run services proxy mcp-images --region us-central1 --port=3000
Then point the client at http://localhost:3000/mcp. Automated callers can send an OIDC ID token as Authorization: Bearer <token>, with the audience set to the service's run.app URL. Callers that run on Cloud Run have more options, including a sidecar, standard service-to-service authentication or Cloud Service Mesh. A public, consumer-facing server needs --allow-unauthenticated plus OAuth inside your app, and that is a decision to make on purpose, not by default.
Warm Instances and Timeouts
Cloud Run scales to zero by default. Add --min-instances 1 if cold starts hurt, and budget for the idle instance. Requests can run up to 60 minutes with --timeout (the default is 5 minutes), the longest ceiling of the three, and HTTP response streaming needs no extra switch.
Side by Side Comparison
Question
AWS Lambda
Azure Functions
Cloud Run
Packaging
Container image with the Web Adapter, or zip plus layer
Custom handler plus host.json
Container image or source deploy
Stateful servers
Avoid
Not in the self-hosted preview
Avoid
Longest request
15 minutes
Set by the Flex Consumption plan
60 minutes
Sign-in options
IAM Function URL, Cognito or Lambda authorizer
Built-in authentication with Entra ID
Invoker role or OIDC ID token
Warm instances
Provisioned concurrency
Always-ready instances
--min-instances
Status for self-hosted MCP
Works through the adapter
Public preview
Documented hosting path
Which Host Fits Which Team?
Already on AWS with spiky traffic: Lambda. You pay per request and nothing while idle.
Microsoft shop with Entra ID: Azure Functions. The built-in authentication saves you from writing an OAuth layer, as long as a preview feature is acceptable.
Small team with long tool calls: Cloud Run. Lowest ceremony and the longest timeout.
If you cannot decide, build one container image first. It runs on Cloud Run as is, runs on Lambda through the adapter, and the same code runs behind the Azure custom handler.
Test the Endpoint Before Clients Do
Run the MCP Inspector with npx @modelcontextprotocol/inspector, choose Streamable HTTP, paste your /mcp URL and list the tools. Then run the test people skip: call the URL without credentials.
curl -i -X POST "$URL/mcp" -H "Content-Type: application/json" -d '{}'
A 401 or 403 means the front door holds. Any other response means the request got past your auth, and your upstream account pays for whatever that caller does next.
3 Common Mistakes
Binding to localhost.127.0.0.1 works on a laptop and fails behind every one of these platforms. Bind to 0.0.0.0.
Keeping state in memory. A counter or cache that lives in the process vanishes on the next cold start. Use explicit handles or an external store.
Buffering the stream. Lambda's default invoke mode is buffered, and a proxy in the middle can do the same. If progress messages arrive in one lump, look for a buffer.
Draft and Illustrate With PicassoIA
The same platform that gives your server something to call can also write the code around it and make the images for its docs.
Use Claude Sonnet 5 on PicassoIA
A coding model gets you from the snippets above to a server that matches your own tools. Claude Sonnet 5 handles multi-step coding and tool-use tasks and reads images, so a screenshot of a failed deploy can go straight into the request.
Paste a prompt that names the transport, the tools and the host, for example: "Write a stateless Streamable HTTP MCP server in TypeScript with two tools, start_job and get_job, ready for Cloud Run."
Fill the system prompt once so every reply follows your rules: stateless, bind to 0.0.0.0, read PORT, no in-memory sessions.
Pick an effort level that matches the task, using the table below.
Leave max_tokens at the default 8,192 for single-file answers, and ask for one file at a time if a reply gets cut off.
Attach an image when you have a log screenshot. The max_image_resolution setting defaults to 0.5 megapixels and scales it down before sending.
Parameter
Suggested setting
Use it for
effort
low (default)
Config tweaks and one-line fixes
effort
high or max
Auth flows and bugs that touch several files
max_tokens
8192 (default)
One file per reply
system_prompt
Your hosting rules
Consistent output across a project
image
Error screenshot
Debugging deploy logs
For a second opinion on a tricky auth bug, run the same prompt through GPT 5.6 Sol and compare the two answers.
Generate Your Own Images
Once the server is live, it needs a README header, a diagram background and a social card. PicassoIA Image turns a plain prompt into a finished picture in seconds, with seven aspect ratios from 1:1 to 16:9, a lockable seed for reproducible results, JPG, PNG or WebP output and up to two variations per run. It is described as unlimited, with no per-image cap, so you can iterate freely. When a still deserves motion, PicassoIA Video animates it into a short clip.
Try this prompt: a quiet loft office at dusk, laptop open on an oak desk, soft window light, 35mm photograph, film grain. Change one detail, lock the seed, generate again and compare the two. Open Picasso IA, run your first prompt and see what your next image looks like. Every model lives at picassoia.com/en/all-models, so there is plenty to experiment with.