Large Language ModelsGenerate imagesGenerate videos

Deploy an MCP Server to AWS Lambda, Azure and Cloud Run Side by Side

One stateless TypeScript MCP server, three hosts. See the exact Lambda Web Adapter setup, the Azure Functions host.json for self-hosted servers and the Cloud Run deploy command, plus the auth options, timeouts and cold start trade-offs that decide which platform fits your project.

Deploy an MCP Server to AWS Lambda, Azure and Cloud Run Side by Side
Cristian Da Conceicao
Founder of Picasso IA

Your MCP server runs fine on your laptop over stdio. Then a teammate asks for a URL they can paste into a client, and the real work begins. A remote server needs HTTPS, authentication, a transport that survives load balancers, and a host that does not bill you while nobody is calling it. This article takes one small TypeScript server and puts it on three platforms: AWS Lambda, Azure Functions and Google Cloud Run. You get the config that matters on each one, the access control that keeps strangers out, and a plain comparison so you can pick a host in ten minutes instead of a week.

💡 Scope: every snippet below assumes the Streamable HTTP transport. Stdio is for local child processes, so a server that only speaks stdio needs an HTTP front end before any of these hosts can run it.

Pick the Transport Before the Cloud

Why Stateless Wins on Serverless

Serverless platforms start and stop instances whenever they like. Request one lands on instance A, request two on instance B, and request three triggers a cold start on instance C. If your server keeps a session in memory, that sequence breaks it.

The fix is a stateless Streamable HTTP server: one /mcp endpoint that accepts a POST, answers, and forgets. All three platforms are built around that shape. Azure's self-hosted preview only accepts stateless servers on the streamable-http transport. Cloud Run documents SSE and Streamable HTTP as its two remote options, with built-in HTTP response streaming. Lambda behaves the same way once you put a web adapter in front of it.

What the July 2026 Spec Changed

The 2026-07-28 revision of the MCP specification pushed in the same direction:

  • No protocol-level sessions. The Mcp-Session-Id header is gone from Streamable HTTP.
  • No handshake. The initialize exchange was removed, and every request now carries its protocol version and client capabilities in _meta.
  • State through handles. A server that needs memory between calls mints an explicit handle and passes it around as an ordinary tool argument.
  • No stream resumability. A broken response stream loses the in-flight request, and the client must send it again with a new request ID.
  • HTTP+SSE is deprecated. New work should use Streamable HTTP.

In practice you can write the server as plain request and response code and let the platform run as many copies as it wants. One caution: SDK releases trail spec revisions, so pin your SDK version and test against the clients you care about before you trust a deploy.

Developer's hands sketching three connected boxes on a glass whiteboard

One Server, Three Targets

The Handler Every Host Shares

The demo server exposes two tools that front the PicassoIA developer API: one starts an image job, the other checks it. The API is Replicate-style and asynchronous, with a base URL of https://api.picassoia.com/v1, Bearer authentication, POST /models/{owner}/{name}/predictions to start a job and GET /predictions/{id} to read it. Splitting the work into start and check keeps every request short, which suits platforms that bill by the millisecond. The job below targets PicassoIA Image through its picassoia/picassoia-image slug.

import express from "express";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StreamableHTTPServerTransport } from "@modelcontextprotocol/sdk/server/streamableHttp.js";
import { z } from "zod";

const API = "https://api.picassoia.com/v1";
const auth = { Authorization: `Bearer ${process.env.PICASSOIA_API_TOKEN}` };

function buildServer() {
  const server = new McpServer({ name: "image-tools", version: "1.0.0" });

  server.registerTool(
    "start_image",
    { description: "Start an image generation", inputSchema: { prompt: z.string().max(4000) } },
    async ({ prompt }) => {
      const res = await fetch(`${API}/models/picassoia/picassoia-image/predictions`, {
        method: "POST",
        headers: { ...auth, "Content-Type": "application/json" },
        body: JSON.stringify({ input: { prompt, aspect_ratio: "16:9" } }),
      });
      const job = await res.json();
      return { content: [{ type: "text", text: JSON.stringify({ id: job.id, status: job.status }) }] };
    }
  );

  server.registerTool(
    "get_image",
    { description: "Check a generation by id", inputSchema: { id: z.string() } },
    async ({ id }) => {
      const res = await fetch(`${API}/predictions/${id}`, { headers: auth });
      return { content: [{ type: "text", text: await res.text() }] };
    }
  );
  return server;
}

const app = express();
app.use(express.json());

app.post("/mcp", async (req, res) => {
  const server = buildServer();
  const transport = new StreamableHTTPServerTransport({ sessionIdGenerator: undefined });
  res.on("close", () => { transport.close(); server.close(); });
  await server.connect(transport);
  await transport.handleRequest(req, res, req.body);
});

app.get("/health", (_req, res) => res.send("ok"));
app.listen(Number(process.env.PORT ?? 8080), "0.0.0.0");

A fresh server and transport per request is the stateless pattern from the SDK examples, and it costs almost nothing because registering two tools is cheap. Method names shift between SDK releases, so match the snippet to the version you install, and check the request and response fields against the PicassoIA API docs.

Secrets Stay Out of the Image

Bake nothing sensitive into the container. Read PICASSOIA_API_TOKEN from the platform's secret store: AWS Secrets Manager or SSM Parameter Store on Lambda, an app setting that points at a managed vault on Azure, and Google Secret Manager on Cloud Run.

💡 PicassoIA accounts allow 5 concurrent predictions, shared across tokens and MCP connections. Cap your host's fan-out with Lambda reserved concurrency, Cloud Run --max-instances and --concurrency, or the Azure maximum instance count, instead of finding the ceiling in production.

Top-down view of an oak desk with a laptop, notebook and tea

Deploy to AWS Lambda

Lambda Web Adapter Setup

The least invasive route runs your Express app unchanged through the Lambda Web Adapter. For a container image it is one extra line:

FROM public.ecr.aws/docker/library/node:22-slim
COPY --from=public.ecr.aws/awsguru/aws-lambda-adapter:1.1.0 /lambda-adapter /opt/extensions/lambda-adapter
ENV PORT=8080 AWS_LWA_INVOKE_MODE=response_stream AWS_LWA_READINESS_CHECK_PATH=/health
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY dist ./dist
CMD ["node", "dist/server.js"]

Prefer zip packages? Attach the adapter layer, set AWS_LAMBDA_EXEC_WRAPPER to /opt/bootstrap, and point the handler at a startup script. The adapter reads the port from AWS_LWA_PORT (it falls back to PORT, default 8080) and probes the readiness path before it forwards traffic.

Function URL and Response Streaming

Put a Function URL in front and set its invoke mode to RESPONSE_STREAM, matching the adapter variable above. The default buffered mode holds the whole response until the tool finishes, which defeats streaming. Lambda gives you up to 15 minutes per invocation and up to 10 GB of memory, far more than a start-and-check tool needs.

Two access paths exist:

  • AWS_IAM Function URL. Callers sign requests with SigV4. Good for service-to-service traffic, awkward for desktop MCP clients.
  • NONE plus your own check. Run OAuth inside the server or place an authorizer in front. A Cognito or Lambda authorizer through API Gateway is the usual choice, but its default integration timeout sits near 30 seconds, so long tool calls favor the Function URL.

The Serverless Framework v4 can wire all of this from a few lines of YAML:

mcp:
  servers:
    images:
      server: index.ts

💡 That post flags two catches: interactive OAuth login needs a custom domain at the root rather than the default execute-api URL, and Cognito has no dynamic client registration.

Long aisle of server racks in a data center with a technician walking away

Deploy to Azure Functions

The host.json That Matters

Azure runs SDK-built servers as custom handlers: the Functions host receives the request and proxies it to your process. Microsoft's self-hosted MCP documentation gives this minimal file for a TypeScript server, and the Node quickstart shows it in a working project:

{
  "version": "2.0",
  "configurationProfile": "mcp-custom-handler",
  "customHandler": {
    "description": {
      "defaultExecutablePath": "npm",
      "arguments": ["run", "start"]
    },
    "port": "8080"
  }
}

The mcp-custom-handler profile turns on HTTP proxying, routes every path ({*route}) to your server and clears the route prefix, so /mcp arrives untouched. Make the port value match the port your server listens on. Test locally with func start, since the F5 debugger is not supported yet, then publish with func azure functionapp publish <APP_NAME>.

Preview Limits and Entra Sign-In

Read the fine print before you commit: this feature is in public preview. It supports stateless streamable-http servers only, written with the Python, TypeScript, C# or Java SDKs, and the app must run on the Flex Consumption plan. If you need state, Microsoft points you to the Functions MCP extension instead. Flex Consumption can keep always-ready instances to trim cold starts, at the cost of paying for idle capacity.

Authentication is where Azure shines. The platform's built-in server authentication implements the MCP authorization requirements for you: it issues the 401 challenge, publishes the Protected Resource Metadata document, and sends clients to Microsoft Entra ID to sign in. The docs' expanded host.json sets defaultAuthorizationLevel to anonymous and leaves sign-in to that platform layer, so switch it on before the URL goes anywhere public.

Hand pressing a yellow fiber patch cable into a network switch

Deploy to Cloud Run

One Command From Source

Cloud Run needs the least ceremony. With a Dockerfile or a Node project in the folder:

gcloud run deploy mcp-images --source . --region us-central1 \
  --set-secrets PICASSOIA_API_TOKEN=picassoia-token:latest \
  --max-instances 3

Already have an image? gcloud run deploy --image IMAGE_URL --port PORT does the job. Cloud Run injects PORT, and the server must bind to 0.0.0.0, which the shared handler already does. The adapter line in the Dockerfile from the Lambda section is just an inert file here, so one image can serve both platforms.

Private by Default

A new Cloud Run URL requires the Cloud Run Invoker (roles/run.invoker) IAM role on every request. For a local client, Google's docs recommend a proxy that injects your identity:

gcloud run services proxy mcp-images --region us-central1 --port=3000

Then point the client at http://localhost:3000/mcp. Automated callers can send an OIDC ID token as Authorization: Bearer <token>, with the audience set to the service's run.app URL. Callers that run on Cloud Run have more options, including a sidecar, standard service-to-service authentication or Cloud Service Mesh. A public, consumer-facing server needs --allow-unauthenticated plus OAuth inside your app, and that is a decision to make on purpose, not by default.

Warm Instances and Timeouts

Cloud Run scales to zero by default. Add --min-instances 1 if cold starts hurt, and budget for the idle instance. Requests can run up to 60 minutes with --timeout (the default is 5 minutes), the longest ceiling of the three, and HTTP response streaming needs no extra switch.

Low-angle view of white clouds over a green hillside with a wind turbine

Side by Side Comparison

QuestionAWS LambdaAzure FunctionsCloud Run
PackagingContainer image with the Web Adapter, or zip plus layerCustom handler plus host.jsonContainer image or source deploy
Stateful serversAvoidNot in the self-hosted previewAvoid
Longest request15 minutesSet by the Flex Consumption plan60 minutes
Sign-in optionsIAM Function URL, Cognito or Lambda authorizerBuilt-in authentication with Entra IDInvoker role or OIDC ID token
Warm instancesProvisioned concurrencyAlways-ready instances--min-instances
Status for self-hosted MCPWorks through the adapterPublic previewDocumented hosting path

Which Host Fits Which Team?

  • Already on AWS with spiky traffic: Lambda. You pay per request and nothing while idle.
  • Microsoft shop with Entra ID: Azure Functions. The built-in authentication saves you from writing an OAuth layer, as long as a preview feature is acceptable.
  • Small team with long tool calls: Cloud Run. Lowest ceremony and the longest timeout.

If you cannot decide, build one container image first. It runs on Cloud Run as is, runs on Lambda through the adapter, and the same code runs behind the Azure custom handler.

Two engineers comparing printed sheets at a high wooden table

Test the Endpoint Before Clients Do

Run the MCP Inspector with npx @modelcontextprotocol/inspector, choose Streamable HTTP, paste your /mcp URL and list the tools. Then run the test people skip: call the URL without credentials.

curl -i -X POST "$URL/mcp" -H "Content-Type: application/json" -d '{}'

A 401 or 403 means the front door holds. Any other response means the request got past your auth, and your upstream account pays for whatever that caller does next.

3 Common Mistakes

  1. Binding to localhost. 127.0.0.1 works on a laptop and fails behind every one of these platforms. Bind to 0.0.0.0.
  2. Keeping state in memory. A counter or cache that lives in the process vanishes on the next cold start. Use explicit handles or an external store.
  3. Buffering the stream. Lambda's default invoke mode is buffered, and a proxy in the middle can do the same. If progress messages arrive in one lump, look for a buffer.

Extreme close-up of fingers typing during an endpoint test

Draft and Illustrate With PicassoIA

The same platform that gives your server something to call can also write the code around it and make the images for its docs.

Use Claude Sonnet 5 on PicassoIA

A coding model gets you from the snippets above to a server that matches your own tools. Claude Sonnet 5 handles multi-step coding and tool-use tasks and reads images, so a screenshot of a failed deploy can go straight into the request.

  1. Open Claude Sonnet 5 on PicassoIA.
  2. Paste a prompt that names the transport, the tools and the host, for example: "Write a stateless Streamable HTTP MCP server in TypeScript with two tools, start_job and get_job, ready for Cloud Run."
  3. Fill the system prompt once so every reply follows your rules: stateless, bind to 0.0.0.0, read PORT, no in-memory sessions.
  4. Pick an effort level that matches the task, using the table below.
  5. Leave max_tokens at the default 8,192 for single-file answers, and ask for one file at a time if a reply gets cut off.
  6. Attach an image when you have a log screenshot. The max_image_resolution setting defaults to 0.5 megapixels and scales it down before sending.
ParameterSuggested settingUse it for
effortlow (default)Config tweaks and one-line fixes
efforthigh or maxAuth flows and bugs that touch several files
max_tokens8192 (default)One file per reply
system_promptYour hosting rulesConsistent output across a project
imageError screenshotDebugging deploy logs

For a second opinion on a tricky auth bug, run the same prompt through GPT 5.6 Sol and compare the two answers.

Designer at a wide desk with a monitor showing a mountain photograph

Generate Your Own Images

Once the server is live, it needs a README header, a diagram background and a social card. PicassoIA Image turns a plain prompt into a finished picture in seconds, with seven aspect ratios from 1:1 to 16:9, a lockable seed for reproducible results, JPG, PNG or WebP output and up to two variations per run. It is described as unlimited, with no per-image cap, so you can iterate freely. When a still deserves motion, PicassoIA Video animates it into a short clip.

Try this prompt: a quiet loft office at dusk, laptop open on an oak desk, soft window light, 35mm photograph, film grain. Change one detail, lock the seed, generate again and compare the two. Open Picasso IA, run your first prompt and see what your next image looks like. Every model lives at picassoia.com/en/all-models, so there is plenty to experiment with.

Person at a window table in a quiet co-working space at dusk

Share this article