Large Language ModelsGenerate imagesGenerate videos
Host MCP Server Free: Vercel, Cloudflare and Docker Options
Three free ways to put an MCP server online: Vercel Hobby, Cloudflare Workers and Docker hosts such as Render, Cloud Run and Hugging Face Spaces. Real limits checked in October 2026, working code, token security and a table showing which option breaks first.
Your MCP server runs perfectly on your laptop. Claude Desktop starts it over stdio, every tool answers, and then you ask a harder question: how do you let Cursor on another machine, a teammate, or a ChatGPT connector reach the same tools without paying for a server? You need a public HTTPS URL, and you would like the bill to read $0.
You can host an MCP server free on three very different platforms: Vercel, Cloudflare Workers and any Docker host with a free tier. Each one fits a different kind of server, and each one has a wall you will hit eventually. Below you get the real numbers, checked against the vendors' own docs in October 2026, working code for each path, and a short list of what breaks first.
💡 Free tiers change often. Every limit in this article comes from the official docs as of October 2026. Recheck the pricing page before you launch anything public.
Why Local Servers Stop Being Enough
Stdio versus remote HTTP
A stdio server is a child process. The client launches it, talks over standard input and output, and shuts it down when the session ends. Nothing leaves your machine, which makes stdio perfect for reading local files and useless for sharing.
A remote server listens at one URL, usually /mcp or /api/mcp, and speaks Streamable HTTP. Any client that has the address and the right token can call it from anywhere. The older HTTP+SSE transport is fading: Vercel's mcp-handler v2 removed it, and Cloudflare's current docs describe only Streamable HTTP on /mcp. Build for that transport and you will not face a rewrite next year.
What a free host must provide
Before comparing platforms, write down what your server actually needs:
A stable HTTPS URL you can paste into a client config
Fast wake-up, because a client with a short timeout fails the handshake if your server needs a minute to boot
Secret storage for the tokens your tools use
Enough CPU time per request for your tool logic, not for running a model
Readable logs for the day a call fails and nobody knows why
Hosting is a networking problem before it is a code problem: one address, one open port, one process that answers. Every option below solves those three things in a different way.
The Three Free Paths at a Glance
Vercel and Cloudflare run your code as serverless functions: no machine to manage, billed by usage, free up to a cap. Docker hosts run your code as a container: a full process with its own filesystem and any language you like. That one difference decides most trade-offs.
Option
Free allowance
Wake-up
Best for
Main catch
Vercel Hobby
1M invocations, 4 CPU hours, 300 s per call
Short, instances are reused
Next.js teams, fast deploys
Personal, non-commercial use only
Cloudflare Workers Free
100,000 requests per day, 10 ms CPU each
Negligible
Light tools, many calls
10 ms CPU, 50 subrequests
Render Free
750 instance hours per month
About 1 minute after 15 idle minutes
Any Dockerfile
Sleeps when idle
Google Cloud Run
2M requests, 180,000 vCPU seconds
Scales to zero, slower first call
Containers with bursty traffic
Needs a billing account
Hugging Face Spaces
Free CPU Basic (2 vCPU, 16 GB RAM)
Idle Spaces can pause
Public demos
Disk resets on restart
Read the table by rows, not by columns. The wake-up column predicts how the server feels inside a chat: Workers and Vercel answer in milliseconds, while a sleeping container makes the first tool call hang. The free allowance column predicts how long you stay free: 100,000 requests a day sounds huge until a chatty agent loops through a tool forty times per task.
💡 All five options cost $0 on day one, so price is the wrong tiebreaker. Choose by what a single tool call does: a quick lookup, a heavy computation, or a long wait on another API.
Vercel: Fastest Route From Repo to URL
If your project already lives in a Next.js App Router app, an MCP server is one extra file. Vercel runs it as a Function with Fluid compute, which bills active CPU separately from memory time. That suits MCP traffic well: long idle stretches, then a burst of calls.
Install and write the route
You need Node.js 20 or later. Vercel's docs pin these versions:
npm i mcp-handler@2.1.1 @modelcontextprotocol/server@2 zod@4
Create app/api/mcp/route.ts:
import { createMcpHandler } from 'mcp-handler';
import { z } from 'zod';
const handler = createMcpHandler((server) => {
server.registerTool(
'roll_dice',
{
description: 'Roll an N-sided die',
inputSchema: z.object({ sides: z.number().int().min(2) }),
},
async ({ sides }) => {
const value = 1 + Math.floor(Math.random() * sides);
return { content: [{ type: 'text', text: `You rolled a ${value}` }] };
},
);
});
export { handler as GET, handler as POST };
Push to Git, import the repo on Vercel, and the server answers at https://your-app.vercel.app/api/mcp. In Cursor the config is a single line: {"mcpServers": {"dice": {"url": "https://your-app.vercel.app/api/mcp"}}}. Test locally first with npx @modelcontextprotocol/inspector@latest, choose Streamable HTTP, and point it at http://localhost:3000/api/mcp.
Hobby limits that matter
1,000,000 function invocations per month
4 Active CPU hours and 360 GB-hours of provisioned memory
300 seconds maximum duration per function
Non-commercial, personal use only, per Vercel's fair use rules
Go over a limit and you usually wait 30 days before that feature works again
A tool that mostly waits on another API barely touches the 4-hour CPU allowance. A tool that parses large files burns through it fast. Plan around the 300-second ceiling too: return a job ID quickly and let a second tool check it, instead of holding one call open for minutes.
💡 If a client receives a login page instead of JSON, check Deployment Protection. Hobby projects can have Vercel Authentication switched on, and it blocks anonymous callers, including your MCP client.
Cloudflare Workers: Free With Edge Speed
Workers run inside V8 isolates, so there is no container to boot. That matches the stateless request and response shape of most MCP servers, and it is why the wake-up time is close to zero.
The template serves Streamable HTTP on /mcp. Cloudflare's docs now recommend createMcpHandler for new servers: it is stateless and builds one MCP server per request. The older McpAgent class is documented as deprecated and feature-frozen, so skip it for anything new.
Wrap the handler inside a fetch method. If you export the callable directly, Wrangler reads a function default export as a WorkerEntrypoint class and the deploy misbehaves.
Free plan numbers
100,000 requests per day, reset at midnight UTC, with a burst limit of 1,000 requests per minute
10 ms of CPU time per HTTP request
50 subrequests per request
128 MB of memory per isolate
Up to 100 Workers per account
Waiting on a fetch() call does not count toward CPU time, so a tool that calls an outside API and returns the answer fits comfortably. Heavy parsing or image processing does not. Past the daily cap, Cloudflare returns Error 1027.
Auth is already solved here. The OAuth templates support GitHub, Google, Slack, Auth0, Stytch, WorkOS and Cloudflare Access, with a KV namespace holding sessions. State, if you need it, can live in KV or in SQLite-backed Durable Objects, which are available on the free plan.
Docker: When You Need a Real Process
Pick a container when your server needs native libraries, local files, a long-running job, or a language other than TypeScript. You trade instant wake-up for freedom.
A minimal Dockerfile
FROM node:22-slim AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build
FROM node:22-slim
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY --from=build /app/dist ./dist
ENV PORT=8080
EXPOSE 8080
CMD ["node", "dist/http.js"]
Two rules matter. Listen on the port the host gives you by reading PORT, and bind to 0.0.0.0, not localhost. Then run the server in stateless Streamable HTTP mode so any instance can answer any request. Hosts that scale to zero will drop in-memory sessions without warning.
Render, Cloud Run and Spaces
Render Free: 750 instance hours per workspace each month, a single instance, no persistent disk, no SSH, and no outbound SMTP. Free Postgres databases expire after 30 days.
Google Cloud Run: the always-free tier gives 2 million requests, 180,000 vCPU-seconds, 360,000 GiB-seconds and 1 GB of North America egress each month. It scales to zero by default, but you must attach a billing account.
Hugging Face Spaces: add sdk: docker and app_port: 7860 to the YAML header of your README. The container runs as user 1000 and anything written to disk is lost on restart.
Cold starts break handshakes
Render's free service spins down after 15 minutes without traffic, and waking it takes about a minute. A client with a short timeout fails the first request, then works on the second. That bug is miserable to chase if you do not know the cause.
You have three options:
Accept it and retry the first call
Split the server: put latency-sensitive tools on Workers and heavy ones in the container
Ping a health endpoint every 10 minutes. One service then stays awake inside the 750 hours (a 31-day month has 744), but check the host's rules before you rely on it
Lock It Down Before Sharing the URL
An MCP URL is a remote control for whatever your tools can do. Anyone who finds it can press the buttons, and free tiers have no budget cap that stops a stranger from burning your quota.
Add a bearer token
On Vercel, wrap the handler with withMcpAuth, set required: true, and return the verified client from verifyToken. Requests without a valid token get a 401. A token missing a required scope gets a 403. The docs example reads a demo token from an environment variable, and says production should check issuer, audience, expiration and scopes against a real authorization server. mcp-handler does not issue tokens itself.
To satisfy the MCP spec, also publish OAuth protected resource metadata at /.well-known/oauth-protected-resource, so compliant clients can find your authorization server.
Keep secrets out of the repo
Vercel: project Environment Variables
Cloudflare:npx wrangler secret put PICASSOIA_TOKEN
Render and Spaces: dashboard secrets, read as environment variables at runtime
Never paste a token into a tool description or a prompt. Clients show tool descriptions to the model, and the model can repeat them. Add a per-token call counter too, since the free quota is shared by every caller.
Pick the Right Free Option
Pick Vercel if the project is already on Next.js, the traffic is personal, and you want a preview URL for every pull request.
Pick Cloudflare if you want near-zero wake-up, up to 100,000 calls a day, and tools that mostly forward requests to other APIs.
Pick a container host if you need native binaries, files on disk, or a job that runs longer than a quick request.
Nothing stops you from mixing them. A common setup is a Worker that answers cheap lookups and forwards heavy requests to a container, so the client always gets a fast first response. Start with the simplest platform that fits, measure for a week, and move only when a limit actually bites.
What breaks first
Platform
First wall you hit
Warning sign
Vercel Hobby
4 CPU hours, or the non-commercial rule
Paused features, 30 day wait
Workers Free
10 ms CPU on heavy tools
CPU limit errors on parsing
Render Free
Wake-up delay
First call fails after idle time
Cloud Run
Usage beyond the free allowance
Charges on the invoice
Spaces
Disk resets
State vanishes after a restart
Try It With Picasso IA
A hosted server is only useful if its tools do something worth calling. Image generation makes a good first tool, because you can see at once whether the whole loop works.
A tool that calls an image model
PicassoIA exposes a Replicate-style API at https://api.picassoia.com/v1. You authenticate with a Bearer token that starts with pia_sk_, create a prediction, then poll it. The PicassoIA Image model takes an input object with a prompt and an aspect_ratio:
On Workers, read the token from env instead of process.env. Predictions are asynchronous, so split the work into two tools: make_image returns the prediction ID, and check_image calls GET /v1/predictions/{id} and returns the status and output URLs. The statuses are starting, processing, succeeded, failed and canceled. Each call finishes in milliseconds, which fits inside the Workers 10 ms CPU budget and never holds a request open. The account limit is 5 predictions at once, shared across tokens and MCP connections.
💡 The API docs say predictions are currently free and use no credits, and they also say an Infinite plan is needed to create them. Check the pricing page before you build a public tool on top of it.
Deploy the dice server tonight, then open Picasso IA and make something with it. Write one prompt, run it through Seedream 5 Pro and Nano Banana 2 Lite, and compare the two results side by side. Once your own MCP tool returns an image URL on a server that costs nothing, you have the full loop working, and every new tool after that takes minutes. Try a few prompts, break a few limits on purpose, and see which free host holds up for your project.