Large Language ModelsGenerate imagesGenerate videos

Host MCP Server Free: Vercel, Cloudflare and Docker Options

Three free ways to put an MCP server online: Vercel Hobby, Cloudflare Workers and Docker hosts such as Render, Cloud Run and Hugging Face Spaces. Real limits checked in October 2026, working code, token security and a table showing which option breaks first.

Host MCP Server Free: Vercel, Cloudflare and Docker Options
Cristian Da Conceicao
Founder of Picasso IA

Your MCP server runs perfectly on your laptop. Claude Desktop starts it over stdio, every tool answers, and then you ask a harder question: how do you let Cursor on another machine, a teammate, or a ChatGPT connector reach the same tools without paying for a server? You need a public HTTPS URL, and you would like the bill to read $0.

You can host an MCP server free on three very different platforms: Vercel, Cloudflare Workers and any Docker host with a free tier. Each one fits a different kind of server, and each one has a wall you will hit eventually. Below you get the real numbers, checked against the vendors' own docs in October 2026, working code for each path, and a short list of what breaks first.

Developer typing on a laptop at a kitchen table on a rainy evening

💡 Free tiers change often. Every limit in this article comes from the official docs as of October 2026. Recheck the pricing page before you launch anything public.

Why Local Servers Stop Being Enough

Stdio versus remote HTTP

A stdio server is a child process. The client launches it, talks over standard input and output, and shuts it down when the session ends. Nothing leaves your machine, which makes stdio perfect for reading local files and useless for sharing.

A remote server listens at one URL, usually /mcp or /api/mcp, and speaks Streamable HTTP. Any client that has the address and the right token can call it from anywhere. The older HTTP+SSE transport is fading: Vercel's mcp-handler v2 removed it, and Cloudflare's current docs describe only Streamable HTTP on /mcp. Build for that transport and you will not face a rewrite next year.

What a free host must provide

Before comparing platforms, write down what your server actually needs:

  • A stable HTTPS URL you can paste into a client config
  • Fast wake-up, because a client with a short timeout fails the handshake if your server needs a minute to boot
  • Secret storage for the tokens your tools use
  • Enough CPU time per request for your tool logic, not for running a model
  • Readable logs for the day a call fails and nobody knows why

Hand plugging an ethernet cable into a small black mini PC next to a white router

Hosting is a networking problem before it is a code problem: one address, one open port, one process that answers. Every option below solves those three things in a different way.

The Three Free Paths at a Glance

Vercel and Cloudflare run your code as serverless functions: no machine to manage, billed by usage, free up to a cap. Docker hosts run your code as a container: a full process with its own filesystem and any language you like. That one difference decides most trade-offs.

OptionFree allowanceWake-upBest forMain catch
Vercel Hobby1M invocations, 4 CPU hours, 300 s per callShort, instances are reusedNext.js teams, fast deploysPersonal, non-commercial use only
Cloudflare Workers Free100,000 requests per day, 10 ms CPU eachNegligibleLight tools, many calls10 ms CPU, 50 subrequests
Render Free750 instance hours per monthAbout 1 minute after 15 idle minutesAny DockerfileSleeps when idle
Google Cloud Run2M requests, 180,000 vCPU secondsScales to zero, slower first callContainers with bursty trafficNeeds a billing account
Hugging Face SpacesFree CPU Basic (2 vCPU, 16 GB RAM)Idle Spaces can pausePublic demosDisk resets on restart

Read the table by rows, not by columns. The wake-up column predicts how the server feels inside a chat: Workers and Vercel answer in milliseconds, while a sleeping container makes the first tool call hang. The free allowance column predicts how long you stay free: 100,000 requests a day sounds huge until a chatty agent loops through a tool forty times per task.

💡 All five options cost $0 on day one, so price is the wrong tiebreaker. Choose by what a single tool call does: a quick lookup, a heavy computation, or a long wait on another API.

Vercel: Fastest Route From Repo to URL

If your project already lives in a Next.js App Router app, an MCP server is one extra file. Vercel runs it as a Function with Fluid compute, which bills active CPU separately from memory time. That suits MCP traffic well: long idle stretches, then a burst of calls.

Install and write the route

You need Node.js 20 or later. Vercel's docs pin these versions:

npm i mcp-handler@2.1.1 @modelcontextprotocol/server@2 zod@4

Create app/api/mcp/route.ts:

import { createMcpHandler } from 'mcp-handler';
import { z } from 'zod';

const handler = createMcpHandler((server) => {
  server.registerTool(
    'roll_dice',
    {
      description: 'Roll an N-sided die',
      inputSchema: z.object({ sides: z.number().int().min(2) }),
    },
    async ({ sides }) => {
      const value = 1 + Math.floor(Math.random() * sides);
      return { content: [{ type: 'text', text: `You rolled a ${value}` }] };
    },
  );
});

export { handler as GET, handler as POST };

Push to Git, import the repo on Vercel, and the server answers at https://your-app.vercel.app/api/mcp. In Cursor the config is a single line: {"mcpServers": {"dice": {"url": "https://your-app.vercel.app/api/mcp"}}}. Test locally first with npx @modelcontextprotocol/inspector@latest, choose Streamable HTTP, and point it at http://localhost:3000/api/mcp.

Low-angle view of a developer working in a bright loft with white brick walls

Hobby limits that matter

  • 1,000,000 function invocations per month
  • 4 Active CPU hours and 360 GB-hours of provisioned memory
  • 300 seconds maximum duration per function
  • Non-commercial, personal use only, per Vercel's fair use rules
  • Go over a limit and you usually wait 30 days before that feature works again

A tool that mostly waits on another API barely touches the 4-hour CPU allowance. A tool that parses large files burns through it fast. Plan around the 300-second ceiling too: return a job ID quickly and let a second tool check it, instead of holding one call open for minutes.

💡 If a client receives a login page instead of JSON, check Deployment Protection. Hobby projects can have Vercel Authentication switched on, and it blocks anonymous callers, including your MCP client.

Cloudflare Workers: Free With Edge Speed

Workers run inside V8 isolates, so there is no container to boot. That matches the stateless request and response shape of most MCP servers, and it is why the wake-up time is close to zero.

Scaffold with one command

npm create cloudflare@latest -- remote-mcp-server-authless --template=cloudflare/ai/demos/remote-mcp-authless

The template serves Streamable HTTP on /mcp. Cloudflare's docs now recommend createMcpHandler for new servers: it is stateless and builds one MCP server per request. The older McpAgent class is documented as deprecated and feature-frozen, so skip it for anything new.

export default {
  fetch(request, env, ctx) {
    return createMcpHandler(createServer)(request, env, ctx);
  },
} satisfies ExportedHandler;

Wrap the handler inside a fetch method. If you export the callable directly, Wrangler reads a function default export as a WorkerEntrypoint class and the deploy misbehaves.

Symmetrical view down an aisle of server racks in a clean data center

Free plan numbers

  • 100,000 requests per day, reset at midnight UTC, with a burst limit of 1,000 requests per minute
  • 10 ms of CPU time per HTTP request
  • 50 subrequests per request
  • 128 MB of memory per isolate
  • Up to 100 Workers per account

Waiting on a fetch() call does not count toward CPU time, so a tool that calls an outside API and returns the answer fits comfortably. Heavy parsing or image processing does not. Past the daily cap, Cloudflare returns Error 1027.

Auth is already solved here. The OAuth templates support GitHub, Google, Slack, Auth0, Stytch, WorkOS and Cloudflare Access, with a KV namespace holding sessions. State, if you need it, can live in KV or in SQLite-backed Durable Objects, which are available on the free plan.

Docker: When You Need a Real Process

Pick a container when your server needs native libraries, local files, a long-running job, or a language other than TypeScript. You trade instant wake-up for freedom.

A minimal Dockerfile

FROM node:22-slim AS build
WORKDIR /app
COPY package*.json ./
RUN npm ci
COPY . .
RUN npm run build

FROM node:22-slim
WORKDIR /app
COPY package*.json ./
RUN npm ci --omit=dev
COPY --from=build /app/dist ./dist
ENV PORT=8080
EXPOSE 8080
CMD ["node", "dist/http.js"]

Two rules matter. Listen on the port the host gives you by reading PORT, and bind to 0.0.0.0, not localhost. Then run the server in stateless Streamable HTTP mode so any instance can answer any request. Hosts that scale to zero will drop in-memory sessions without warning.

Aerial view of colorful shipping containers stacked at a harbour in golden light

Render, Cloud Run and Spaces

  • Render Free: 750 instance hours per workspace each month, a single instance, no persistent disk, no SSH, and no outbound SMTP. Free Postgres databases expire after 30 days.
  • Google Cloud Run: the always-free tier gives 2 million requests, 180,000 vCPU-seconds, 360,000 GiB-seconds and 1 GB of North America egress each month. It scales to zero by default, but you must attach a billing account.
  • Hugging Face Spaces: add sdk: docker and app_port: 7860 to the YAML header of your README. The container runs as user 1000 and anything written to disk is lost on restart.

Cold starts break handshakes

Render's free service spins down after 15 minutes without traffic, and waking it takes about a minute. A client with a short timeout fails the first request, then works on the second. That bug is miserable to chase if you do not know the cause.

Finger pressing the power button on a small circuit board on a workbench

You have three options:

  1. Accept it and retry the first call
  2. Split the server: put latency-sensitive tools on Workers and heavy ones in the container
  3. Ping a health endpoint every 10 minutes. One service then stays awake inside the 750 hours (a 31-day month has 744), but check the host's rules before you rely on it

Lock It Down Before Sharing the URL

An MCP URL is a remote control for whatever your tools can do. Anyone who finds it can press the buttons, and free tiers have no budget cap that stops a stranger from burning your quota.

Add a bearer token

On Vercel, wrap the handler with withMcpAuth, set required: true, and return the verified client from verifyToken. Requests without a valid token get a 401. A token missing a required scope gets a 403. The docs example reads a demo token from an environment variable, and says production should check issuer, audience, expiration and scopes against a real authorization server. mcp-handler does not issue tokens itself.

To satisfy the MCP spec, also publish OAuth protected resource metadata at /.well-known/oauth-protected-resource, so compliant clients can find your authorization server.

Brass padlock on a steel chain with raindrops on a wooden gate

Keep secrets out of the repo

  • Vercel: project Environment Variables
  • Cloudflare: npx wrangler secret put PICASSOIA_TOKEN
  • Render and Spaces: dashboard secrets, read as environment variables at runtime

Never paste a token into a tool description or a prompt. Clients show tool descriptions to the model, and the model can repeat them. Add a per-token call counter too, since the free quota is shared by every caller.

Pick the Right Free Option

Top-down view of a notebook with hand-drawn grid lines beside a coffee cup

  • Pick Vercel if the project is already on Next.js, the traffic is personal, and you want a preview URL for every pull request.
  • Pick Cloudflare if you want near-zero wake-up, up to 100,000 calls a day, and tools that mostly forward requests to other APIs.
  • Pick a container host if you need native binaries, files on disk, or a job that runs longer than a quick request.

Nothing stops you from mixing them. A common setup is a Worker that answers cheap lookups and forwards heavy requests to a container, so the client always gets a fast first response. Start with the simplest platform that fits, measure for a week, and move only when a limit actually bites.

What breaks first

PlatformFirst wall you hitWarning sign
Vercel Hobby4 CPU hours, or the non-commercial rulePaused features, 30 day wait
Workers Free10 ms CPU on heavy toolsCPU limit errors on parsing
Render FreeWake-up delayFirst call fails after idle time
Cloud RunUsage beyond the free allowanceCharges on the invoice
SpacesDisk resetsState vanishes after a restart

Try It With Picasso IA

A hosted server is only useful if its tools do something worth calling. Image generation makes a good first tool, because you can see at once whether the whole loop works.

A tool that calls an image model

PicassoIA exposes a Replicate-style API at https://api.picassoia.com/v1. You authenticate with a Bearer token that starts with pia_sk_, create a prediction, then poll it. The PicassoIA Image model takes an input object with a prompt and an aspect_ratio:

server.registerTool(
  'make_image',
  {
    description: 'Start an image generation from a text prompt',
    inputSchema: z.object({ prompt: z.string().min(10) }),
  },
  async ({ prompt }) => {
    const res = await fetch(
      'https://api.picassoia.com/v1/models/picassoia/picassoia-image/predictions',
      {
        method: 'POST',
        headers: {
          Authorization: `Bearer ${process.env.PICASSOIA_TOKEN}`,
          'Content-Type': 'application/json',
        },
        body: JSON.stringify({ input: { prompt, aspect_ratio: '16:9' } }),
      },
    );
    if (!res.ok) {
      return { isError: true, content: [{ type: 'text', text: `API error ${res.status}` }] };
    }
    const prediction = await res.json();
    return {
      content: [{ type: 'text', text: `Prediction ${prediction.id} is ${prediction.status}` }],
    };
  },
);

On Workers, read the token from env instead of process.env. Predictions are asynchronous, so split the work into two tools: make_image returns the prediction ID, and check_image calls GET /v1/predictions/{id} and returns the status and output URLs. The statuses are starting, processing, succeeded, failed and canceled. Each call finishes in milliseconds, which fits inside the Workers 10 ms CPU budget and never holds a request open. The account limit is 5 predictions at once, shared across tokens and MCP connections.

💡 The API docs say predictions are currently free and use no credits, and they also say an Infinite plan is needed to create them. Check the pricing page before you build a public tool on top of it.

Models worth wiring in

Photographer reviewing printed photographs pinned to a wall in a bright studio

Deploy the dice server tonight, then open Picasso IA and make something with it. Write one prompt, run it through Seedream 5 Pro and Nano Banana 2 Lite, and compare the two results side by side. Once your own MCP tool returns an image URL on a server that costs nothing, you have the full loop working, and every new tool after that takes minutes. Try a few prompts, break a few limits on purpose, and see which free host holds up for your project.

Share this article