Large Language ModelsGenerate imagesGenerate videos
Cloudflare MCP Hosting: Code Mode, Server Portals and Setup
Deploy a remote MCP server on Cloudflare Workers with McpAgent and OAuth, shrink tool context with the search and execute pattern of Code Mode, then group your servers behind an MCP Server Portal with Zero Trust Access policies, tool filtering and access logs.
Your Model Context Protocol (MCP) server runs fine on your laptop. Then a teammate asks for the URL, a second editor needs it on another machine, and someone from security asks who can call which tool. A local process cannot answer any of that. Cloudflare gives you three pieces for this moment: Workers to host the server, Code Mode to shrink what the model has to read, and MCP Server Portals to put every server behind one controlled door. Below you will find each piece in order, with the commands, config and traps that matter, so you can go from an empty folder to a governed setup without guessing.
Why Host MCP on Cloudflare
Local servers hit a ceiling
A stdio server is a child process of one client on one machine. That works for a weekend project. It stops working the moment a second person shows up: everyone installs their own copy, secrets sit in local config files, and nobody can see which tools get called. A remote server flips every one of those problems. You get one URL, one deploy and one place to read logs.
Local still wins in one case: a tool that touches files on a single person's machine, such as a private notes folder. Remote hosting is for tools that several people or several agents share.
What Workers bring to the table
Workers run your code on Cloudflare's edge network, close to whoever is calling. For MCP specifically, Cloudflare ships three building blocks:
McpAgent, a class in the Agents SDK that handles the remote transport. The SDK serves Streamable HTTP for you.
workers-oauth-provider, an OAuth 2.1 provider library that wraps your Worker and adds authorization to its endpoints, MCP endpoints included.
mcp-remote, an adapter that lets clients which only speak stdio connect to a remote server.
Need
Local stdio server
Remote server on Workers
Who can use it
One machine
Anyone with the URL and a login
Updating
Reinstall on every machine
One wrangler deploy
Secrets
Local config files
Worker secrets
Per-session state
Process memory
Durable Objects
Visibility
None built in
Portal access logs
💡 Remember: remote does not mean public. Treat the URL as an internet-facing API from the first deploy.
Deploy Your First Remote Server
Scaffold from the template
Cloudflare maintains a template for a server with no login, which is the fastest way to see the moving parts:
npm create cloudflare@latest -- my-mcp-server --template=cloudflare/ai/demos/remote-mcp-authless
cd my-mcp-server
npm start
Your server now listens locally at http://localhost:8788/mcp. Nothing else to install.
Write the McpAgent class
The heart of the project is a class that extends McpAgent. You register tools inside init(), exactly like you would with the official TypeScript SDK:
import { McpAgent } from "agents/mcp";
import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { z } from "zod";
export class MyMCP extends McpAgent {
server = new McpServer({ name: "math", version: "1.0.0" });
async init() {
this.server.tool("add", { a: z.number(), b: z.number() }, async ({ a, b }) => ({
content: [{ type: "text", text: String(a + b) }],
}));
}
}
export default MyMCP.serve("/mcp");
That class is also a Durable Object, which is why the project config declares a binding and a migration for it. Durable Objects give each MCP session its own state without a database on your side. If your tools only return pure results, you will never touch that state. The moment you track a cart, a draft or a long conversation, you will be glad it exists.
Test and ship it
Run the MCP Inspector in a second terminal and point it at the local URL:
npx @modelcontextprotocol/inspector@latest
Call your tool from its web interface. When it behaves, deploy:
npx wrangler@latest deploy
Your server is then live at https://my-mcp-server.<your-account>.workers.dev/mcp. Clients that only speak stdio, such as Claude Desktop, connect through the adapter:
You can also paste the URL into the Cloudflare AI Playground or the Inspector to test the deployed version.
Add Login With OAuth
An authless server is fine for a demo and a bad idea for anything that touches real data. Cloudflare's second template wires in GitHub as the identity provider:
Register two GitHub OAuth apps, one for local development and one for production, so a leaked dev secret never touches prod.
Store the credentials as Worker secrets with npx wrangler secret put GITHUB_CLIENT_ID and npx wrangler secret put GITHUB_CLIENT_SECRET, plus the cookie encryption secret named in the template's README.
Create the session store with npx wrangler kv namespace create "OAUTH_KV".
Paste the returned namespace ID into wrangler.jsonc, then deploy.
Under the hood, workers-oauth-provider wraps your Worker, so your tools receive already-authenticated user details as a parameter. You do not hand-roll token checks, and that is the whole point.
💡 Tip: GitHub is only one option. The same provider library can sit in front of any OAuth identity provider, which matters once you plan to put the server behind a portal.
How Code Mode Cuts Token Costs
Long tool lists burn context
Every tool definition you expose is text the model must read before it does anything useful. That is manageable with ten tools. It collapses with a whole platform. Cloudflare reports that exposing its API of more than 2,500 endpoints as ordinary MCP tools would take over 1.17 million tokens. With Code Mode, the same reach fits in roughly 1,000 tokens.
There is a second cost that gets less attention. In a normal tool-calling loop, every intermediate result travels back through the model. If step two needs the output of step one, the model reads it, restates it and sends it onward. Code Mode lets the model write a short program instead. Dependent calls run inside the sandbox, the intermediate data stays there, and only the final answer returns to the conversation. Fewer round trips means less text to read and fewer chances to copy a value wrong.
Search and execute in practice
The large-API pattern, openApiMcpServer(), exposes just two tools:
search runs model-written code against an OpenAPI document inside a sandbox and returns only the operations, parameters or schemas the task needs.
execute runs model-written code with an authenticated request function that your Worker supplies.
As the docs put it, only the returned subset enters the model's context. The model asks a narrow question, gets a narrow answer, then acts.
Picture a request like list the DNS records for my zone. The model first writes a small snippet for search that filters the OpenAPI paths down to the DNS operations, and gets back a handful of matches instead of thousands. Then it writes a snippet for execute that calls the right operation through your request function and returns only the fields it needs. Two short round trips replace a tool list the size of a phone book.
To build one, you need a Workers project, an OpenAPI 3.x document and a host-side method for authenticating requests.
The sandbox keeps code contained
Model-written code runs in an isolated Worker, and direct outbound network access is blocked by default. The generated code can only reach the outside world through upstream MCP tools or the request callback you provide. That is a strong default, but it does not do your authorization for you:
Enforce permissions inside your tool handlers or request callback before any side effect happens.
Never place credentials in tool results or in the OpenAPI document.
Treat the callback as the one place where a bad request can actually do damage.
Pick the right pattern
codeMcpServer()
openApiMcpServer()
Best for
Wrapping an existing MCP server with a manageable tool set
Large API catalogs
What the model sees
One code tool with TypeScript definitions of every upstream operation
Two tools: search and execute
How calls happen
Through a codemode namespace, so dependent calls compose inside the sandbox
Selected operations called through a host-provided request function
Context cost
Grows with the number of upstream tools
Bounded, because only results of search come back
💡 Rule of thumb: wrap what you already have with codeMcpServer(). Reach for openApiMcpServer() when your tool list would be a catalog, not a toolbox.
Set Up an MCP Server Portal
MCP Server Portals launched in open beta in August 2025 as part of Cloudflare One. The idea is simple: route every MCP request through one portal endpoint, apply Zero Trust policies there, and log everything.
Check the prerequisites first
Before you open the dashboard, confirm three things:
You have an active Cloudflare domain, with a full or partial (CNAME) setup.
An identity provider is configured in Cloudflare Zero Trust.
Your servers are reachable over HTTP. Stdio-only servers are not supported unless you wrap them. A portal holds up to 80 servers.
Add servers, then build the portal
In the dashboard, go to Zero Trust > Access controls > MCP Portals and open the MCP servers tab.
Select Add MCP server. Enter a name, an optional custom Server ID, the full HTTP URL of the server, and the Access policies that decide who sees it.
For OAuth-enabled servers, use automatic Dynamic Client Registration (recommended) or enter credentials manually. Add the dashboard's callback URL to your OAuth provider's allowlist.
Back on the MCP Portals page, select Add MCP server portal. Set a name, a custom domain with an optional subdomain, the servers to attach and the access policies for users.
Connect clients to https://<subdomain>.<domain>/mcp.
A server only appears in the portal for people who match an Allow policy. Menu labels can shift while a feature is in beta, so trust the current dashboard over any screenshot.
Trim tools and set auth
Inside the portal settings you can switch off the toggle next to any tool or prompt you want to hide. Each server shows a Tools authorized count so you can see how much you have exposed. A few controls worth knowing:
Require user auth decides whether people sign in with their own credentials or the admin credential handles access.
Namespacing shows tools as {server_id}_{tool_name}, so two servers can both have a search tool without a collision.
Aliases rename tools and prompts at the portal or server level.
Code Mode can be switched on for the portal to reduce token use.
Gateway routing can add optional DLP inspection for sensitive data.
Read the access logs
Portal logs record time, status, server name, capability and duration, per portal or per server. They can be exported with Logpush to third-party storage or a SIEM. For a team rollout, this is where you answer the question security asked in the first paragraph: who called what, and when.
A rollout order that keeps surprises small:
Deploy one server with OAuth and test it in the Inspector.
Add it to a portal with an Allow policy for a pilot group only.
Switch off any tool the pilot group does not need.
Check the logs after a few days for unexpected callers or failing calls.
Widen the policy, then attach the next server.
Mistakes That Cost Hours
Sharing the workers.dev URL with no login. It works, which is exactly the danger. Add OAuth before anyone outside your machine sees the address.
An empty Allow policy. Users who sign in to the portal and see "No allowed servers available, check your Zero Trust Policies" almost always lack a matching Allow policy on the portal or the server.
Forgetting the callback URL. OAuth-enabled servers fail to connect until the dashboard's callback URL is on your provider's allowlist.
Putting credentials where the model can read them. With Code Mode, anything in a tool result or the OpenAPI document is visible to model-written code.
Expecting stdio inside a portal. Wrap the server behind HTTP first, or host it on Workers.
Skipping the Inspector. A tool that works in your editor can still fail on the deployed URL. Test the live /mcp endpoint before you add it to a portal.
💡 Quick test: open the portal as a user who is not in your Allow policy. If you see any server, your policy is wrong.
Pair It With PicassoIA Models
Choose a model for the client
Whatever client calls your server needs a capable model behind it. These PicassoIA language models are all worth testing against your tools:
They are also handy before you deploy anything: ask one to draft tool descriptions, write the TypeScript you would feed to Code Mode, or review your OpenAPI document for operations you would rather not expose.
Add image tools to your server
A Worker can call any HTTP API, so an MCP tool can call PicassoIA's. The PicassoIA API lives at https://api.picassoia.com/v1, takes a bearer token that starts with pia_sk_, and follows a Replicate-style pattern: POST /v1/models/{owner}/{name}/predictions creates a job and GET /v1/predictions/{id} reads its status. Jobs are asynchronous, and an account runs up to 5 predictions at once.
That shape maps neatly onto two tools: one that starts a generation and returns an id, and one that polls for the result. Store the token with npx wrangler secret put PICASSOIA_API_TOKEN, and never print it in a tool result. If you would rather not build anything, PicassoIA also offers its own MCP connection, which gives your client the same image and video models directly.
Try It on PicassoIA Today
You now have the full path: a Worker that serves MCP, OAuth in front of it, Code Mode to keep the context small, and a portal to govern all of it. The fastest reward is to make the server do something visual.
Open Seedream 5 Pro, GPT Image 2 or FLUX 2 Pro and write a prompt for the photo you wish your last project had. Then push it further with Seedance 2.0 or Veo 3.1 Fast and turn the still into motion. Experiment with lighting, lens and angle until the result looks like a real shoot.