Large Language ModelsGenerate imagesGenerate videos

Convert API to MCP Server: REST and OpenAPI Step by Step

Wrap an existing REST API as an MCP server that agents can call without guessing. Generate tools from an OpenAPI file with FastMCP, hand-build them in TypeScript, handle auth and slow image or video jobs, then test with the Inspector and ship over stdio or HTTP.

Convert API to MCP Server: REST and OpenAPI Step by Step
Cristian Da Conceicao
Founder of Picasso IA

Your REST API already works, and agents still fumble it. They guess at parameter names, choke on 40 KB JSON responses and call DELETE when they meant GET. The fix is not a smarter model. It is a thin layer in between: an MCP server that tells the agent exactly which actions exist, what each input looks like and what comes back. This tutorial shows how to convert an API to an MCP server from its OpenAPI spec, first with generated code, then by hand, including auth, slow jobs, testing and deployment.

You need three things before you begin: an API you can already call with curl, its OpenAPI 3.x file (or the patience to write one), and an MCP client such as Claude Desktop, Cursor or VS Code to try the result. A first working version takes an afternoon. The polish is where the real time goes, and it is also where the quality comes from.

💡 Short version: MCP wraps your API in tools. Each tool has a name, a description and a JSON Schema for its input. The agent chooses tools by reading those descriptions, so the descriptions matter more than the HTTP plumbing.

Why Wrap an API as MCP

REST was designed for developers who read documentation once and write code against it. An agent works differently. It reads whatever the server lists at the start of a session, then decides from that list alone which call to make. If the list is vague, it guesses. If the list is huge, it burns its context window before the user has typed a word.

Operator connecting patch cords on a vintage telephone switchboard

An MCP server solves both problems the way a switchboard operator does: it receives a clear request, routes it to the right line and returns a clean answer.

What the Agent Actually Sees

When a client connects, it asks the server for its tool list. Every entry carries a name, a description, an inputSchema written in JSON Schema and, optionally, an outputSchema and a set of annotations. That is the entire surface. The agent never sees your routes, your HTTP verbs or your status codes. It sees names, sentences and schemas.

REST to MCP at a Glance

Every part of an OpenAPI operation has a home on the MCP side:

REST / OpenAPIMCP tool
operationIdTool name
summary and descriptionTool description
Path, query and body parametersinputSchema, one flat JSON Schema object
200 response schemaoutputSchema and structured content
4xx and 5xx responsesResult with isError: true and a readable message
Security schemeServer config: environment token, or OAuth for remote servers
Pagination linksExplicit cursor and limit inputs

Structured output and output schemas arrived with the 2025-06-18 revision of the spec, so check that your SDK version supports them before you rely on outputSchema.

Map OpenAPI Operations to Tools

Open the spec and resist the urge to expose everything. A 120-endpoint API becomes a 120-tool server, and the tool list alone can eat thousands of tokens in every conversation. Start small, name things well and describe them the way a colleague would.

Highlighted printed API reference pages with a pencil circling one line

Pick Operations, Not Endpoints

Ask four questions of every endpoint before it becomes a tool:

  • Would a person ask an assistant to do this in plain language?
  • Is it safe to call twice if the agent retries?
  • Does the response fit in a few kilobytes, or can you trim it until it does?
  • Does it belong to another audience, such as admin, billing or internal tooling?

Anything that fails the first or last question stays out. Five to ten well-chosen tools beat a hundred raw ones. Multi-step flows deserve one tool: if "create a cart, add items, check out" always happens in sequence, the agent should see a single place_order action.

Name and Describe Each Tool

Start from operationId, then rewrite it as a verb and a noun. A good description answers three questions: what the tool does, when to use it instead of its neighbors, and what it returns.

GeneratedRewritten
NameOrdersController_findAllsearch_orders
Description"Find all""Search orders by customer email, status or date range. Returns up to 20 orders with id, status and total. Use get_order for line items."
Parameterqemail: "Customer email, for example ana@example.com"

Turn Parameters Into Schemas

Flatten path, query and body parameters into a single object. Keep enums, mark required fields, give every property a short description with an example value, and set limits such as maximum and maxLength so the model cannot ask for 10,000 rows. An OpenAPI operation like this:

/orders/{orderId}:
  get:
    operationId: getOrder
    summary: Fetch one order
    parameters:
      - name: orderId
        in: path
        required: true
        schema: { type: string }

becomes this tool definition:

{
  "name": "get_order",
  "description": "Fetch one order by id. Returns status, total and line items. Use search_orders when you only have an email.",
  "inputSchema": {
    "type": "object",
    "properties": {
      "order_id": { "type": "string", "description": "Order id, for example ord_8f2c1" }
    },
    "required": ["order_id"]
  }
}

Two Ways to Build the Server

You can generate a server straight from the spec in minutes, or write every tool by hand. Most teams do both: generate first to see the shape, then hand-tune the five tools that matter.

Developer typing code at a cafe table beside a rainy window

Generate With FastMCP

The Python library FastMCP can build a server directly from an OpenAPI document:

import os
import httpx
from fastmcp import FastMCP
from fastmcp.server.providers.openapi import RouteMap, MCPType

client = httpx.AsyncClient(
    base_url="https://api.example.com",
    headers={"Authorization": f"Bearer {os.environ['ORDERS_API_TOKEN']}"},
)
spec = httpx.get("https://api.example.com/openapi.json").json()

mcp = FastMCP.from_openapi(
    openapi_spec=spec,
    client=client,
    name="Orders API",
    route_maps=[
        RouteMap(pattern=r"^/admin/.*", mcp_type=MCPType.EXCLUDE),
        RouteMap(tags={"internal"}, mcp_type=MCPType.EXCLUDE),
    ],
)

if __name__ == "__main__":
    mcp.run()

The server reads the spec, creates the tools and forwards each call through the httpx client you pass in, which is also where the auth header lives. The route maps drop admin and internal routes before the agent ever sees them.

⚠️ Version check: the default mapping differs between FastMCP major versions. Recent releases turn every operation into a tool, while older 2.x releases mapped some GET routes to resources. Pin your version, set route maps explicitly and confirm the import path in the docs for the release you installed.

When Generation Falls Short

FastMCP's own documentation warns that curated servers give models noticeably better results than auto-converted ones, especially for APIs with many endpoints and parameters. You will see why in the first test run:

  • Names such as get_orders_by_id_using_get that no human would write
  • Descriptions copied from developer docs, written for readers who already know the system
  • Responses that return every field, internal flags included
  • Four tools that should have been one

Fix them in this order: prune, rename, rewrite descriptions, trim responses, merge flows.

Hand-Build in TypeScript

For the tools that matter, the official TypeScript SDK gives you full control. Install @modelcontextprotocol/sdk and zod, then register each tool with a schema and a handler:

import { McpServer } from "@modelcontextprotocol/sdk/server/mcp.js";
import { StdioServerTransport } from "@modelcontextprotocol/sdk/server/stdio.js";
import { z } from "zod";

const API = "https://api.example.com";
const TOKEN = process.env.ORDERS_API_TOKEN;

const server = new McpServer({ name: "orders", version: "1.0.0" });

server.registerTool(
  "get_order",
  {
    title: "Get order",
    description:
      "Fetch one order by id. Returns status, total and line items. Use search_orders when you only have an email.",
    inputSchema: { order_id: z.string().describe("Order id, for example ord_8f2c1") },
    annotations: { readOnlyHint: true },
  },
  async ({ order_id }) => {
    const res = await fetch(`${API}/orders/${encodeURIComponent(order_id)}`, {
      headers: { Authorization: `Bearer ${TOKEN}` },
    });
    if (!res.ok) {
      return {
        isError: true,
        content: [{ type: "text", text: `Orders API returned ${res.status}. Check the id and try again.` }],
      };
    }
    const order = await res.json();
    return { content: [{ type: "text", text: JSON.stringify(order) }] };
  }
);

await server.connect(new StdioServerTransport());

Two details do the heavy lifting. The readOnlyHint annotation tells the client this call is safe to run without a confirmation prompt, and the error branch returns isError: true with a readable message, so the agent can try again instead of stalling.

Auth, Secrets and Guardrails

Brass padlock beside a hardware security token on a wooden ledge

The server holds the credentials. The model never does.

Keep Tokens Out of Prompts

Read the API token from an environment variable or a secret manager when the process starts. Never accept it as a tool argument, never echo it in an error message and never log request headers. Create the narrowest token the API allows: a read-only token for a read-only server. For remote servers, MCP's authorization flow is built on OAuth 2.1, so each user signs in with their own account and every call carries their own permissions instead of a shared superuser.

Annotate Risky Tools

Annotations are hints that help clients decide when to ask the user for confirmation:

AnnotationSet it when
readOnlyHint: trueThe tool only reads, such as a GET or a search
destructiveHint: trueThe tool deletes or overwrites data
idempotentHint: trueRepeating the call with the same input changes nothing more
openWorldHint: trueThe tool reaches systems outside your own, such as the open web

Treat them as hints, not enforcement, because a client should not trust annotations from a server it does not know. The real protection is on your side: ship the first release read-only, add write tools one at a time and give destructive ones a dry_run or confirm input so the agent has to be explicit.

Slow Jobs, Polling and Media

Chef sliding a plated dish onto a kitchen pass beside a rail of order tickets

Image generation, video rendering and report exports share a pattern: the API answers immediately with a job id, and the result arrives seconds or minutes later. A tool that blocks for three minutes will time out in most clients. A kitchen rail of order tickets solves the same problem in a restaurant: take the order, hand over a ticket, call the number when the dish is ready.

Create, Poll, Fetch

Split the job into three tools: one starts it, one checks it and one cancels it. The start tool returns an id and a hint for when to check back. The check tool returns a small status object, queued, running, succeeded or failed, plus a URL once there is something to fetch. Return links, not file bytes: a 5 MB image pasted into the context helps nobody.

server.registerTool(
  "get_render",
  {
    description:
      "Check a render started with start_render. Call again after next_poll_in_seconds until status is succeeded or failed.",
    inputSchema: { render_id: z.string() },
    annotations: { readOnlyHint: true },
  },
  async ({ render_id }) => {
    const job = await api(`/renders/${render_id}`); // api() is your fetch helper
    const done = job.status === "succeeded" || job.status === "failed";
    const body = {
      status: job.status,
      url: job.output?.[0] ?? null,
      next_poll_in_seconds: done ? null : 5,
    };
    return { content: [{ type: "text", text: JSON.stringify(body) }] };
  }
);

A Real Image and Video Example

PicassoIA's own connector follows this design. Its tools generate_image, edit_image, generate_video_picassoia and generate_video_seedance return a predict_id as soon as a GPU accepts the job, together with an estimated time. The agent then calls get_generation after the returned next_poll_in_seconds, and repeats until the status is succeeded or failed. A cancel_generation tool stops a running job, with one honest caveat in its instructions: a video the GPU is already rendering can no longer be canceled.

Underneath sits a Replicate-style REST API at https://api.picassoia.com/v1 with bearer-token auth. POST /v1/models/{owner}/{name}/predictions creates a job, GET /v1/predictions/{id} reads it and POST /v1/predictions/{id}/cancel stops it. That makes it a textbook conversion target, and the four models behind the connector line up with its four generation tools:

ModelWhat it returnsConnector tool
PicassoIA ImageText-to-image, seven aspect ratiosgenerate_image
PicassoIA Image Editor ProEdits with up to three reference imagesedit_image
Picasso IA Video5-second clips at 24 fps with audio, 480p or 720pgenerate_video_picassoia
Seedance 2.5 LiteClips of 5 or 10 seconds with audiogenerate_video_seedance

Limits belong in the tool descriptions too. The API allows 5 concurrent predictions per account, shared across tokens and MCP connections, with prompts up to 4,000 characters, so a good description tells the agent to wait for a running job before it launches a sixth. Check the current plan terms on the PicassoIA site before you build a product on top of the API.

Test, Then Ship

Two engineers reviewing a laptop screen at a wooden table

An agent is an unforgiving tester: it uses your tools in ways you did not plan. Test with proper tooling before you hand it over.

Run the MCP Inspector

The MCP Inspector is the official debugging interface. Point it at your server, for example npx @modelcontextprotocol/inspector node dist/server.js, and it lists every tool, lets you call each one with raw JSON and shows the exact result a client would receive. Then connect a real client and run ten realistic prompts. For each one, check three things: did the agent pick the right tool, did it fill the arguments correctly, and did the response give it enough to answer?

Choose stdio or HTTP

stdioStreamable HTTP
Runs asLocal process started by the clientRemote service behind a URL
AuthEnvironment variables on the user's machineOAuth or bearer tokens
Best forPersonal tools and developmentTeams and shared APIs
Watch outNever print logs to stdoutTLS, rate limits, horizontal scaling

Low-angle view down an aisle of server racks in a data hall

Streamable HTTP replaced the older HTTP plus SSE transport in the 2025-03-26 revision of the spec, and the protocol keeps moving, so pin your SDK version and read its changelog before upgrading.

Five Common Mistakes

  1. Exposing every endpoint. Tool lists cost tokens on every single turn.
  2. Logging to stdout on a stdio server. Stdout carries the protocol itself. Send logs to stderr.
  3. Returning the whole upstream payload. Trim to the fields the agent needs and paginate the rest.
  4. Throwing exceptions. Return an isError result with a message that says what to try next.
  5. Overlapping descriptions. If two tools sound alike, the agent flips a coin. Say when to prefer each one.

Use Claude Sonnet 5 on PicassoIA

Writing twenty tool descriptions by hand is tedious, and an LLM does the job well when you give it rules. On PicassoIA, Claude Sonnet 5 is a good fit: its model page lists multi-step coding and tool-use tasks among its strengths, it accepts a system prompt and it lets you choose how much thinking it does.

Open notebook with a pencil sketch of connected boxes beside a coffee mug

  1. Open the model. Go to the Claude Sonnet 5 page on PicassoIA.
  2. Set the system prompt once. For example: You write MCP tool definitions. For each OpenAPI operation return a verb_noun name, a description that says what the tool does, when to use it and what it returns, and a flat JSON Schema with example values. Never copy internal parameter names.
  3. Paste one operation at a time into the prompt field, or a small group of related ones. A whole 5 MB spec produces muddy output.
  4. Pick the effort level. low is the fastest and turns thinking off, medium suits a batch of simple operations, and high pays off for nested request bodies.
  5. Leave max tokens at 8192, the default, for batches of five to eight operations.
  6. Attach a screenshot if all you have is rendered docs. The image field accepts one.
  7. Review before shipping. Run every draft through the Inspector and fix names that overlap.

💡 Need output that must parse as JSON every time? GPT 5 Structured is built to return clean JSON, which suits schema drafts you plan to load straight into code.

Build Your Own Agent Toolkit

Pick one API, five tools and a free afternoon. Ship the read-only version first, test it with ten real prompts and only then add the tools that write or delete.

Hands resting on a laptop at a window desk at dusk

Agents need things to show, not just data to read. Try it yourself on Picasso IA: write a prompt in PicassoIA Image, pick 16:9, then refine the result with PicassoIA Image Editor Pro, which accepts up to three reference images. When the still looks right, animate it with Picasso IA Video or stretch the shot to ten seconds with Seedance 2.5 Lite. Experiment with your own scenes, and when you are ready, build the MCP tool that lets your agent do the same.

Share this article