Large Language ModelsGenerate imagesGenerate videos

How to Build an MCP Server From Scratch in Python, Step by Step

A working Python MCP server, from an empty folder to Claude Desktop. Write tools, resources and prompts with the official SDK 2.x, test them in the Inspector, add a real image tool with async API calls, and deploy it over Streamable HTTP.

How to Build an MCP Server From Scratch in Python, Step by Step
Cristian Da Conceicao
Founder of Picasso IA

Most MCP tutorials stop at a weather tool that returns a hardcoded string. This one builds a server you can actually keep. You will write it from scratch with the official Python SDK, test it without any AI client, connect it to Claude Desktop and Claude Code, and finish with a real tool that calls an image API and waits for the result. Everything below needs Python 3.10 or newer and matches MCP Python SDK 2.x (2.3.0 on PyPI at the time of writing, October 2026).

If you copied code from a 2025 tutorial and hit ModuleNotFoundError: No module named 'mcp.server.fastmcp', you are in the right place. The main class was renamed, and the setup section shows the one-line fix.

What an MCP Server Actually Does

The Model Context Protocol (MCP) is a standard way for an AI application to call your code. Three roles matter. The host is the app the person talks to, such as Claude Desktop or an IDE. The client lives inside the host and speaks the protocol. The server is the part you build. Your server never talks to the model directly, it only answers requests from a client.

Three Primitives, Three Owners

A server exposes exactly three kinds of capability, and what separates them is who decides to use them:

PrimitiveWho triggers itWhat it isExample
ToolThe modelA function that takes an actionGenerate an image, write a database row
ResourceThe applicationData loaded into the model's contextA file, a config, a catalog
PromptThe userA reusable message templateA slash command

If you have built a web API, the mapping is quick. A resource behaves like a GET, a tool behaves like a POST, and a prompt is a saved query the user runs by name.

Hand-drawn notebook sketch of three connected blocks on an oak desk

💡 Rule of thumb: if the model should decide when to run it, make it a tool. If the app should attach it, make it a resource. If a person should pick it from a menu, make it a prompt.

Pick a Transport Early

The transport is how bytes move between client and server. You choose it with one argument to mcp.run().

TransportHow it worksUse it for
stdioThe host launches your file as a subprocess and talks over its stdin and stdoutLocal servers, and the default
streamable-httpA real HTTP server on a port, with the endpoint at /mcpAnything you deploy
sseThe older HTTP transportNothing new, it was superseded in the 2025-03-26 protocol revision

Network patch cables routed into a switch in a small server closet

Start with stdio. You will switch to Streamable HTTP near the end, and the tool code stays exactly the same.

Set Up Python in Five Minutes

Developer's hands typing in a code editor on a silver laptop

Install uv and the SDK

You need Python 3.10 or newer and uv. Create a project and add the SDK:

uv init mcp-image-studio
cd mcp-image-studio
uv add "mcp[cli]" httpx

The cli extra installs the mcp command with mcp dev, mcp run and mcp install. Plain pip install "mcp[cli]" works too. The Inspector is a Node.js app, so npx must be on your PATH.

A Rename That Breaks Old Code

In SDK 1.x the high-level class was FastMCP. In 2.x it is MCPServer, and it lives in a different module:

# SDK 1.x, seen in older tutorials
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Demo")

# SDK 2.x, used in this article
from mcp.server import MCPServer
mcp = MCPServer("Demo")

One more change trips people up: transport settings such as port moved off the constructor and onto run(). Passing port= to MCPServer(...) raises a TypeError.

Write Your First Server

Create server.py. One file, three decorators, and every primitive is registered:

from mcp.server import MCPServer

mcp = MCPServer("Demo")


@mcp.tool()
def add(a: int, b: int) -> int:
    """Add two numbers."""
    return a + b


@mcp.resource("greeting://{name}")
def greeting(name: str) -> str:
    """Greet someone by name."""
    return f"Hello, {name}!"


@mcp.prompt()
def summarize(text: str) -> str:
    """Summarize a piece of text in one sentence."""
    return f"Summarize the following text in one sentence:\n\n{text}"


if __name__ == "__main__":
    mcp.run()

That is a working server. The if __name__ guard matters: mcp dev, mcp run, mcp install and your tests all import this file, and an unguarded run() would start a server the moment anything loaded it.

Add a Tool

The SDK reads three things from your function. The name becomes the tool name, the docstring becomes the description the model sees, and the type hints become the argument schema. There is no JSON Schema to write, because a: int, b: int is the schema. If a client sends a string where you declared an integer, the SDK rejects the call before your function runs.

Give a parameter a default and it becomes optional. For tighter limits, wrap the type in Annotated with a Pydantic Field:

from typing import Annotated, Literal
from pydantic import Field


@mcp.tool()
def search_books(
    query: str,
    limit: Annotated[int, Field(ge=1, le=50, description="Maximum results")] = 10,
    genre: Literal["fiction", "non-fiction", "poetry"] = "fiction",
) -> str:
    """Search the catalog by title or author."""
    return f"Found 3 books matching {query!r} (up to {limit})."

The bounds land in the schema as minimum and maximum, and the Literal becomes an enum the model must pick from.

Add a Resource and a Prompt

A {param} in a resource URI turns it into a resource template, so greeting://{name} has no single entry to list until someone supplies a name. A prompt is simpler still: the string it returns becomes a user message. Both read their descriptions from the docstring, just like tools.

Raise Errors the Model Can Read

When a tool fails, raise ToolError. Never return an error string, because a returned string has is_error=False and looks like a successful answer.

from mcp.server.mcpserver.exceptions import ToolError

CATALOG = {"Dune": "Frank Herbert", "Neuromancer": "William Gibson"}


@mcp.tool()
def get_author(title: str) -> str:
    """Look up the author of a book in the catalog."""
    if title not in CATALOG:
        raise ToolError(f"No book titled {title!r} in the catalog.")
    return CATALOG[title]

The model reads that message, realizes it guessed the title wrong, and calls again with a better one. One raise gives you a self-correcting agent. Any other exception counts as a crash: the model only sees that the call failed, and your log gets the traceback.

💡 Declare a tool async def whenever it does I/O, such as an API call, a file read or a database query. Use plain def for everything else.

Test and Connect It

Run the MCP Inspector

Before any AI client touches your server, run it under the Inspector:

uv run mcp dev server.py

Open the URL it prints. The Inspector launches server.py as a subprocess over stdio, exactly like a real host would. Walk through the tabs in order:

  • Tools: add appears with a form built from your type hints. Call it with a=1 and b=2 and you get 3.
  • Resources: the list is empty, and greeting sits under Resource Templates. Give it World and you read Hello, World!.
  • Prompts: summarize has one required text argument and returns a single user message.

Two engineers reviewing code on a laptop at a standing desk

Write an In-Memory Test

The SDK's Client class also connects in memory: hand it the server object and there is no subprocess and no port. Add pytest with uv add --dev pytest, then create test_server.py:

import pytest
from mcp import Client

from server import mcp


@pytest.fixture
def anyio_backend():
    return "asyncio"


@pytest.mark.anyio
async def test_add():
    async with Client(mcp, raise_exceptions=True) as client:
        result = await client.call_tool("add", {"a": 1, "b": 2})
        assert result.structured_content == {"result": 3}

Keep raise_exceptions=True in tests only. It makes the real error message visible instead of the sanitized Internal server error a remote caller would see.

Connect Claude Desktop and Claude Code

Every host needs the same thing: the command that starts your server. This one works from any directory, with no virtual environment to activate:

uv run --with "mcp[cli]" mcp run /absolute/path/to/server.py

Claude Desktop is the one host the SDK can configure for you:

uv run mcp install server.py

That writes an entry into claude_desktop_config.json, found in ~/Library/Application Support/Claude/ on macOS and %APPDATA%\Claude\ on Windows. Quit Claude Desktop fully, not just the window, then reopen it. The app starts your server with its own environment, so pass secrets with -v NAME=value or -f .env.

Claude Code needs no file at all. Register the server with the CLI, then run /mcp inside a session to confirm it is connected:

claude mcp add image-studio -- uv run --with "mcp[cli]" mcp run /absolute/path/to/server.py

Cursor reads .cursor/mcp.json under the mcpServers field, and VS Code reads .vscode/mcp.json under servers with "type": "stdio". The command inside is identical.

Laptop showing a chat window beside a terminal on a wooden table

Fix a Server That Won't Show Up

First, run the launch command yourself. A healthy stdio server prints nothing and waits for a host to speak first. A traceback or an instant exit is your real bug. If it waits quietly, check these three causes:

SymptomCauseFix
Server never startsA relative path, because the host launches from its own working directoryUse absolute paths, including the one to uv (where uv on Windows, which uv elsewhere)
Edits have no effectHosts read their config at launchQuit the host fully and reopen it
Connection drops at onceSomething wrote to stdout, which is the protocol wireLog with the logging module, which writes to stderr, and never rely on print()

Claude Desktop keeps one log per server, named mcp-server-<NAME>.log, in ~/Library/Logs/Claude on macOS and %APPDATA%\Claude\logs on Windows. That file is your server's stderr.

Build a Real Image Tool

A server earns its place when a tool does work the model cannot do alone. This one takes a prompt, asks the PicassoIA API for an image, and returns the URL.

Design the Tool Contract

The API is Replicate-style: you create a prediction, poll it, then read the output. These are the facts the tool depends on:

DetailValue
Base URLhttps://api.picassoia.com/v1
AuthAuthorization: Bearer pia_sk_..., created on the API page
Create a jobPOST /v1/models/{owner}/{name}/predictions with {"input": {"prompt": "..."}}
Model used herePicassoIA Image, slug picassoia/picassoia-image
Prompt length1 to 4,000 characters
Status valuesstarting, processing, succeeded, failed, canceled
Concurrency5 predictions per account, shared across every token and MCP connection

The API docs state that an Infinite plan is required, and that a request without it returns 403 plan_required. Predictions are described as free and use no credits. Check your plan before you debug anything else.

Woman sketching an API flow on a whiteboard in a sunlit loft

Keep the contract small: one tool, two parameters, one URL back. Every failure path raises ToolError, so the model always gets a readable message.

Handle Async Jobs Without Blocking

Because the job runs on a remote GPU, the tool must wait without freezing the server. That means async def, httpx.AsyncClient and asyncio.sleep:

import asyncio
import os
from typing import Literal

import httpx
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError

API = "https://api.picassoia.com/v1"
MODEL = "picassoia/picassoia-image"
DONE = ("succeeded", "failed", "canceled")

mcp = MCPServer("Image Studio")
slots = asyncio.Semaphore(4)


@mcp.tool()
async def generate_image(
    prompt: str,
    aspect_ratio: Literal["1:1", "16:9", "9:16", "4:3"] = "16:9",
) -> str:
    """Generate an image from a text prompt and return its URL."""
    token = os.environ.get("PICASSOIA_API_TOKEN")
    if not token:
        raise ToolError("PICASSOIA_API_TOKEN is not set for this server.")

    headers = {"Authorization": f"Bearer {token}"}
    body = {"input": {"prompt": prompt, "aspect_ratio": aspect_ratio}}

    async with slots, httpx.AsyncClient(headers=headers, timeout=30) as http:
        response = await http.post(f"{API}/models/{MODEL}/predictions", json=body)
        prediction = response.json()
        if not response.is_success:
            raise ToolError(f"{prediction.get('code')}: {prediction.get('detail')}")

        while prediction["status"] not in DONE:
            eta = prediction.get("eta") or {}
            await asyncio.sleep(eta.get("next_poll_in_seconds", 2))
            prediction = (await http.get(prediction["urls"]["get"])).json()

    if prediction["status"] != "succeeded":
        raise ToolError(prediction.get("error") or prediction["status"])

    output = prediction["output"]
    return output[0] if isinstance(output, list) else output


if __name__ == "__main__":
    mcp.run()

Pen pointing at API response notes next to a laptop

Three details make this safe to leave running:

  1. Poll on the server's schedule. The response carries eta.next_poll_in_seconds, so you sleep exactly as long as the API asks.
  2. Cap your own concurrency. The Semaphore(4) keeps one chatty model from taking all 5 account slots.
  3. Read the token from the environment. Register it with mcp install server.py -v PICASSOIA_API_TOKEN=pia_sk_..., and never paste it into the file.

💡 The output field can be a list of URLs, a single URL, or null. The last two lines handle the first two, and the succeeded check above them rules out null in practice.

How to Use Sonnet 5 on PicassoIA

Writing tool code is where a language model saves the most time. Claude Sonnet 5 on PicassoIA handles multi-step coding and tool-use tasks, so you can paste the working generate_image tool and ask for the next one.

  1. Open the model page for Claude Sonnet 5.
  2. Write the prompt. Paste your server and ask: Add a second tool that lists my recent predictions with GET /v1/predictions. Reuse the same error handling. The prompt field is the only required one.
  3. Set a system prompt to stop the rename problem before it starts: You write Python for MCP SDK 2.x. Import MCPServer from mcp.server and never use FastMCP.
  4. Choose the effort. The default low skips extended thinking and is fastest. Use high or max for a bug that touches several files.
  5. Keep max tokens at 8192 for full server files, or lower it for a quick snippet.
  6. Attach a screenshot of an Inspector error if you have one. The model reads images, and max_image_resolution (default 0.5 megapixels) scales them down.
  7. Generate, copy, and test. Paste the result into server.py and run it through mcp dev before trusting it.
ParameterRequiredDefaultWhat it does
promptYesnoneYour request
system_promptNoemptyFixes role and constraints for the session
effortNolowThinking depth, from fastest to deepest
max_tokensNo8192Output length cap
imageNononeA screenshot or diagram as context

Prefer another model? Kimi K2.6 sits in the same category and is described for building agents and writing code.

Ship It Over HTTP

Technician walking through a server rack corridor

Switch the Transport

Change one line at the bottom of server.py:

if __name__ == "__main__":
    mcp.run(transport="streamable-http", port=3001)

Clients now connect to http://127.0.0.1:3001/mcp. You can also leave the file alone and run uv run mcp run server.py --transport streamable-http. The run() call accepts these options:

  • host and port, defaulting to 127.0.0.1 and 8000
  • streamable_http_path, defaulting to /mcp
  • json_response=True to answer each POST with one JSON body
  • stateless_http=True for a fresh transport per request

Register the remote server in Claude Code with claude mcp add --transport http image-studio https://mcp.example.com/mcp.

Lock It Down

Once your server leaves localhost, three things change:

  • The Host allowlist. The default accepts only 127.0.0.1, localhost and [::1]. Behind a real hostname, every request fails with 421 Misdirected Request and Invalid Host header. Fix it with transport_security= and list both "mcp.example.com" and "mcp.example.com:*" in allowed_hosts.
  • Authorization. Your server is an OAuth 2.1 resource server. Implement TokenVerifier with one async verify_token method that returns an access token or None, and pass token_verifier= together with auth=.
  • TLS behind a proxy. When a load balancer ends TLS, start uvicorn with --proxy-headers so it trusts the forwarded headers.

💡 A 421 is a plain HTTP response, not a protocol error, so the client only shows a generic transport failure. The offending hostname appears in the server's log. A freshly deployed server that refuses every connection is a Host allowlist problem until proven otherwise.

Build Your Own Images With Picasso IA

Developer leaning back in a chair after finishing work

You now have a server that registers tools, resources and prompts, passes an in-memory test, runs inside Claude, and can be deployed behind a real hostname. The part worth your next hour is the tool itself: swap the model slug, add an edit_image tool, or wire a video tool next to it.

Try the model your server just called. PicassoIA Image turns a prompt into a finished picture in seconds, and you can test any prompt in the browser before you automate it. When a still is not enough, PicassoIA Video and Seedance 2.5 Lite animate a prompt or a photo into short clips.

Open Picasso IA, pick a model, and run the same prompt you would give your tool. Then wire it into your server and let Claude do the clicking.

Share this article