Large Language ModelsGenerate imagesGenerate videos
How to Build an MCP Server From Scratch in Python, Step by Step
A working Python MCP server, from an empty folder to Claude Desktop. Write tools, resources and prompts with the official SDK 2.x, test them in the Inspector, add a real image tool with async API calls, and deploy it over Streamable HTTP.
Most MCP tutorials stop at a weather tool that returns a hardcoded string. This one builds a server you can actually keep. You will write it from scratch with the official Python SDK, test it without any AI client, connect it to Claude Desktop and Claude Code, and finish with a real tool that calls an image API and waits for the result. Everything below needs Python 3.10 or newer and matches MCP Python SDK 2.x (2.3.0 on PyPI at the time of writing, October 2026).
If you copied code from a 2025 tutorial and hit ModuleNotFoundError: No module named 'mcp.server.fastmcp', you are in the right place. The main class was renamed, and the setup section shows the one-line fix.
What an MCP Server Actually Does
The Model Context Protocol (MCP) is a standard way for an AI application to call your code. Three roles matter. The host is the app the person talks to, such as Claude Desktop or an IDE. The client lives inside the host and speaks the protocol. The server is the part you build. Your server never talks to the model directly, it only answers requests from a client.
Three Primitives, Three Owners
A server exposes exactly three kinds of capability, and what separates them is who decides to use them:
Primitive
Who triggers it
What it is
Example
Tool
The model
A function that takes an action
Generate an image, write a database row
Resource
The application
Data loaded into the model's context
A file, a config, a catalog
Prompt
The user
A reusable message template
A slash command
If you have built a web API, the mapping is quick. A resource behaves like a GET, a tool behaves like a POST, and a prompt is a saved query the user runs by name.
💡 Rule of thumb: if the model should decide when to run it, make it a tool. If the app should attach it, make it a resource. If a person should pick it from a menu, make it a prompt.
Pick a Transport Early
The transport is how bytes move between client and server. You choose it with one argument to mcp.run().
Transport
How it works
Use it for
stdio
The host launches your file as a subprocess and talks over its stdin and stdout
Local servers, and the default
streamable-http
A real HTTP server on a port, with the endpoint at /mcp
Anything you deploy
sse
The older HTTP transport
Nothing new, it was superseded in the 2025-03-26 protocol revision
Start with stdio. You will switch to Streamable HTTP near the end, and the tool code stays exactly the same.
Set Up Python in Five Minutes
Install uv and the SDK
You need Python 3.10 or newer and uv. Create a project and add the SDK:
uv init mcp-image-studio
cd mcp-image-studio
uv add "mcp[cli]" httpx
The cli extra installs the mcp command with mcp dev, mcp run and mcp install. Plain pip install "mcp[cli]" works too. The Inspector is a Node.js app, so npx must be on your PATH.
A Rename That Breaks Old Code
In SDK 1.x the high-level class was FastMCP. In 2.x it is MCPServer, and it lives in a different module:
# SDK 1.x, seen in older tutorials
from mcp.server.fastmcp import FastMCP
mcp = FastMCP("Demo")
# SDK 2.x, used in this article
from mcp.server import MCPServer
mcp = MCPServer("Demo")
One more change trips people up: transport settings such as port moved off the constructor and onto run(). Passing port= to MCPServer(...) raises a TypeError.
Write Your First Server
Create server.py. One file, three decorators, and every primitive is registered:
from mcp.server import MCPServer
mcp = MCPServer("Demo")
@mcp.tool()
def add(a: int, b: int) -> int:
"""Add two numbers."""
return a + b
@mcp.resource("greeting://{name}")
def greeting(name: str) -> str:
"""Greet someone by name."""
return f"Hello, {name}!"
@mcp.prompt()
def summarize(text: str) -> str:
"""Summarize a piece of text in one sentence."""
return f"Summarize the following text in one sentence:\n\n{text}"
if __name__ == "__main__":
mcp.run()
That is a working server. The if __name__ guard matters: mcp dev, mcp run, mcp install and your tests all import this file, and an unguarded run() would start a server the moment anything loaded it.
Add a Tool
The SDK reads three things from your function. The name becomes the tool name, the docstring becomes the description the model sees, and the type hints become the argument schema. There is no JSON Schema to write, because a: int, b: intis the schema. If a client sends a string where you declared an integer, the SDK rejects the call before your function runs.
Give a parameter a default and it becomes optional. For tighter limits, wrap the type in Annotated with a Pydantic Field:
from typing import Annotated, Literal
from pydantic import Field
@mcp.tool()
def search_books(
query: str,
limit: Annotated[int, Field(ge=1, le=50, description="Maximum results")] = 10,
genre: Literal["fiction", "non-fiction", "poetry"] = "fiction",
) -> str:
"""Search the catalog by title or author."""
return f"Found 3 books matching {query!r} (up to {limit})."
The bounds land in the schema as minimum and maximum, and the Literal becomes an enum the model must pick from.
Add a Resource and a Prompt
A {param} in a resource URI turns it into a resource template, so greeting://{name} has no single entry to list until someone supplies a name. A prompt is simpler still: the string it returns becomes a user message. Both read their descriptions from the docstring, just like tools.
Raise Errors the Model Can Read
When a tool fails, raiseToolError. Never return an error string, because a returned string has is_error=False and looks like a successful answer.
from mcp.server.mcpserver.exceptions import ToolError
CATALOG = {"Dune": "Frank Herbert", "Neuromancer": "William Gibson"}
@mcp.tool()
def get_author(title: str) -> str:
"""Look up the author of a book in the catalog."""
if title not in CATALOG:
raise ToolError(f"No book titled {title!r} in the catalog.")
return CATALOG[title]
The model reads that message, realizes it guessed the title wrong, and calls again with a better one. One raise gives you a self-correcting agent. Any other exception counts as a crash: the model only sees that the call failed, and your log gets the traceback.
💡 Declare a tool async def whenever it does I/O, such as an API call, a file read or a database query. Use plain def for everything else.
Test and Connect It
Run the MCP Inspector
Before any AI client touches your server, run it under the Inspector:
uv run mcp dev server.py
Open the URL it prints. The Inspector launches server.py as a subprocess over stdio, exactly like a real host would. Walk through the tabs in order:
Tools:add appears with a form built from your type hints. Call it with a=1 and b=2 and you get 3.
Resources: the list is empty, and greeting sits under Resource Templates. Give it World and you read Hello, World!.
Prompts:summarize has one required text argument and returns a single user message.
Write an In-Memory Test
The SDK's Client class also connects in memory: hand it the server object and there is no subprocess and no port. Add pytest with uv add --dev pytest, then create test_server.py:
import pytest
from mcp import Client
from server import mcp
@pytest.fixture
def anyio_backend():
return "asyncio"
@pytest.mark.anyio
async def test_add():
async with Client(mcp, raise_exceptions=True) as client:
result = await client.call_tool("add", {"a": 1, "b": 2})
assert result.structured_content == {"result": 3}
Keep raise_exceptions=True in tests only. It makes the real error message visible instead of the sanitized Internal server error a remote caller would see.
Connect Claude Desktop and Claude Code
Every host needs the same thing: the command that starts your server. This one works from any directory, with no virtual environment to activate:
uv run --with "mcp[cli]" mcp run /absolute/path/to/server.py
Claude Desktop is the one host the SDK can configure for you:
uv run mcp install server.py
That writes an entry into claude_desktop_config.json, found in ~/Library/Application Support/Claude/ on macOS and %APPDATA%\Claude\ on Windows. Quit Claude Desktop fully, not just the window, then reopen it. The app starts your server with its own environment, so pass secrets with -v NAME=value or -f .env.
Claude Code needs no file at all. Register the server with the CLI, then run /mcp inside a session to confirm it is connected:
claude mcp add image-studio -- uv run --with "mcp[cli]" mcp run /absolute/path/to/server.py
Cursor reads .cursor/mcp.json under the mcpServers field, and VS Code reads .vscode/mcp.json under servers with "type": "stdio". The command inside is identical.
Fix a Server That Won't Show Up
First, run the launch command yourself. A healthy stdio server prints nothing and waits for a host to speak first. A traceback or an instant exit is your real bug. If it waits quietly, check these three causes:
Symptom
Cause
Fix
Server never starts
A relative path, because the host launches from its own working directory
Use absolute paths, including the one to uv (where uv on Windows, which uv elsewhere)
Edits have no effect
Hosts read their config at launch
Quit the host fully and reopen it
Connection drops at once
Something wrote to stdout, which is the protocol wire
Log with the logging module, which writes to stderr, and never rely on print()
Claude Desktop keeps one log per server, named mcp-server-<NAME>.log, in ~/Library/Logs/Claude on macOS and %APPDATA%\Claude\logs on Windows. That file is your server's stderr.
Build a Real Image Tool
A server earns its place when a tool does work the model cannot do alone. This one takes a prompt, asks the PicassoIA API for an image, and returns the URL.
Design the Tool Contract
The API is Replicate-style: you create a prediction, poll it, then read the output. These are the facts the tool depends on:
Detail
Value
Base URL
https://api.picassoia.com/v1
Auth
Authorization: Bearer pia_sk_..., created on the API page
Create a job
POST /v1/models/{owner}/{name}/predictions with {"input": {"prompt": "..."}}
5 predictions per account, shared across every token and MCP connection
The API docs state that an Infinite plan is required, and that a request without it returns 403 plan_required. Predictions are described as free and use no credits. Check your plan before you debug anything else.
Keep the contract small: one tool, two parameters, one URL back. Every failure path raises ToolError, so the model always gets a readable message.
Handle Async Jobs Without Blocking
Because the job runs on a remote GPU, the tool must wait without freezing the server. That means async def, httpx.AsyncClient and asyncio.sleep:
import asyncio
import os
from typing import Literal
import httpx
from mcp.server import MCPServer
from mcp.server.mcpserver.exceptions import ToolError
API = "https://api.picassoia.com/v1"
MODEL = "picassoia/picassoia-image"
DONE = ("succeeded", "failed", "canceled")
mcp = MCPServer("Image Studio")
slots = asyncio.Semaphore(4)
@mcp.tool()
async def generate_image(
prompt: str,
aspect_ratio: Literal["1:1", "16:9", "9:16", "4:3"] = "16:9",
) -> str:
"""Generate an image from a text prompt and return its URL."""
token = os.environ.get("PICASSOIA_API_TOKEN")
if not token:
raise ToolError("PICASSOIA_API_TOKEN is not set for this server.")
headers = {"Authorization": f"Bearer {token}"}
body = {"input": {"prompt": prompt, "aspect_ratio": aspect_ratio}}
async with slots, httpx.AsyncClient(headers=headers, timeout=30) as http:
response = await http.post(f"{API}/models/{MODEL}/predictions", json=body)
prediction = response.json()
if not response.is_success:
raise ToolError(f"{prediction.get('code')}: {prediction.get('detail')}")
while prediction["status"] not in DONE:
eta = prediction.get("eta") or {}
await asyncio.sleep(eta.get("next_poll_in_seconds", 2))
prediction = (await http.get(prediction["urls"]["get"])).json()
if prediction["status"] != "succeeded":
raise ToolError(prediction.get("error") or prediction["status"])
output = prediction["output"]
return output[0] if isinstance(output, list) else output
if __name__ == "__main__":
mcp.run()
Three details make this safe to leave running:
Poll on the server's schedule. The response carries eta.next_poll_in_seconds, so you sleep exactly as long as the API asks.
Cap your own concurrency. The Semaphore(4) keeps one chatty model from taking all 5 account slots.
Read the token from the environment. Register it with mcp install server.py -v PICASSOIA_API_TOKEN=pia_sk_..., and never paste it into the file.
💡 The output field can be a list of URLs, a single URL, or null. The last two lines handle the first two, and the succeeded check above them rules out null in practice.
Writing tool code is where a language model saves the most time. Claude Sonnet 5 on PicassoIA handles multi-step coding and tool-use tasks, so you can paste the working generate_image tool and ask for the next one.
Write the prompt. Paste your server and ask: Add a second tool that lists my recent predictions with GET /v1/predictions. Reuse the same error handling. The prompt field is the only required one.
Set a system prompt to stop the rename problem before it starts: You write Python for MCP SDK 2.x. Import MCPServer from mcp.server and never use FastMCP.
Choose the effort. The default low skips extended thinking and is fastest. Use high or max for a bug that touches several files.
Keep max tokens at 8192 for full server files, or lower it for a quick snippet.
Attach a screenshot of an Inspector error if you have one. The model reads images, and max_image_resolution (default 0.5 megapixels) scales them down.
Generate, copy, and test. Paste the result into server.py and run it through mcp dev before trusting it.
Parameter
Required
Default
What it does
prompt
Yes
none
Your request
system_prompt
No
empty
Fixes role and constraints for the session
effort
No
low
Thinking depth, from fastest to deepest
max_tokens
No
8192
Output length cap
image
No
none
A screenshot or diagram as context
Prefer another model? Kimi K2.6 sits in the same category and is described for building agents and writing code.
Ship It Over HTTP
Switch the Transport
Change one line at the bottom of server.py:
if __name__ == "__main__":
mcp.run(transport="streamable-http", port=3001)
Clients now connect to http://127.0.0.1:3001/mcp. You can also leave the file alone and run uv run mcp run server.py --transport streamable-http. The run() call accepts these options:
host and port, defaulting to 127.0.0.1 and 8000
streamable_http_path, defaulting to /mcp
json_response=True to answer each POST with one JSON body
stateless_http=True for a fresh transport per request
Register the remote server in Claude Code with claude mcp add --transport http image-studio https://mcp.example.com/mcp.
Lock It Down
Once your server leaves localhost, three things change:
The Host allowlist. The default accepts only 127.0.0.1, localhost and [::1]. Behind a real hostname, every request fails with 421 Misdirected Request and Invalid Host header. Fix it with transport_security= and list both "mcp.example.com" and "mcp.example.com:*" in allowed_hosts.
Authorization. Your server is an OAuth 2.1 resource server. Implement TokenVerifier with one async verify_token method that returns an access token or None, and pass token_verifier= together with auth=.
TLS behind a proxy. When a load balancer ends TLS, start uvicorn with --proxy-headers so it trusts the forwarded headers.
💡 A 421 is a plain HTTP response, not a protocol error, so the client only shows a generic transport failure. The offending hostname appears in the server's log. A freshly deployed server that refuses every connection is a Host allowlist problem until proven otherwise.
Build Your Own Images With Picasso IA
You now have a server that registers tools, resources and prompts, passes an in-memory test, runs inside Claude, and can be deployed behind a real hostname. The part worth your next hour is the tool itself: swap the model slug, add an edit_image tool, or wire a video tool next to it.
Try the model your server just called. PicassoIA Image turns a prompt into a finished picture in seconds, and you can test any prompt in the browser before you automate it. When a still is not enough, PicassoIA Video and Seedance 2.5 Lite animate a prompt or a photo into short clips.
Open Picasso IA, pick a model, and run the same prompt you would give your tool. Then wire it into your server and let Claude do the clicking.