Large Language ModelsGenerate imagesGenerate videos
MCP Server Architecture Explained With Diagrams: From Host to Tool Call
An MCP server sits between an AI application and the systems it needs to reach. This article draws the whole path in diagrams: host, client, server, transport, JSON-RPC handshake, tool call, error handling and security, with a real image and video connector as the worked example.
Your AI assistant can draft a contract in seconds, yet it cannot read your calendar, query your database or resize a photo on its own. The Model Context Protocol, shortened to MCP, closes that gap with one shared contract between AI applications and the systems around them. An MCP server is the small program on the far side of that contract: it announces what it can do, waits for requests and returns results in a predictable shape. This article draws that architecture piece by piece. Every diagram is plain text, so it survives copy and paste into a README, a design doc or a pull request.
Why MCP Exists
Before MCP, every AI application that wanted to reach a database, a calendar or a file system needed its own custom connector. Each connector carried its own authentication, error format and bugs. Three apps and three tools already meant nine integrations, and the grid grows with every new product on either side. MCP replaces the grid with a shared protocol, so the arithmetic changes from apps times tools to apps plus tools.
Before MCP: one custom connector for every pair
App A ──► Database App B ──► Database App C ──► Database
App A ──► Calendar App B ──► Calendar App C ──► Calendar
App A ──► Files App B ──► Files App C ──► Files
3 apps x 3 tools = 9 connectors to build and maintain
With MCP: one shared protocol in the middle
App A ──┐ ┌── Database server
App B ──┼──── MCP (JSON-RPC) ─────┼── Calendar server
App C ──┘ └── Files server
3 clients + 3 servers = 6 pieces
The word server misleads people. An MCP server is not a language model and it does not think. It is an ordinary program, written in TypeScript, Python or any language with a JSON library, that wraps a real capability and describes it in a format every compliant client can read. The model never needs to know how your database driver works. It only needs to know that a tool called run_query exists and what arguments it accepts.
The Three Roles in One Diagram
MCP defines three roles, and mixing them up causes most of the confusion. Here is the whole picture before the details.
┌──────────── HOST (the AI application) ────────────┐
│ The LLM picks a tool, the host routes the call │
│ │
│ ┌────────────┐ ┌────────────┐ ┌────────────┐ │
│ │ client 1 │ │ client 2 │ │ client 3 │ │
│ └─────┬──────┘ └─────┬──────┘ └─────┬──────┘ │
└────────┬───────────────┬───────────────┬──────────┘
│ session │ session │ session
┌─────┴──────┐ ┌─────┴──────┐ ┌─────┴──────┐
│ Server A │ │ Server B │ │ Server C │
│ files │ │ GitHub │ │ image tool │
└────────────┘ └────────────┘ └────────────┘
The Host Owns the Conversation
The host is the application a person actually uses: a desktop chat app, an IDE assistant or a custom agent. It runs the language model, decides which servers to connect to and shows the user what is about to happen. Reasoning models such as Claude Sonnet 5, GPT 5.6 Sol and Gemini 3.1 Pro sit inside the host, but they never speak MCP themselves. The host translates between the model's tool-calling format and the protocol, which is why one server works with many different models.
The Client Holds One Session
Inside the host, one MCP client is created per server. Each client keeps a stateful, one-to-one session with exactly one server, remembers the capabilities both sides agreed on and routes every message. A host connected to three servers runs three clients, as the diagram shows. If one session drops, the other two keep working.
The Server Does the Work
The MCP server wraps a real system: a Postgres database, a GitHub account, a folder of documents or an image API. It stays deliberately small. It declares what it offers, validates input, performs the action and returns structured output. It never sees the full conversation, only the requests sent to it, a design choice that protects privacy and keeps servers reusable.
A quick recap you can pin above your desk:
Host: owns the model, the user interface and the consent prompts.
Client: one per server, speaks the protocol and holds session state.
Server: exposes capabilities, runs the action and returns results.
What a Server Exposes
A server offers three building blocks, called primitives. They differ in one detail that shapes the whole design: who decides when each one is used.
Primitive
Controlled by
Typical use
Example methods
Tools
The model
Run an action or compute a result
tools/list, tools/call
Resources
The application
Supply read-only context such as files or records
resources/list, resources/read
Prompts
The user
Reusable templates, often shown as slash commands
prompts/list, prompts/get
Tools Run Actions
A tool has a name, a plain-language description and an inputSchema written in JSON Schema. The model reads the description to decide whether the tool fits the request, and the schema keeps the arguments valid. Descriptions deserve real effort, because a vague one makes the model guess. Name tools like verbs, keep each one narrow and return only the fields the model needs. A tool called search_orders with three typed arguments beats a single do_anything tool with one free-text field. Tool results are not limited to text: they can carry images, audio or links, which is why media generators fit the protocol so naturally. An image tool might wrap Flux 2 Pro or Seedream 4.5, and a video tool might wrap Veo 3.1 or Kling v3 Video.
Resources Supply Context
Resources are read-only data addressed by URI, such as file:///reports/q3.md or postgres://db/customers/schema. The host chooses which ones to attach to the model's context, so resources suit documents, schemas and logs. After a client subscribes, a server can announce changes with notifications/resources/updated.
Prompts Offer Templates
Prompts are parameterized message templates that a user picks on purpose, usually from a slash-command menu. A prompt such as review this pull request returns a ready-made list of messages, so every teammate starts from the same wording and the same checklist.
Traffic also flows the other way. Servers can ask the client for help through sampling (request a completion from the host's model), roots (ask which directories are in scope) and elicitation (ask the user for missing input). Hosts decide whether to allow each one.
💡 Rule of thumb: if the action changes something in the world, build a tool. If it only supplies information, start with a resource.
Transports: stdio or Streamable HTTP
The transport decides how bytes travel between client and server. The messages stay identical either way: JSON-RPC 2.0 requests, responses and notifications. Two transports are standard.
stdio for Local Servers
stdio: the host starts the server as a child process
┌────────┐ stdin: requests ┌───────────┐
│ Client │ ─────────────────────► │ Server │
│ │ ◄───────────────────── │ process │
└────────┘ stdout: responses └───────────┘
stderr: logs only
With stdio, the host launches the server as a child process and exchanges newline-delimited JSON-RPC messages over standard input and output. Setup is one command line, latency is tiny and credentials arrive through environment variables. One rule trips up many first-time authors: the server must never print anything except protocol messages to stdout. Logs go to stderr, or the stream corrupts and the session dies.
Streamable HTTP for Remote Ones
Streamable HTTP: one URL, many clients
┌──────────┐ POST /mcp ┌────────────┐
│ Client A │ ───────────────────► │ │
└──────────┘ ◄─────────────────── │ Server │
┌──────────┐ JSON or SSE reply │ (web app) │
│ Client B │ ───────────────────► │ │
└──────────┘ ◄─────────────────── └────────────┘
Streamable HTTP serves many clients from one endpoint. The client sends each message as an HTTP POST, and the server answers with plain JSON or opens a Server-Sent Events stream when it needs to send several messages. A session identifier travels in the Mcp-Session-Id header. This transport replaced the older HTTP plus SSE design in the 2025-03-26 revision of the specification, and it is the right pick for hosted servers, multi-user products and anything behind a load balancer. Authorization is built on OAuth, so a server can answer 401 and point the client toward its authorization server. Servers that still run the previous design can stay compatible by serving both endpoints during a migration, but a new project should start on Streamable HTTP.
Question
stdio
Streamable HTTP
Where does the server run?
On the same machine as the host
Anywhere reachable by URL
Users per server
One
Many
Credentials
Environment variables
OAuth or HTTP headers
Best for
Developer tools, local files
Hosted products, shared services
Main pitfall
Stray output on stdout
Session handling behind proxies
A practical split: ship a stdio build for developers who want to try the server in a minute, and a Streamable HTTP build for everyone else. The tool code stays the same. Only the entry point changes.
One Tool Call, Traced Step by Step
Here is a single request from the moment the user types to the moment the answer appears.
User Host + LLM MCP client MCP server
│ │ │ │
├─ asks for image ───► │ │
│ │ picks a tool │ │
│ ├─ tool request ────► │
│ │ ├─ tools/call ──────►
│ │ │ │ does the work
│ │ ◄─ text, isError ───┤
│ ◄─ result ──────────┤ │
◄─ answer + URL ─────┤ │ │
│ │ │ │
Step 1: The Handshake
Every session opens with initialize. The client sends the protocol version it supports plus its own capabilities. The server replies with the version it chose and the capabilities it offers. The client then sends a notifications/initialized message and normal traffic begins. If the versions cannot be reconciled, the client disconnects instead of guessing.
The client sends tools/list, the host hands the schemas to the model, and the model decides whether to call one. When a server's tool list changes at runtime, it sends notifications/tools/list_changed so the client can refresh. A call looks like this:
MCP separates two kinds of failure. A protocol error is a JSON-RPC error object, for example code -32602 for invalid parameters or an unknown tool name. A tool execution error is a normal result with isError: true and a message the model can read. The second kind matters most. When a tool answers prompt too long, the model can shorten the prompt and retry, but only if the error reaches it as text instead of a crashed session. Either side can also send notifications/cancelled to abandon a slow request.
Run the server under the MCP Inspector (npx @modelcontextprotocol/inspector) before any model touches it. The Inspector lists the tools, lets you fire calls by hand and shows the raw JSON-RPC traffic, so you can separate protocol bugs from prompt bugs.
3 Common Design Mistakes
Most production trouble traces back to the same three choices:
One giant tool. A tool named do_anything with a free-text argument forces the model to guess. Split it into narrow verbs such as search_orders and refund_order, each with typed arguments.
Chatty results. Returning a 40,000-token blob burns the model's context window. Return the fields the model needs and link to a resource for the rest.
Hidden state. If a tool only works after another tool has run, say so in its description, or the model will call them in the wrong order.
A Real Server: Image and Video Tools
A concrete case shows why architecture choices matter. PicassoIA offers an MCP connector whose tools generate and edit images and videos on its own GPUs. Image and video jobs take seconds to minutes, which breaks the naive picture of call a tool, wait for the answer. Hosts and SDKs usually enforce a request timeout, so a server that blocks until the render finishes would fail exactly when the work is almost done.
Async Jobs Behind a Simple Tool
The connector exposes tools named generate_image, edit_image, generate_video_picassoia, generate_video_seedance, get_generation, cancel_generation, list_models, list_generations and get_account. A generate call returns a prediction ID and an estimated time at once. The model then calls get_generation after the suggested wait, and again after each new hint, until the status reads succeeded or failed. The tool call stays short while the heavy work runs as a background job on a GPU worker.
A server that fronts a GPU queue has to protect it. PicassoIA lists a ceiling of five concurrent predictions per account, shared across API credentials and MCP connections. A well-built server turns a ceiling like that into a clear tool result, such as five jobs running, retry in 30 seconds, instead of letting requests pile up behind it. The model can read that message and wait, which is far better than a timeout.
Four habits make an async server pleasant to use from an agent:
Return fast. Hand back an ID within a second and never block for minutes.
Hint the wait. Tell the model when to poll next so it does not spam the status tool.
Make failure final. A failed job stays failed, and the message says why.
Offer a cancel tool. Users change their minds, and queued work costs money.
Where Security Belongs
MCP moves capability, so it also moves risk. The protocol sets the shape of the conversation, but the host and the server carry the responsibility. Six checks catch most problems before launch:
User consent. The host should show which tool is about to run and ask before any action with side effects.
Least privilege. Give a read-only tool a read-only credential, and keep powerful actions in a separate server.
Input validation. Treat every argument as untrusted. Validate against the schema on the server too, not only in the client.
Prompt injection. Text inside a tool result or a resource can contain instructions. The host should treat it as data, never as a command from the user.
Secret handling. Never place credentials in tool descriptions or results. stdio servers read them from the environment, and HTTP servers use OAuth.
Audit logs. Write structured logs with a request ID to stderr or a log service, so every tool call can be traced afterward.
Deployment follows the same logic. Pin the protocol version in your tests, run the server under the Inspector in CI, and put remote servers behind a gateway that handles TLS, rate limits and OAuth, so the tool code stays focused on the tools.
Try Image and Video Generation Yourself
Diagrams stick better when you can watch a tool call produce something real. Open PicassoIA, write a prompt for a photo of your own desk, a server room or a whiteboard sketch, and generate it with Seedream 4.5 or GPT Image 2. Then animate your favorite frame with Veo 3.1 or Kling v3 Video. If your host supports MCP connections, add the PicassoIA connector and let your assistant run the job while you follow the polling steps from the diagram above.
Try three prompts and change one thing each time: the camera angle, the lighting direction or the lens. The differences show how much a precise prompt matters, the same way a precise tool description matters to a model. Browse every available model at picassoia.com/en/all-models and start your first generation today.