Large Language ModelsGenerate imagesGenerate videos
How Does MCP Work? Under the Hood With Claude and AI Agents
Model Context Protocol looks like magic when Claude reads a file or calls an image API on its own. This article opens it up: hosts, clients, and servers, the JSON-RPC handshake, tools versus resources versus prompts, the agent loop that repeats tool calls, and the security checks that keep it safe.
You type a request, Claude reads a file, queries a database, or renders an image, and a few seconds later the result is sitting in the chat. Nothing inside the model's weights knows how to reach your database. Something sits in the middle, and that something is the Model Context Protocol, shortened to MCP. If you have ever asked how does MCP work behind the curtain, here is the short version: a shared message format lets an AI application ask outside programs what they can do, then call them on the model's behalf. The rest of this article follows one request from the first handshake to the final tool result, with real message shapes, so you can picture every hop.
What MCP Actually Is
MCP is an open standard that Anthropic introduced in November 2024. It defines how an AI application connects to external tools and data through one shared set of rules, instead of a custom integration for every pair of products. Think of how USB-C lets a single cable serve a microphone, a drive, and a monitor. MCP plays that role between language models and the software around them.
The Problem It Solves
Before MCP, every AI app that wanted to read a calendar, search a codebase, or call an image API needed its own glue code. With N applications and M tools, teams faced up to N × M separate integrations, each with different authentication quirks, error formats, and update schedules. A bug fixed in one connector never helped the others, and every new model or tool multiplied the work.
Why One Protocol Wins
A shared protocol changes the math from N × M to N + M. A tool author writes one server. An application author writes one client. Everything else connects without extra code.
Reuse: one server for a database works in Claude Desktop, Claude Code, an editor, or a custom agent.
Swap freedom: change the model behind your application and the tool connections keep working.
Isolation: each server runs as its own process, so a crash or a bug stays contained.
Runtime lookup: clients ask servers what they offer when they connect, so new tools appear without a new release of the host.
💡 Worth remembering: MCP does not make a model smarter. It gives the model hands. The reasoning still happens inside the model, and the protocol only carries requests out and results back.
The Three Players in Every Session
Picture a restaurant. The dining room is the host, the waiter is the client, and each kitchen station is a server. The diner never talks to the stove, and the stove never talks to the diner. Everything travels through a defined hand-off, which is exactly the job MCP does.
Hosts
The host is the application you actually open: Claude Desktop, Claude Code, an editor with AI features, or an agent you wrote yourself. It owns the conversation with the model, draws the permission prompts, and decides which servers to start.
Clients
Inside the host, one client object manages one connection to one server. Connect to five servers and the host holds five clients. Each client keeps its own session state, negotiates capabilities, and translates between the host's internal calls and protocol messages.
Servers
A server is a small program that exposes capabilities. It can run on your laptop as a child process or on another machine behind a URL. One might wrap a file system, a GitHub account, a database, or an image generator.
Role
Where it lives
What it is responsible for
Example
Host
Your device or a cloud app
Model conversation, consent prompts, server startup
Claude Desktop, Claude Code
Client
Inside the host
One connection and one session per server
A connection object from an MCP SDK
Server
Local process or remote URL
Exposing tools, resources, and prompts
A filesystem server, a database server
Inside the Protocol Messages
Every conversation between a client and a server is a stream of small, boring, predictable messages. That predictability is the whole point.
JSON-RPC Over the Wire
Every MCP message is JSON-RPC 2.0, which gives the protocol exactly three message shapes:
A request carries an id and a method, and expects an answer.
A response repeats that id and carries either a result or an error.
A notification has no id and expects nothing back.
Here is a tool call as it travels from client to server:
The shared id is how the client matches an answer to its question, even when several calls are in flight at once.
The Handshake in Order
Before any tool runs, client and server agree on how to talk. The sequence is short and always the same:
initialize (request): the client sends the protocol version it wants, the capabilities it supports, and its own name and version.
initialize (response): the server answers with the version it will use, its own capabilities, and optional written instructions for the model.
notifications/initialized: the client confirms, and the session is open.
tools/list, resources/list, prompts/list: the client asks what the server offers, limited to the capabilities both sides declared.
Normal traffic: calls, reads, and the occasional notifications/tools/list_changed when a server adds or removes a tool mid-session.
💡 Version mismatch: if a server does not support the version a client asked for, it replies with one it does support. The client either accepts that version or disconnects cleanly.
stdio vs Streamable HTTP
The same JSON-RPC messages can travel over two official transports.
With stdio, the host launches the server as a child process and exchanges newline-delimited JSON through standard input and output. Logs must go to standard error, because one stray print to standard output corrupts the stream. With Streamable HTTP, the server sits behind a single URL, the client sends messages with POST requests, and the server answers with plain JSON or opens a streaming response when it needs to send several messages back. Streamable HTTP replaced the older HTTP plus SSE transport in the March 2025 revision of the specification.
Feature
stdio
Streamable HTTP
Where the server runs
Your machine, as a child process
Anywhere reachable by URL
Setup
One command in a config file
A deployed endpoint
Authentication
Inherits your user permissions
OAuth based authorization
Typical users
Solo developers, local tools
Teams, hosted services, shared connectors
Concurrent clients
One
Many
A typical Claude Desktop entry for a local server looks like this:
Servers expose three building blocks. They differ mainly in who decides when to use them.
Primitive
Controlled by
Purpose
Example
Methods
Tools
The model
Perform an action or run a computation
generate_image, run_query
tools/list, tools/call
Resources
The application
Supply read-only context
A file, a database schema
resources/list, resources/read
Prompts
The user
Offer reusable templates
A "review this pull request" command
prompts/list, prompts/get
Tools Do Things
A tool has a name, a description, and an inputSchema written in JSON Schema. The description is what the model actually reads when it decides whether to call the tool, so a vague description produces wrong or missing calls. Write it like a short manual for a colleague who has never seen your system.
Real connectors show the pattern well. The claude.ai PicassoIA connector exposes tools such as generate_image, edit_image, generate_video_picassoia, generate_video_seedance, get_generation, and list_models. A generation tool returns a prediction id right away, and the model then calls get_generation again and again until the status reads succeeded. That is one user request turning into several chained tool calls, and the protocol never needed a special feature for it.
Resources Supply Context
Resources are addressed by URI, such as file:///project/README.md. The host decides which ones to attach to the conversation, and resources/read returns the contents with a MIME type. Servers can also publish resource templates with parameters, and clients can subscribe to changes so the context stays current.
Prompts Package Workflows
Prompts are templates the user triggers on purpose, usually as slash commands. A prompt can take arguments and return a ready-made set of messages, so a team can ship its best "summarize this incident" or "write release notes" wording once and reuse it everywhere.
What Happens When Claude Calls a Tool
Here is the part people most often get wrong: the model never speaks MCP. The host does. Claude only sees tool definitions in its own tool-use format and emits structured requests. Everything else is plumbing.
From Question to Tool Call
You ask. "Resize the hero photo and save it to my project folder."
The host sends context. It forwards your message plus the tool definitions gathered from every connected server to the model API.
Claude decides. Instead of final text, it returns a tool_use block naming a tool and its arguments.
The host checks consent. It may show an approval prompt, then routes the call to the client that owns that tool.
The client calls the server. A tools/call request goes out, the server does the work, and it returns content blocks.
The host reports back. The result goes to Claude as a tool_result block, together with the conversation so far.
Claude continues. It calls another tool or writes the final answer.
Where the Agent Loop Lives
Steps 3 to 7 repeat until Claude stops asking for tools. That repetition is the agent loop, and it lives in the host, not in the protocol. MCP defines the doors. The host decides how many times to walk through them. An AI agent is simply a host that runs this loop with longer plans, fewer interruptions, and sometimes helper agents of its own.
💡 Hidden cost: every tool definition takes space in the context window. Fifty tools can eat thousands of tokens before the first word of your question is read. Good hosts load schemas lazily or let you switch servers off per project.
Security and Permissions
A server that can run commands or touch files is powerful, which is precisely why it needs guardrails.
Consent Comes First
The specification asks hosts to get explicit user consent before invoking tools or sharing data, and mature hosts follow through with approval prompts and per-tool allow lists. Remember that a local stdio server runs with your user permissions, so treat installing one the way you would treat installing any program. Remote servers add OAuth, which lets you grant narrow access and revoke it later.
Common Failure Points
Prompt injection through results: a web page, ticket, or email returned by a tool can contain text that tries to give the model new orders.
Poisoned tool descriptions: a malicious server can hide instructions inside its own descriptions.
Over-broad credentials: a token that can delete everything will eventually be used to delete everything. Prefer read-only access.
Noisy standard output: debugging prints in a stdio server break the JSON stream.
Too many tools: with dozens available, the model picks the wrong one more often. Keep each server focused.
Long jobs: a single call that runs for ten minutes times out. Return an id and let the model poll, as image and video generators do.
Use Claude Sonnet 5 on PicassoIA
If you want to build your own server, Claude Sonnet 5 on PicassoIA is a handy pair programmer. It is built for multi-step coding and tool-use tasks, and it reads images, so a screenshot of an error works as input.
Open the model page. Go to the Claude Sonnet 5 page in the large language models collection.
Write the prompt. This is the only required field. Be specific about the language, the SDK, and the tool you want.
Set the effort. The default is low, which skips extended thinking for fast answers. Raise it for a bug that spans several files.
Adjust the limits.max_tokens defaults to 8192, and an optional system_prompt locks in a role or coding style for the session.
Attach an image if useful. The image field accepts a screenshot, and max_image_resolution (default 0.5 megapixels) keeps it light.
Generate and review. Copy the code into your project and test it with the MCP Inspector before trusting it.
A prompt that works well:
Write a minimal MCP server in TypeScript using the official SDK. It exposes one tool,
word_count, that takes a string and returns the number of words. Use the stdio transport
and log only to stderr. Include the claude_desktop_config.json entry to register it.
PicassoIA also speaks MCP itself. Its claude.ai connector lists four models: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video, and Seedance 2.5 Lite, which produces video with audio. Predictions are asynchronous, and an account can run five at a time across all of its connections, so keep each batch small. Check the pricing page for what your plan includes.
Build Your Own Visuals Today
Behind every smooth "Claude, make me an image" moment there is a handshake, a schema, and a loop. The fastest way to feel it is to produce something. Open PicassoIA Image and describe a scene, switch to Seedream 4.5 or FLUX 2 Pro for a different look, then hand the result to Seedance 2.0 to turn a still into a short clip. Connect the PicassoIA connector to Claude and you can ask for all of it in plain language, then watch the tool calls from this article play out in real time.
Pick one idea from the sections above, write a single sentence about what you want to see, and run it. Your first image is a few seconds away at picassoia.com.