Large Language ModelsGenerate imagesGenerate videos
MCP Stateless Spec: Stateless vs Stateful Servers Explained
The 2026-07-28 MCP spec removed the initialize handshake and the Mcp-Session-Id header. This article compares stateless and stateful servers, shows where tool state goes now with handles, tasks and multi-round-trip requests, and lists the migration steps.
Every MCP server you built before this summer probably opens the same way: a client connects, sends initialize, waits for the reply, sends initialized, and only then gets to do real work. The 2026-07-28 revision of the Model Context Protocol deletes that ritual. There is no handshake, no Mcp-Session-Id header, and no protocol-level session pinned to one process. If you run an MCP server behind a load balancer, or you postponed that deployment because sticky sessions felt like a trap, this is the change you were waiting for.
This article breaks down the MCP stateless spec, what stateless vs stateful servers mean in practice, and, most important, what happens to the state your tools still need. You will see the exact fields and headers that changed, a migration checklist, and a short section on how an image generation connection fits the same pattern.
What Changed in the 2026-07-28 Spec
The Agentic AI Foundation, the Linux Foundation project that now stewards MCP, summarized the release in its migration post. The short version: the protocol stopped assuming a long conversation between one client and one server, and started treating every call like a normal HTTP request.
The Handshake Is Gone
Under the 2025-era protocol, the first thing a client did was negotiate. The initialize request carried a protocol version and a list of capabilities, the server answered with its own, and the client confirmed with initialized. Everything after that depended on what had been agreed on that specific connection.
The new revision removes the exchange entirely. The server no longer builds a private memory of each client, so two requests from the same client can be answered by two different machines without either one noticing.
Every Request Carries Its Own Context
Because nothing is negotiated up front, each request announces itself. A _meta object in the JSON-RPC envelope carries the protocol version, the client identity, and the capability flags. Client info looks like this:
The practical effect is a server that reads everything it needs from the request in front of it. Here is what moved:
Concern
2025-era protocol
2026-07-28 protocol
Version agreement
Negotiated once in initialize
Sent in _meta on every request
Client identity
Stored on the session
Sent in _meta on every request
Capabilities
Negotiated at connect time
Sent per request, plus an optional lookup call
Session tracking
Mcp-Session-Id header
Removed
List endpoints
Could vary per connection
Same answer for every caller
💡 Tip: A client can still fetch a server's capabilities up front when it wants to. It is an optional call now, not a mandatory first step.
Routing Headers for Gateways
On Streamable HTTP, the spec also defines headers that let infrastructure route traffic without parsing a JSON body:
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
A gateway can send tools/call for search to one pool and everything else to another, using headers alone. Rate limiting and logging get simpler for the same reason.
What Stays the Same
Nothing about the model's view of your server changes. Tools, resources, and prompts are still the three primitives, requests are still JSON-RPC, and a tool still receives arguments and returns a result. The difference is behind the curtain: the list endpoints no longer vary per connection, so tools/list returns the same answer to every caller instead of a per-session variation. That one rule is what makes caching a tool list at the edge safe.
Stateful vs Stateless in Plain Terms
The terms get thrown around loosely, so here is the working definition. A stateful server keeps something between requests, and the next request only makes sense if it reaches the same place. A stateless server keeps nothing between requests, and every request contains everything needed to answer it.
A letter works the opposite way. The address, the return address, and the stamp are all on the outside, so any clerk in any branch can sort it without calling anyone. That is a stateless server, and it is exactly how a 2026-07-28 MCP request behaves: protocol version, client identity, and capabilities all travel with the call.
The Cost Behind a Load Balancer
The Streamable HTTP transport introduced in the 2025-03-26 revision let a server issue an Mcp-Session-Id during initialization. The client echoed it on every later request, and the server used it to find the right page in its ledger: negotiated capabilities, per-user context, sometimes open subscriptions. Servers on the stdio transport were stateful in an even simpler way, because the process itself was the session.
Once that ledger sits in one process's memory, your load balancer has to keep sending the same client to the same process. Teams solved it in two ways. Sticky sessions skew traffic and break whenever a node restarts. A shared store such as Redis adds latency and a new single point of failure. Neither is free.
With sessions gone, any request can land on any instance behind a plain round-robin balancer, like cars filling identical toll lanes. Here is the comparison side by side:
Question
Stateful server
Stateless server
Where does memory live?
In the process or a session store
In the request, or in your own database
Load balancer
Sticky routing or shared store
Plain round robin
A node dies
Sessions on that node are lost
The next request goes elsewhere
Scaling out
Add nodes plus session plumbing
Add nodes
Debugging
Replay a whole session
Replay one request
Serverless fit
Awkward
Natural
Where Your State Goes Now
Removing sessions from the protocol does not make your application stateless. A shopping cart, a browser tab, and a half-finished workflow still exist. The difference is that the state now lives where it belongs, in your own storage, and the protocol no longer hides it.
💡 Rule of thumb: If the model needs to continue something later, give it a handle. If the user must answer something mid-call, use a multi-round-trip request. If the work is slow, use a task.
Explicit Handles
The recommended pattern is the one REST APIs have used for decades. One tool call mints an identifier and returns it, and the model passes it back as an argument on later calls. The paper tag on that basket plays the same role as a basket_id.
The server looks the basket up in a database on each call. Any instance can serve any step, a restart loses nothing, and the model can resume the work in a brand new conversation as long as it still holds the ID.
Multi-Round-Trip Requests
Sometimes a tool needs a confirmation halfway through, such as "delete these 40 files?" In a session world the server would pause and wait on an open connection. Under SEP-2322 the response instead carries resultType: "input_required" and an opaque requestState token. The client retries the same call with the answers in inputResponses.
Because the progress rides inside that token, whichever instance receives the retry can pick up exactly where the last one stopped.
Tasks for Slow Work
Long jobs follow the claim ticket model. With the Tasks extension (SEP-2663), the client receives a taskId immediately and polls tasks/get until the work finishes. The call is decoupled from its execution, so a ten minute render never holds a connection open, and each poll can reach any instance.
Situation
Pattern
What travels between calls
Work that continues later
Explicit handle
An ID such as basket_id
Confirmation mid-call
Multi-round-trip request
requestState and inputResponses
Job that takes minutes
Tasks extension
A taskId to poll
Migrating a Server Without Pain
Audit What You Store
Start by finding every place your server remembers something about a client between requests. Typical culprits:
Auth context saved at initialize time
Per-session caches or rate counters held in memory
Capability checks that read negotiated flags instead of _meta
Tool lists that change depending on who connected
Subscriptions tied to an open connection
Each one needs a new home: the request itself, your database, or an explicit handle.
Use the SDK Codemod
The TypeScript SDK v2 splits into side-specific packages and ships a codemod for the mechanical changes:
The v2 SDKs keep speaking the 2025-era protocol by default, and serving 2026-07-28 is an explicit opt-in. That lets you ship the code change first and flip the protocol switch when your clients are ready. The v1.x line keeps receiving bug and security fixes for at least six months after v2.
Mind the Deprecation Clock
The maintainers promise at least twelve months between deprecation and removal, and the earliest removal date for deprecated features is July 28, 2027. Migration write-ups list Roots, Sampling, and Logging among the deprecated features (SEP-2577), with server-initiated flows moving to multi-round-trip requests. Plan the work, but this is not an emergency.
Security Gets Stricter
A session let a server say "this client logged in earlier." That shortcut disappears. Every request must carry credentials, and every request must be checked. If validation cost worries you, hold the result in memory for a few seconds, but never trust a request because the previous one looked fine.
Make handles unguessable. A handle is only an ID, and IDs are classic attack targets. b_47f2 works in a diagram. In production, generate long random values, store the owner next to the record, and verify that the caller owns the handle on every call. Expire the ones you no longer need.
The new routing headers help here too. Because Mcp-Method and Mcp-Name are visible without opening the body, a gateway can apply a policy per tool, such as stricter rate limits on a payment tool or an allow list for a destructive one, before the request ever reaches your code. Defense in depth is easier when the outer layer can read the label on the envelope.
💡 Rule of thumb: Treat every handle like a public URL parameter. Assume someone will try the next one.
Should You Go Stateless?
For most servers, yes. Read-only lookups, CRUD over a database, search, and anything that wraps a REST API have nothing to remember, so the move is mostly deleting code. You gain simpler deployments, a natural fit for serverless platforms, and failures that hit one request instead of a whole conversation.
Some tools hold something live: a browser page, a shell, a render in progress. Keep that state, but keep it behind a handle with an expiry, in a store every instance can reach. The protocol is stateless. Your backend does not have to be.
A quick test tells you how ready a server is. Pick any request in flight, kill the instance handling it, and retry the same call against a different instance. If the answer is identical, you are stateless where it counts. If the retry fails, asks the client to start over, or returns something subtly different, there is still a ledger hiding in memory, and that is the code to move into a database or behind a handle first.
Server type
Best fit
Why
Read-only data lookups
Fully stateless
Nothing to remember
Database CRUD
Stateless, record ID as handle
The database already holds the truth
Browser or shell control
Stateless protocol, stateful backend
Live resource stays behind an expiring handle
Long renders and batch jobs
Tasks extension
Polling replaces open connections
Generating Images Over MCP
MCP is also how assistants reach creative tools, and image generation happens to be a clean example of the handle pattern. PicassoIA exposes its models through a developer API and an MCP connection. The API lives at https://api.picassoia.com/v1, uses a Bearer token that begins with pia_sk_, and follows a Replicate style layout: POST /v1/models/{owner}/{name}/predictions to start a job, GET /v1/predictions/{id} to check it, and POST /v1/predictions/{id}/cancel to stop it.
Prediction IDs Are Handles Too
Generation is asynchronous. Starting a job returns a prediction ID immediately, and the assistant calls the status tool with that ID after the suggested wait, again and again, until the status reads succeeded or failed. No open connection sits idle while the GPU works, and the ID carries all the continuity. It is the basket_id pattern applied to pixels.
A few limits worth knowing when you plan a workflow: 5 concurrent predictions per account, shared across tokens and MCP connections, prompts up to 4,000 characters, and a 3 hour timeout per job.
Models You Can Call
These four models are available through both the API and the MCP connection:
The photos in this article came from P Image, one of the many text to image models on the platform. When you are writing the tool descriptions and JSON schemas your own MCP server will expose, a large language model saves time. Claude Sonnet 5, GPT 5.6 Sol, and Gemini 3.5 Flash are all available for drafting and reviewing that kind of structured text.
When you are ready for more, browse every model on the all models page, turn a favorite image into a short clip with a video model, and keep experimenting. Each prompt you write is a small request that carries everything it needs, which is exactly the point of this spec.