Large Language ModelsGenerate imagesGenerate videos

MCP Stateless Spec: Stateless vs Stateful Servers Explained

The 2026-07-28 MCP spec removed the initialize handshake and the Mcp-Session-Id header. This article compares stateless and stateful servers, shows where tool state goes now with handles, tasks and multi-round-trip requests, and lists the migration steps.

MCP Stateless Spec: Stateless vs Stateful Servers Explained
Cristian Da Conceicao
Founder of Picasso IA

Every MCP server you built before this summer probably opens the same way: a client connects, sends initialize, waits for the reply, sends initialized, and only then gets to do real work. The 2026-07-28 revision of the Model Context Protocol deletes that ritual. There is no handshake, no Mcp-Session-Id header, and no protocol-level session pinned to one process. If you run an MCP server behind a load balancer, or you postponed that deployment because sticky sessions felt like a trap, this is the change you were waiting for.

This article breaks down the MCP stateless spec, what stateless vs stateful servers mean in practice, and, most important, what happens to the state your tools still need. You will see the exact fields and headers that changed, a migration checklist, and a short section on how an image generation connection fits the same pattern.

Rows of identical server racks in a bright data center aisle

What Changed in the 2026-07-28 Spec

The Agentic AI Foundation, the Linux Foundation project that now stewards MCP, summarized the release in its migration post. The short version: the protocol stopped assuming a long conversation between one client and one server, and started treating every call like a normal HTTP request.

The Handshake Is Gone

Under the 2025-era protocol, the first thing a client did was negotiate. The initialize request carried a protocol version and a list of capabilities, the server answered with its own, and the client confirmed with initialized. Everything after that depended on what had been agreed on that specific connection.

The new revision removes the exchange entirely. The server no longer builds a private memory of each client, so two requests from the same client can be answered by two different machines without either one noticing.

Every Request Carries Its Own Context

Because nothing is negotiated up front, each request announces itself. A _meta object in the JSON-RPC envelope carries the protocol version, the client identity, and the capability flags. Client info looks like this:

{
  "_meta": {
    "io.modelcontextprotocol/clientInfo": {
      "name": "my-app",
      "version": "1.0"
    }
  }
}

The practical effect is a server that reads everything it needs from the request in front of it. Here is what moved:

Concern2025-era protocol2026-07-28 protocol
Version agreementNegotiated once in initializeSent in _meta on every request
Client identityStored on the sessionSent in _meta on every request
CapabilitiesNegotiated at connect timeSent per request, plus an optional lookup call
Session trackingMcp-Session-Id headerRemoved
List endpointsCould vary per connectionSame answer for every caller

💡 Tip: A client can still fetch a server's capabilities up front when it wants to. It is an optional call now, not a mandatory first step.

Routing Headers for Gateways

On Streamable HTTP, the spec also defines headers that let infrastructure route traffic without parsing a JSON body:

  • MCP-Protocol-Version: 2026-07-28
  • Mcp-Method: tools/call
  • Mcp-Name: search

A gateway can send tools/call for search to one pool and everything else to another, using headers alone. Rate limiting and logging get simpler for the same reason.

What Stays the Same

Nothing about the model's view of your server changes. Tools, resources, and prompts are still the three primitives, requests are still JSON-RPC, and a tool still receives arguments and returns a result. The difference is behind the curtain: the list endpoints no longer vary per connection, so tools/list returns the same answer to every caller instead of a per-session variation. That one rule is what makes caching a tool list at the edge safe.

Stateful vs Stateless in Plain Terms

The terms get thrown around loosely, so here is the working definition. A stateful server keeps something between requests, and the next request only makes sense if it reaches the same place. A stateless server keeps nothing between requests, and every request contains everything needed to answer it.

The Café That Remembers You

A barista handing a coffee to a smiling regular customer

Picture a café where the barista knows your order, your name, and that you skip the sugar. Ordering takes three words because the context lives in her head. That is a stateful server. It is fast and friendly, right up until she goes on break and the replacement has no idea who you are.

The Post Office That Doesn't

Hands sorting envelopes that carry their own address labels

A letter works the opposite way. The address, the return address, and the stamp are all on the outside, so any clerk in any branch can sort it without calling anyone. That is a stateless server, and it is exactly how a 2026-07-28 MCP request behaves: protocol version, client identity, and capabilities all travel with the call.

The Cost Behind a Load Balancer

A hotel concierge reading from a thick guest ledger

The Streamable HTTP transport introduced in the 2025-03-26 revision let a server issue an Mcp-Session-Id during initialization. The client echoed it on every later request, and the server used it to find the right page in its ledger: negotiated capabilities, per-user context, sometimes open subscriptions. Servers on the stdio transport were stateful in an even simpler way, because the process itself was the session.

Once that ledger sits in one process's memory, your load balancer has to keep sending the same client to the same process. Teams solved it in two ways. Sticky sessions skew traffic and break whenever a node restarts. A shared store such as Redis adds latency and a new single point of failure. Neither is free.

Aerial view of a toll plaza with traffic spread evenly across identical lanes

With sessions gone, any request can land on any instance behind a plain round-robin balancer, like cars filling identical toll lanes. Here is the comparison side by side:

QuestionStateful serverStateless server
Where does memory live?In the process or a session storeIn the request, or in your own database
Load balancerSticky routing or shared storePlain round robin
A node diesSessions on that node are lostThe next request goes elsewhere
Scaling outAdd nodes plus session plumbingAdd nodes
DebuggingReplay a whole sessionReplay one request
Serverless fitAwkwardNatural

Where Your State Goes Now

Removing sessions from the protocol does not make your application stateless. A shopping cart, a browser tab, and a half-finished workflow still exist. The difference is that the state now lives where it belongs, in your own storage, and the protocol no longer hides it.

💡 Rule of thumb: If the model needs to continue something later, give it a handle. If the user must answer something mid-call, use a multi-round-trip request. If the work is slow, use a task.

Explicit Handles

A wicker basket at a market stall with a numbered paper tag on the handle

The recommended pattern is the one REST APIs have used for decades. One tool call mints an identifier and returns it, and the model passes it back as an argument on later calls. The paper tag on that basket plays the same role as a basket_id.

create_basket()                           -> {"basket_id": "b_47f2"}
add_item(basket_id="b_47f2", sku="widget-123")
checkout(basket_id="b_47f2")

The server looks the basket up in a database on each call. Any instance can serve any step, a restart loses nothing, and the model can resume the work in a brand new conversation as long as it still holds the ID.

Multi-Round-Trip Requests

Sometimes a tool needs a confirmation halfway through, such as "delete these 40 files?" In a session world the server would pause and wait on an open connection. Under SEP-2322 the response instead carries resultType: "input_required" and an opaque requestState token. The client retries the same call with the answers in inputResponses.

Because the progress rides inside that token, whichever instance receives the retry can pick up exactly where the last one stopped.

Tasks for Slow Work

A dry cleaner clerk handing a customer a numbered claim ticket

Long jobs follow the claim ticket model. With the Tasks extension (SEP-2663), the client receives a taskId immediately and polls tasks/get until the work finishes. The call is decoupled from its execution, so a ten minute render never holds a connection open, and each poll can reach any instance.

SituationPatternWhat travels between calls
Work that continues laterExplicit handleAn ID such as basket_id
Confirmation mid-callMulti-round-trip requestrequestState and inputResponses
Job that takes minutesTasks extensionA taskId to poll

Migrating a Server Without Pain

Audit What You Store

A developer at a standing desk reviewing server code

Start by finding every place your server remembers something about a client between requests. Typical culprits:

  • Auth context saved at initialize time
  • Per-session caches or rate counters held in memory
  • Capability checks that read negotiated flags instead of _meta
  • Tool lists that change depending on who connected
  • Subscriptions tied to an open connection

Each one needs a new home: the request itself, your database, or an explicit handle.

Use the SDK Codemod

The TypeScript SDK v2 splits into side-specific packages and ships a codemod for the mechanical changes:

npm install @modelcontextprotocol/server
npx @modelcontextprotocol/codemod@latest v1-to-v2 .

The v2 SDKs keep speaking the 2025-era protocol by default, and serving 2026-07-28 is an explicit opt-in. That lets you ship the code change first and flip the protocol switch when your clients are ready. The v1.x line keeps receiving bug and security fixes for at least six months after v2.

Mind the Deprecation Clock

The maintainers promise at least twelve months between deprecation and removal, and the earliest removal date for deprecated features is July 28, 2027. Migration write-ups list Roots, Sampling, and Logging among the deprecated features (SEP-2577), with server-initiated flows moving to multi-round-trip requests. Plan the work, but this is not an emergency.

Security Gets Stricter

A session let a server say "this client logged in earlier." That shortcut disappears. Every request must carry credentials, and every request must be checked. If validation cost worries you, hold the result in memory for a few seconds, but never trust a request because the previous one looked fine.

Make handles unguessable. A handle is only an ID, and IDs are classic attack targets. b_47f2 works in a diagram. In production, generate long random values, store the owner next to the record, and verify that the caller owns the handle on every call. Expire the ones you no longer need.

The new routing headers help here too. Because Mcp-Method and Mcp-Name are visible without opening the body, a gateway can apply a policy per tool, such as stricter rate limits on a payment tool or an allow list for a destructive one, before the request ever reaches your code. Defense in depth is easier when the outer layer can read the label on the envelope.

💡 Rule of thumb: Treat every handle like a public URL parameter. Assume someone will try the next one.

Should You Go Stateless?

Two engineers sketching boxes and arrows on a whiteboard

For most servers, yes. Read-only lookups, CRUD over a database, search, and anything that wraps a REST API have nothing to remember, so the move is mostly deleting code. You gain simpler deployments, a natural fit for serverless platforms, and failures that hit one request instead of a whole conversation.

Some tools hold something live: a browser page, a shell, a render in progress. Keep that state, but keep it behind a handle with an expiry, in a store every instance can reach. The protocol is stateless. Your backend does not have to be.

A quick test tells you how ready a server is. Pick any request in flight, kill the instance handling it, and retry the same call against a different instance. If the answer is identical, you are stateless where it counts. If the retry fails, asks the client to start over, or returns something subtly different, there is still a ledger hiding in memory, and that is the code to move into a database or behind a handle first.

Server typeBest fitWhy
Read-only data lookupsFully statelessNothing to remember
Database CRUDStateless, record ID as handleThe database already holds the truth
Browser or shell controlStateless protocol, stateful backendLive resource stays behind an expiring handle
Long renders and batch jobsTasks extensionPolling replaces open connections

Generating Images Over MCP

MCP is also how assistants reach creative tools, and image generation happens to be a clean example of the handle pattern. PicassoIA exposes its models through a developer API and an MCP connection. The API lives at https://api.picassoia.com/v1, uses a Bearer token that begins with pia_sk_, and follows a Replicate style layout: POST /v1/models/{owner}/{name}/predictions to start a job, GET /v1/predictions/{id} to check it, and POST /v1/predictions/{id}/cancel to stop it.

Prediction IDs Are Handles Too

Generation is asynchronous. Starting a job returns a prediction ID immediately, and the assistant calls the status tool with that ID after the suggested wait, again and again, until the status reads succeeded or failed. No open connection sits idle while the GPU works, and the ID carries all the continuity. It is the basket_id pattern applied to pixels.

A few limits worth knowing when you plan a workflow: 5 concurrent predictions per account, shared across tokens and MCP connections, prompts up to 4,000 characters, and a 3 hour timeout per job.

Models You Can Call

These four models are available through both the API and the MCP connection:

The photos in this article came from P Image, one of the many text to image models on the platform. When you are writing the tool descriptions and JSON schemas your own MCP server will expose, a large language model saves time. Claude Sonnet 5, GPT 5.6 Sol, and Gemini 3.5 Flash are all available for drafting and reviewing that kind of structured text.

Your Turn: Create Images With Picasso IA

You have seen the whole picture now: sessions out, handles in, any instance can answer any request. The fastest way to feel the pattern is to use it. Open Picasso IA, pick a text to image model, and write a prompt for the scene you wish your own architecture diagram looked like. Try a cozy café, a busy toll plaza, or a quiet row of servers, then change one detail at a time and watch how the result shifts.

When you are ready for more, browse every model on the all models page, turn a favorite image into a short clip with a video model, and keep experimenting. Each prompt you write is a small request that carries everything it needs, which is exactly the point of this spec.

Share this article