Large Language ModelsGenerate imagesGenerate videos

MCP Sampling Deprecated: Sampling vs Elicitation and What to Use Now

MCP sampling was deprecated on 2026-07-28 (SEP-2577), but elicitation was not. This article compares the two, explains the new Multi Round-Trip Requests flow, and shows what to use instead of sampling, with a dated migration checklist for servers.

MCP Sampling Deprecated: Sampling vs Elicitation and What to Use Now
Cristian Da Conceicao
Founder of Picasso IA

If your MCP server calls sampling/createMessage, the specification now tells you to stop building on it. On 2026-07-28 the Model Context Protocol deprecated Sampling (SEP-2577), together with Roots and Logging. Elicitation, the feature most people lump in with it, was not deprecated. It was rebuilt on a new delivery mechanism instead. This article separates what is going away from what stays, with the dates, field names and replacement paths you need before your next release.

💡 Short version: Sampling is deprecated, and the spec says to integrate directly with LLM provider APIs. Elicitation is active and now travels through Multi Round-Trip Requests. Nothing can be removed before a revision released on or after 2027-07-28.

A developer at a bright whiteboard filled with hand-drawn boxes and arrows, planning an MCP protocol change

What Changed on 2026-07-28

The 2026-07-28 revision rewrites how clients and servers talk, and sampling's deprecation is one entry in a long changelog. Three changes matter here:

  • Stateless core. The initialize handshake and the Mcp-Session-Id header are gone. Every request carries its protocol version and client capabilities in _meta.
  • Multi Round-Trip Requests (MRTR). A server can no longer push sampling/createMessage or elicitation/create down an open connection. It returns an interim result, and the client retries the original call with the answer attached.
  • A deprecation registry. A new lifecycle policy defines Active, Deprecated and Removed states, with a minimum twelve-month window before anything can be removed.

The first change explains the second. Without a persistent session there is no channel to push a request into, so server-to-client requests had to be redesigned. Sampling was redesigned and deprecated in the same release. Elicitation was only redesigned.

What Is Deprecated at a Glance

FeatureStatusReplacementEarliest removal
SamplingDeprecatedCall LLM provider APIs directlyFirst revision on or after 2027-07-28
RootsDeprecatedTool parameters, resource URIs or server configurationFirst revision on or after 2027-07-28
LoggingDeprecatedstderr for stdio, OpenTelemetry for observabilityFirst revision on or after 2027-07-28
includeContext values "thisServer" and "allServers"DeprecatedOmit the field or use "none"No later than Sampling
Dynamic Client RegistrationDeprecatedClient ID Metadata DocumentsFirst revision on or after 2027-07-28
ElicitationActiveStays, now delivered through MRTRNot scheduled

What Deprecated Means Here

"Deprecated" is not "removed." During the window, wire-level behavior is unchanged, capability negotiation still works and existing implementations keep running. New implementations should not adopt the feature, and existing ones should migrate. Removal is a Core Maintainer decision made during release preparation, so it can land later than the earliest date. The SEP also asks implementations to emit a warning whenever a deprecated capability is negotiated, which is why SDKs have started printing deprecation notices. Treat those warnings as your to-do list.

⚠️ Gotcha: The Python SDK deprecation page says old session-style calls such as ctx.session.create_message() still work on sessions negotiated at 2025-11-25 or earlier. On a 2026-07-28 connection they warn and then raise, because there is no back channel left to send on.

Why Sampling Got Deprecated

SEP-2577 gives three reasons. They stack.

A weathered brass padlock on a rusted latch of an old oak gate with raindrops on the metal

Too Heavy for Most Clients

Sampling lets a server ask the client's model for a generation. Doing that correctly means human-in-the-loop approval, model selection logic, security handling and, since SEP-1577, a tool loop. That is a lot to build for a feature a server may never call. The SEP points out that few clients support sampling even though it has been in the spec since the November 2024 revision.

A Large Attack Surface

The SEP calls sampling the most security-sensitive of the three deprecated features. A server that can get the client's model to run its prompts creates room for prompt injection and data exfiltration, so every client has to get review screens and rate limits right.

Direct APIs Give More Control

A server that needs a model can call a provider itself. It picks the model, sets the parameters and streams the output. The old pitch for sampling was that servers needed no credentials of their own. The trade-off now is plain: you hold the credentials, you pay the bill, and you decide what happens to the data. Budget for retries, rate limits and a fallback model, because those are now your operations rather than the client's.

Elicitation Was Not Deprecated

A woman at a sunlit kitchen table about to tap a confirm button on a tablet showing a simple form

Check the evidence in the spec itself. The deprecated features registry lists Roots, Sampling, Logging, Dynamic Client Registration, the includeContext values and HTTP+SSE. Elicitation is not on it. The elicitation page carries no deprecation warning, and the changelog files elicitation under MRTR. Some write-ups group all the server-to-client features together, so check the registry before repeating that claim.

Elicitation did lose two details: the out-of-band finish notification and the elicitationId field in URL mode. A server that must match a retry to an earlier request now encodes its own identifier inside requestState.

Form Mode

Form mode collects structured data in band, so the client sees the answer. The requestedSchema is a flat object with primitive properties only:

  • strings, with email, uri, date or date-time formats
  • numbers and integers
  • booleans
  • single-select and multi-select enums

Nested objects and arrays of objects are intentionally unsupported. The user answers with one of three actions: accept with content, decline, or cancel. A server has to handle all three, including offering alternatives on a decline.

⚠️ Servers must not ask for passwords, API tokens, access tokens or payment details in form mode.

URL Mode

URL mode sends the user to a page out of band, and the data never passes through the client. It is the route for sensitive handoffs and third-party OAuth. The client must show the full URL, highlight the domain, ask for consent and never pre-fetch the page. An accept response only means the user agreed to open it. It does not mean the interaction finished.

My read on why it survived: sampling lets a server skip work it could do itself by calling a provider. Elicitation is the only route to the person. Confirmation gates before deleting, publishing or paying have no provider API to fall back on, so removing the feature would leave servers guessing.

Sampling vs Elicitation Side by Side

An aerial view of a forest trail splitting in two at a fork with blank wooden signposts

QuestionSamplingElicitation
Who answers?The client's model, with a human able to denyThe human user
Methodsampling/createMessageelicitation/create
Status on 2026-07-28DeprecatedActive
Result shapeA model message, possibly with tool_use blocksaccept, decline or cancel, plus content
Delivered throughinputRequests in an InputRequiredResultinputRequests in an InputRequiredResult
Best forServer-side text generationDecisions, missing inputs, sensitive handoffs
ReplacementProvider API, or let the host model reasonNone needed

The two are often called substitutes, and they are not. Sampling delegates text generation. Elicitation gets an answer from a person. Replacing a "summarize this page" sampling call with a form would be the wrong fix, since the user never wanted to write that summary. Going the other way is just as wrong: using a model's guess where a human decision is needed.

Take an image server as a worked case. Before it spends a generation, it needs a style and a go-ahead. Both are human choices, so it elicits. It also wants a richer prompt than the user typed. Rewriting is text generation, which used to be a sampling call. Now the server calls a provider itself, or the tool description tells the host model to send a fuller prompt in the first place.

In this revision both still ride the same mechanism, because each shows up as an entry inside inputRequests. The difference is who reads the entry and what the spec says about its lifespan.

How the New Round Trip Works

Two hands passing a cream envelope sealed with red wax across a walnut counter

The flow has four steps. The client calls tools/call. The server returns an InputRequiredResult with resultType: "input_required". The client gathers the answer and retries the original call with inputResponses attached. The server then produces the real result. Because the retry carries everything the server needs, any instance behind a load balancer can handle it.

Here is the server's interim response for an image tool that needs an aspect ratio:

{
  "jsonrpc": "2.0",
  "id": 1,
  "result": {
    "resultType": "input_required",
    "inputRequests": {
      "aspect": {
        "method": "elicitation/create",
        "params": {
          "mode": "form",
          "message": "Which aspect ratio should the image use?",
          "requestedSchema": {
            "type": "object",
            "properties": {
              "ratio": { "type": "string", "enum": ["16:9", "1:1", "9:16"] }
            },
            "required": ["ratio"]
          }
        }
      }
    },
    "requestState": "<signed blob>"
  }
}

And the client's retry, with the answer and the echoed state:

{
  "jsonrpc": "2.0",
  "id": 2,
  "method": "tools/call",
  "params": {
    "name": "generate_image",
    "arguments": { "prompt": "A foggy harbor at dawn" },
    "inputResponses": {
      "aspect": { "action": "accept", "content": { "ratio": "16:9" } }
    },
    "requestState": "<signed blob>"
  }
}

The required _meta fields (protocol version, client info, client capabilities) are left out here for brevity. Note that the JSON-RPC id changes between the original request and the retry, because they are independent requests.

Rules for Clients

  • Echo requestState back exactly. Never inspect, parse or modify it.
  • If the result has no requestState, do not send one on the retry.
  • Use a new JSON-RPC id for the retry.
  • If there are no inputRequests, the client may retry immediately.
  • The fields affect only the retry of that one request, never parallel calls.

Rules for Servers

Treat requestState as attacker-controlled input, because it passes through the client. When it influences authorization, resource access or business logic, protect its integrity with an HMAC or AEAD and reject anything that fails verification. Put the authenticated user, a short expiry and an identifier for the original request inside the protected payload. These limit replay but do not make a state single-use, so enforce one-time redemptions on the server.

Two more constraints apply. Every InputRequiredResult needs at least one of inputRequests or requestState, and a server must not send a request type the client never declared support for. MRTR works only on prompts/get, resources/read and tools/call.

What to Use Instead of Sampling

A hand plugging a blue ethernet cable straight into a router port in a tidy network closet

The spec's answer is one line: integrate directly with LLM provider APIs. In practice, the right replacement depends on why you called sampling.

What the server wantedUse now
Rewrite, summarize or classify textA provider API call, or return raw data to the host model
Ask the user to choose or confirmElicitation in form mode
Hand off a secret or run OAuthElicitation in URL mode
Read workspace paths (Roots)Tool arguments or resource URIs
Send log lines to the clientstderr or OpenTelemetry

Call a Provider API Directly

When you need real text generation, call the provider from the server. Compare candidates before you commit one to your code. On PicassoIA you can try Claude Sonnet 5, GPT 5.6 Terra and Gemini 3.5 Flash side by side and see which one handles your prompts best at the speed you need. Then store the credential as a normal server secret, set a timeout and log token use.

Let the Host Model Reason

Many sampling calls were never needed. A tool result already flows into the model that called the tool. Return the raw page, a clear description and sensible structure, and the host model does the summarizing with no extra round trip and no extra credentials. A documentation server that once used sampling to summarize a long changelog can return the changelog in sections and let the calling model pick what matters. This is my own recommendation, not spec text, but it removes the most calls.

Ask the Human for Decisions

When the missing piece is a decision, switch to elicitation. In the Python SDK, one write-up shows the pattern with an affirmative choice:

CONFIRM = ["cancel", "confirm"]
result = await ctx.elicit("Delete 42 drafts?", response_type=CONFIRM)

The same write-up warns that passing response_type=None sends an empty schema, so an automatic accept looks identical to human approval on the wire. Give the user a real value to choose. Check your SDK's current signature before copying.

💡 Rule of thumb: if a person would have to decide, elicit. If a model would have to write, call a provider or let the host model do it.

A Migration Checklist for Servers

A technician holding a clipboard with a checklist in a server room aisle

  1. Find every call. Search for sampling/createMessage, create_message and any sampling capability checks.
  2. Sort each call by purpose. Text generation goes to a provider API or the host model. Decisions go to elicitation. Context goes to tool arguments or resource URIs.
  3. Drop includeContext values. Remove "thisServer" and "allServers". Omit the field or use "none".
  4. Replace Roots and Logging. Pass paths as tool parameters. Log to stderr on stdio and use OpenTelemetry elsewhere.
  5. Return InputRequiredResult. Swap server-initiated requests for interim results and a requestState you sign.
  6. Test on a 2026-07-28 connection. Watch for the SDK's deprecation warnings and for old session calls that raise.
  7. Keep a legacy path. Clients on 2025-11-25 or earlier still use the old behavior during the window.

Timeline to Track

DateWhat happens
2024-11Sampling enters the spec
2025-11-25includeContext values soft-deprecated, URL mode elicitation introduced
2026-07-28Sampling, Roots and Logging deprecated, MRTR introduced
2027-07-28Earliest revision in which Sampling can be removed

Three Mistakes to Avoid

  1. Treating deprecated as removed. Older sessions still work, but 2026-07-28 connections reject the old session-style calls. Know which versions you serve.
  2. Using a form for text a model should write. Elicitation asks a person for input. It is not a cheaper way to draft copy.
  3. Skipping requestState integrity. Unsigned state is an open invitation to tamper with your server logic.

Review Rules That Stay

Two engineers reviewing printed pages with a highlighter at a long maple table

Human oversight is the part of sampling that carries over. Elicitation clients must show which server is asking, offer clear decline and cancel options, and let people review form answers before sending. In URL mode they must show the target domain and get consent before opening anything. Servers must bind each elicitation to the user's identity and verify who opens a URL, which blocks the phishing pattern where an attacker sends their own link to a victim.

Make Your Own Images With Picasso IA

A creator's desk seen from above with printed mountain lake photographs and a tablet

Protocol work gets easier when you can see what you are building. If you are writing an MCP server that generates pictures, the elicitation pattern above fits it well: ask for an aspect ratio or a style before spending a generation.

On PicassoIA you can try the models behind that kind of tool: PicassoIA Image for text to image, PicassoIA Image Editor Pro for edits, and GPT Image 2 when you want another look. PicassoIA also offers a Replicate-style developer API at https://api.picassoia.com/v1 and MCP connections, so the same generators can sit behind your own tools. Check the current limits on its API page before you plan around them.

Open Picasso IA, write one prompt, and generate your first image. Then run it again with a different aspect ratio and compare the two. That small loop is the same ask, answer and retry pattern this article described, with pixels at the end.

Share this article