Large Language ModelsGenerate imagesGenerate videos

MCP Rate Limiting: How to Add Rate Limits to an MCP Server

A working plan for MCP rate limiting in TypeScript. Build a token bucket, identify callers by token or client id, return 429 and Retry-After over HTTP, weigh tools by cost, scale counters with Redis, and test every limit with fake timers before an agent finds the gaps.

MCP Rate Limiting: How to Add Rate Limits to an MCP Server
Cristian Da Conceicao
Founder of Picasso IA

An AI agent never gets bored. Point one at your MCP server with a vague task and it can fire 40 tool calls in ten seconds, retry every failure instantly, and fan out parallel requests nobody planned for. That is great for productivity and terrible for your invoice. MCP rate limiting is how a Model Context Protocol server stays useful under that kind of pressure: each caller gets a fair budget, expensive tools cost more than cheap ones, and anyone who runs out gets a clear message about when to come back.

This article shows how to add rate limits to an MCP server in TypeScript, from a short token bucket to a Redis backed limiter that works across several instances. You will see where the checks belong, how to return errors an agent can act on, and how to test it all without waiting a real minute.

Why MCP Servers Need Rate Limits

Aerial view of a highway toll plaza with cars queuing in orderly lanes at golden hour

A rate limit is a promise about capacity: this caller may use this much, per unit of time, and no more. The security notes for tools in the Model Context Protocol specification list rate limiting tool invocations as a requirement for servers, right beside input validation and access control. The official SDKs handle transports and schemas, but they leave the limiter to you, so every server author ends up writing one.

Agents Retry Without Getting Tired

A person clicking a button is slow and easy to predict. An agent loop is neither. The model calls a tool, reads the result, and decides what to call next, often within milliseconds. Three patterns show up again and again:

  • Retry storms. A tool fails, the model tries again, fails again, and repeats until its context or budget runs out.
  • Parallel fan-out. Clients can send several tool calls at once, so a single prompt can turn into a dozen simultaneous requests.
  • Runaway loops. A vague task plus a tool that never says "done" produces hundreds of calls from one session.

There is a security angle too. A web page or document the agent reads can hide instructions telling it to call a tool over and over. You cannot always stop the injection, but a rate limit caps the damage.

Tools Wrap Paid APIs

Most MCP tools are thin wrappers around something that costs money or has its own quota: a language model, an image generator, a search API, a database. A limiter protects three things at once:

  1. Your budget, because one noisy session should not burn a day of spend.
  2. Your upstream quotas, because providers answer abuse with 429 responses that hit every user of your server.
  3. Other users' latency, because a greedy client that saturates your workers slows everyone else.

💡 Local stdio servers need limits too. Even when only one person runs the server on a laptop, a looping agent can drain the paid API behind it. A per-tool budget costs nothing to add and prevents a very unpleasant surprise.

Pick the Right Algorithm

Glass hourglass on a walnut desk with golden sand halfway through its fall

Six designs handle nearly every case. Here is how they behave once agent traffic hits them:

AlgorithmBurst behaviorMemory per callerBest for
Fixed windowAllows up to 2x at window edgesOne counterSimple quotas such as daily caps
Sliding window logExact, no edge burstsOne timestamp per requestLow volume, strict limits
Sliding window counterClose to exactTwo countersHigh volume HTTP endpoints
Token bucketControlled bursts, steady refillTwo numbersTool calls from agents
Leaky bucketNo bursts, smooth outputA queueFeeding fragile upstream services
Concurrency limitCaps parallel jobsOne counterLong running tools

Fixed and Sliding Windows

A fixed window counter is the simplest design: count requests per minute and reset at the top of the minute. It is cheap and easy to explain, but a caller can send a full quota at 12:00:59 and another full quota at 12:01:00, doubling the burst your server sees.

A sliding window removes that edge by looking at the trailing 60 seconds from the current moment. It either stores every timestamp (exact, but memory hungry) or weights the previous window's counter (close enough and cheap). Reach for windows when you want plain quotas like 1,000 calls per day, where the exact refill moment does not matter.

Why Token Bucket Usually Wins

Low-angle view of a galvanized bucket under a dripping brass tap with ripples on the water

Agent traffic is bursty: nothing for ten seconds, then six tool calls at once, then silence. A token bucket fits that shape. Each caller owns a bucket with a capacity (the largest burst allowed) and a refill rate (the sustained pace). A call takes tokens out, and time puts them back. A bucket of 60 tokens refilling at one per second allows a burst of 60 calls, then one call per second, which is easy to describe to users as "60 now, 60 per minute after that".

Two properties make it ideal for MCP:

  • Lazy refill. You compute the refill when a request arrives, so there are no timers to manage.
  • Cost support. A video tool can take 20 tokens while a lookup takes one, all from the same budget.

Build the Limiter in TypeScript

Over-the-shoulder view of a developer typing at a tidy wooden desk with a blurred monitor

The limiter below works in any TypeScript MCP server built on @modelcontextprotocol/sdk. It has no dependencies and keeps its state in memory.

The Limiter Class

// rate-limit.ts
export type Decision = {
  allowed: boolean;
  remaining: number;
  retryAfterMs: number;
};

type Bucket = { tokens: number; updatedAt: number };

export class TokenBucket {
  private buckets = new Map<string, Bucket>();

  constructor(
    private readonly capacity: number,
    private readonly refillPerSecond: number,
  ) {}

  take(id: string, cost = 1): Decision {
    if (cost > this.capacity) {
      throw new RangeError(`Cost ${cost} is larger than the bucket (${this.capacity})`);
    }

    const now = Date.now();
    const bucket = this.buckets.get(id) ?? { tokens: this.capacity, updatedAt: now };

    const elapsedSeconds = (now - bucket.updatedAt) / 1000;
    bucket.tokens = Math.min(this.capacity, bucket.tokens + elapsedSeconds * this.refillPerSecond);
    bucket.updatedAt = now;
    this.buckets.set(id, bucket);

    if (bucket.tokens >= cost) {
      bucket.tokens -= cost;
      return { allowed: true, remaining: Math.floor(bucket.tokens), retryAfterMs: 0 };
    }

    const missing = cost - bucket.tokens;
    return {
      allowed: false,
      remaining: 0,
      retryAfterMs: Math.ceil((missing / this.refillPerSecond) * 1000),
    };
  }

  // Drop idle buckets so the map cannot grow forever.
  sweep(maxIdleMs = 10 * 60_000) {
    const cutoff = Date.now() - maxIdleMs;
    for (const [id, bucket] of this.buckets) {
      if (bucket.updatedAt < cutoff) this.buckets.delete(id);
    }
  }
}

export const budget = new TokenBucket(60, 1); // burst of 60, refills one per second
setInterval(() => budget.sweep(), 60_000).unref();

Three details deserve a second look. The refill comes from elapsed time, so there is no setInterval per caller. The cost argument lets one limiter serve cheap and expensive tools. And retryAfterMs says exactly how long until the bucket holds enough tokens, which is the number you show the agent.

Wrap Every Tool Handler

Put the check in one wrapper so no tool can forget it:

import { z } from "zod";
import type { CallToolResult } from "@modelcontextprotocol/sdk/types.js";
import { budget } from "./rate-limit.js";

type Extra = { authInfo?: { clientId?: string }; sessionId?: string };

export function limited<Args>(
  tool: string,
  cost: number,
  handler: (args: Args, extra: Extra) => Promise<CallToolResult>,
) {
  return async (args: Args, extra: Extra): Promise<CallToolResult> => {
    const caller = extra.authInfo?.clientId ?? "anonymous";
    const decision = budget.take(caller, cost);

    if (!decision.allowed) {
      const seconds = Math.ceil(decision.retryAfterMs / 1000);
      return {
        isError: true,
        content: [
          {
            type: "text",
            text: `Rate limit reached for ${tool}. Wait ${seconds} seconds before calling it again.`,
          },
        ],
      };
    }
    return handler(args, extra);
  };
}

server.registerTool(
  "generate_image",
  { description: "Generate an image from a prompt", inputSchema: { prompt: z.string().max(4000) } },
  limited("generate_image", 5, async ({ prompt }) => {
    const url = await createImage(prompt);
    return { content: [{ type: "text", text: url }] };
  }),
);

💡 Return the block as a tool result, not a thrown error. The specification separates protocol errors from tool execution errors. A result with isError: true stays inside the model's context, so the agent reads "wait 12 seconds" and adjusts. A thrown exception becomes a JSON-RPC error that many clients show as a failure and nothing more.

Identify Callers and Guard the Endpoint

Close-up of a hand tapping a transit card on a stainless steel turnstile reader

A limit is only as fair as the way you identify the caller. Get it wrong and you either throttle everyone together or let one client dodge the limit by reconnecting.

Pick the Right Identity

Transport and authIdentity to useWatch out for
stdioOne shared budget per toolOne process serves one client, so there is no caller to tell apart
Streamable HTTP with OAuthThe client or user id from authInfoThe best option, since it survives reconnects
Streamable HTTP with a static bearer tokenA hash of the tokenRotate tokens and hash them before storing
Anonymous HTTPIP addressShared offices and mobile networks look like one caller

Resist the temptation to use the Mcp-Session-Id header as the main identity. The server issues it, and a client can simply initialize a new session to get a fresh bucket. Use sessions as a secondary limit, for example to cap in-flight jobs per session, and keep the budget tied to something that survives reconnects.

Return 429 With Retry-After

Telephoto view of a red railway signal above empty tracks in early morning mist

Put a coarse limit at the HTTP layer and a precise one inside the tools. The HTTP layer is cheap and runs before any JSON is parsed or any session is created, so it shields the server from floods. It cannot tell tools/list from an expensive tools/call without reading the body, so keep it generous and let the tool layer do the accurate accounting.

import { createHash } from "node:crypto";
import type { NextFunction, Request, Response } from "express";
import { TokenBucket } from "./rate-limit.js";

const httpBudget = new TokenBucket(120, 2); // 120 burst, 2 per second sustained

function callerId(req: Request): string {
  const auth = req.header("authorization");
  if (auth) return "tok:" + createHash("sha256").update(auth).digest("hex").slice(0, 16);
  return "ip:" + req.ip;
}

export function limitHttp(req: Request, res: Response, next: NextFunction) {
  const decision = httpBudget.take(callerId(req));
  res.setHeader("RateLimit-Remaining", String(decision.remaining));
  if (decision.allowed) return next();

  res.setHeader("Retry-After", String(Math.ceil(decision.retryAfterMs / 1000)));
  res.status(429).json({
    jsonrpc: "2.0",
    error: { code: -32000, message: "Too many requests. Retry after the delay in Retry-After." },
    id: null,
  });
}

// app.set("trust proxy", 1);
// app.post("/mcp", limitHttp, handleMcp);

Hash the bearer token before using it as an identifier, so raw credentials never sit in a Map or a log line. Send Retry-After in whole seconds, because HTTP clients with retry logic read it. And behind a proxy, set trust proxy, or every anonymous caller shares the proxy's address.

Weigh Tools by What They Cost

Chef's hands plating a dish at a kitchen pass with blank order tickets hanging on a steel rail

Walk into a busy restaurant kitchen and you will see order tickets of very different sizes hanging on the same rail. A side salad and a slow braise do not take the same effort, and a good kitchen does not treat them the same. Tool calls work the same way.

Tool typeExampleToken costExtra guard
Read-only lookuplist_articles, get_article1None
Write or publishsave_article2Idempotency check
Text generationSummaries from a language model3Cap output length
Image generationgenerate_image52 concurrent per caller
Video generationgenerate_image_to_video201 concurrent, spaced submissions

With a bucket of 60 tokens refilling at one per second, a caller can run 60 lookups in a burst, or 12 image generations, or 3 video jobs, and the budget refills fully in a minute.

Cost Weights per Tool

Keep the weights in one place and pass them to the wrapper from the previous section:

export const TOOL_COST = {
  list_articles: 1,
  save_article: 2,
  generate_image: 5,
  generate_image_to_video: 20,
} as const;

// limited("generate_image", TOOL_COST.generate_image, handler)

Start with weights proportional to what each call costs in money or upstream seconds, then adjust them from real traffic.

Cap Jobs and Back Off Upstream

A token budget limits how often a caller starts work. It does not limit how much work runs at the same moment. Long tools need a concurrency cap, and shared upstream queues sometimes need spacing between submissions. Both are short:

const inFlight = new Map<string, number>();

export async function withConcurrency<T>(
  caller: string,
  max: number,
  job: () => Promise<T>,
): Promise<T | "busy"> {
  const current = inFlight.get(caller) ?? 0;
  if (current >= max) return "busy";

  inFlight.set(caller, current + 1);
  try {
    return await job();
  } finally {
    const left = (inFlight.get(caller) ?? 1) - 1;
    if (left <= 0) inFlight.delete(caller);
    else inFlight.set(caller, left);
  }
}

// One submission per slot: concurrent callers queue behind each other.
let nextSlot = 0;
export async function waitForSlot(minGapMs = 30_000) {
  const now = Date.now();
  const start = Math.max(now, nextSlot);
  nextSlot = start + minGapMs;
  await new Promise((resolve) => setTimeout(resolve, start - now));
}

// When the upstream API answers 429 or 5xx, wait and retry with jitter.
export async function fetchWithBackoff(
  send: () => Promise<Response>,
  maxAttempts = 4,
): Promise<Response> {
  for (let attempt = 1; ; attempt++) {
    const response = await send();
    const retryable = response.status === 429 || response.status >= 500;
    if (!retryable || attempt >= maxAttempts) return response;

    const retryAfter = Number(response.headers.get("retry-after"));
    const delayMs = retryAfter > 0 ? retryAfter * 1000 : 2 ** attempt * 500;
    await new Promise((resolve) => setTimeout(resolve, delayMs + Math.random() * 250));
  }
}

Return "busy" to the agent as a normal tool error that says a job is already running and suggests polling its status. The jitter in fetchWithBackoff matters: without it, every instance that hit a 429 wakes up at the same instant and hits the upstream again together.

💡 Real limits, published. PicassoIA's own API documents 5 concurrent predictions per account, shared across API tokens and MCP connections, plus 4,000 character prompts, 10 MB request bodies and a 3 hour timeout. Its image and video tools also hand back a predict_id and a next_poll_in_seconds hint, so the client never has to guess how often to poll. Copy that idea: a limit with a published number and a polling hint is a limit agents can obey.

Scale Past One Process

Wide view down an aisle of black server racks in a data center

The in-memory limiter has one flaw: its memory belongs to one process. Run three instances behind a load balancer and each keeps its own counters, so a caller effectively gets three times the limit. Serverless platforms are worse, since every cold start begins with empty buckets. The fix is to move counters into a shared store, and Redis is the usual choice because its operations are atomic and fast.

Shared Counters With Redis

import Redis from "ioredis";
import { RateLimiterRedis, RateLimiterRes } from "rate-limiter-flexible";
import type { Decision } from "./rate-limit.js";

const redis = new Redis(process.env.REDIS_URL!);
const FAIL_OPEN = process.env.RATE_LIMIT_FAIL_OPEN === "true";

const limiter = new RateLimiterRedis({
  storeClient: redis,
  points: 60,   // budget per window
  duration: 60, // window length in seconds
});

export async function takeShared(caller: string, cost: number): Promise<Decision> {
  try {
    const res = await limiter.consume(caller, cost);
    return { allowed: true, remaining: res.remainingPoints, retryAfterMs: 0 };
  } catch (rejection) {
    if (rejection instanceof RateLimiterRes) {
      return { allowed: false, remaining: 0, retryAfterMs: rejection.msBeforeNext };
    }
    // Redis itself failed, so apply the configured failure policy.
    return { allowed: FAIL_OPEN, remaining: 0, retryAfterMs: 5_000 };
  }
}

This library counts in fixed windows, so the edge burst from the algorithm table applies. For most MCP servers that trade is fine. If you need a true token bucket across instances, store the two numbers in a Redis hash and run the refill and the take in a single Lua script, because a read followed by a write from two instances is a race.

Fail Open or Fail Closed

Redis will go down eventually, and your limiter has to pick a side:

  • Fail closed for tools that cost money. A short outage is cheaper than an unlimited video queue.
  • Fail open for cheap reads, where blocking everyone would hurt more than a burst of lookups.
  • Degrade, not disable. The library supports an in-memory insuranceLimiter as a fallback. Limits then apply per instance, which is still far better than none.

Whatever you choose, log every limiter failure loudly, because a silent fail open is how a limit quietly stops existing.

Test and Watch the Limits

Macro view of a brass pressure gauge with its needle resting in the green zone

A limiter that has never rejected anything in a test is a limiter you cannot trust.

Test With Fake Timers

Fake timers let a unit test move through a full minute in a single millisecond:

import { afterEach, beforeEach, describe, expect, it, vi } from "vitest";
import { TokenBucket } from "./rate-limit";

describe("TokenBucket", () => {
  beforeEach(() => vi.useFakeTimers());
  afterEach(() => vi.useRealTimers());

  it("allows a burst, then blocks with a wait time", () => {
    const bucket = new TokenBucket(3, 1);
    for (let i = 0; i < 3; i++) expect(bucket.take("a").allowed).toBe(true);

    const blocked = bucket.take("a");
    expect(blocked.allowed).toBe(false);
    expect(blocked.retryAfterMs).toBeGreaterThan(0);
  });

  it("refills as time passes", () => {
    const bucket = new TokenBucket(1, 1);
    bucket.take("a");
    expect(bucket.take("a").allowed).toBe(false);

    vi.advanceTimersByTime(1000);
    expect(bucket.take("a").allowed).toBe(true);
  });

  it("keeps callers apart", () => {
    const bucket = new TokenBucket(1, 1);
    bucket.take("a");
    expect(bucket.take("b").allowed).toBe(true);
  });
});

After the unit tests, run the real server under the MCP Inspector (npx @modelcontextprotocol/inspector) and call one tool in a loop. The first calls should succeed and the rest should return the rate limit message with a wait time. To test the HTTP layer, send a burst with curl and count the status codes:

for i in $(seq 1 150); do
  curl -s -o /dev/null -w "%{http_code}\n" -X POST http://localhost:3000/mcp \
    -H "Content-Type: application/json" \
    -H "Accept: application/json, text/event-stream" \
    -d '{"jsonrpc":"2.0","id":1,"method":"ping"}'
done | sort | uniq -c

The early responses depend on your session setup, but once the bucket is empty every response should be 429.

Metrics Worth Watching

Count every rejection with the tool name and a caller bucket (never the raw identity). Those few numbers show whether the limits are fair:

MetricWhat it tells youAlert when
Rejection rate per toolLimits too tight, or one noisy callerAbove 5% for 10 minutes
Top callers by rejectionsA looping or abusive agentOne caller owns over half of them
Retry-After, 95th percentileHow long agents are really waitingAbove 60 seconds
In-flight jobsConcurrency saturationAt the cap for 5 minutes
Limiter errorsRedis healthAny

Mistakes that show up in real servers:

  • Using the session id as the only identity, so a reconnect resets the limit.
  • Sharing one global bucket, so one heavy caller starves everyone.
  • Dropping or stalling requests silently instead of returning an error with a wait time.
  • Setting limits once and never reading the rejection numbers again.

Try It on PicassoIA

The limiter above is under a hundred lines, which makes it a good job for a language model: tightly specified, easy to test, and quick to review. PicassoIA puts several coding capable models in one place, so you can draft, compare and fix without juggling accounts.

How to Use Claude Sonnet 5

  1. Open the model. Go to Claude Sonnet 5 on PicassoIA.
  2. Paste a specific prompt. Name the SDK, the transport, the algorithm, the budget, the cost of each tool and the exact error text. For example:
Write a TypeScript token bucket limiter for an MCP server built on
@modelcontextprotocol/sdk. Budget: 60 tokens, refill 1 per second.
Costs: list_articles 1, generate_image 5, generate_image_to_video 20.
Identify callers by authInfo.clientId, fall back to "anonymous".
When blocked, return isError: true with the wait time in seconds.
Include vitest tests that use fake timers.
  1. Ask for tests first. Reading the tests shows what behavior the model assumed before you read a line of implementation.
  2. Run them and feed failures back. Paste the exact error output into the same conversation and ask for a fix.
  3. Request a review. Finish with "Review this limiter for race conditions and memory growth" and read the answer critically.

Keep each request to one concern, and give the model your real numbers instead of "reasonable defaults". Other models are worth a second opinion:

ModelUse it for
GPT 5.6 SolChecking tricky concurrency and Lua script logic
Kimi K2.6Drafting the agent side retry and backoff code
Gemini 3.5 FlashFast iterations and sample test data

Once your server is protected, put it to work. Every photograph in this article started as a plain text prompt written for P-Image, and you can run the same prompt on Flux 2 Pro to compare results. Open Picasso IA, describe a scene in a sentence or two, and generate your first image. Change the lens, the light or the angle, run it again, and then animate your favorite result into a short video. The fastest way to find out what the platform can do is to experiment with your own ideas.

Share this article