Plug in one more MCP server and your agent can suddenly forget how to read a file. No crash, no red banner, just a tool that was there yesterday and is missing today. That is the MCP tool limit at work, and every editor handles it differently. Cursor warns you past 40 tools, VS Code rejects requests past 128, and Claude Code quietly defers most tool definitions until the model asks for them.
This article gives you the real numbers for each client, the messages you will see when you cross the line, the token math behind the limits, and a short list of fixes that keep a busy setup under the ceiling.
๐ก Quick take: plan around 40 tools in Cursor, 128 in VS Code and 100 in Windsurf. Treat Claude Code as "no fixed count, but every definition still costs context unless tool search is on." Limits shift between releases, so confirm against your installed version.
The Short Answer
Here is what the official docs and community threads report as of October 2026.
| Client | Tool limit | What happens past it |
|---|
| Cursor | 40 tools across all enabled MCP servers | A warning appears and only the first 40 tools reach the agent |
| VS Code (Copilot Chat) | 128 tools per chat request | Error: "Cannot have more than 128 tools per request" |
| Windsurf (Cascade) | 100 tools across all servers | Extra tools are unavailable until you toggle some off |
| Claude Code | No fixed count | Tool search defers definitions by default; turning it off loads everything upfront |
| Claude Desktop | No published hard cap | Every enabled tool definition goes into the conversation context |
Two things stand out. First, VS Code is the only client in this table whose number sits in official documentation. Cursor's 40 and Windsurf's 100 come from forum threads and third-party write-ups, which is why they deserve a "check your version" note. Second, the hard cap is rarely your first problem. Slower replies, wrong tool picks and a context window that fills up before you type a prompt show up long before any cap does.
Every Tool Costs Context

An MCP server describes each tool with a name, a plain language description and a JSON schema for its arguments. The client sends that whole bundle to the model with every request, before your prompt and before any file contents. Ten small tools barely register. Eighty verbose ones can eat a real slice of the window you wanted for code.
That is the actual reason clients set ceilings. A cap is a blunt way to protect the context budget, keep latency down and stop the model from drowning in options. Cursor forum threads describe the 40 tool cap in exactly those terms: forwarding dozens of raw tools makes the model burn context judging them, so latency rises and precision drops.
Models Pick Worse With More Choices

Context is only half the story. The other half is choice. When a model faces dozens of tools with overlapping names, like search_issues, search_repos and search_code, it has to guess which one fits, and the guesses go wrong more often as the list grows.
OpenAI's function calling documentation suggests aiming for fewer than 20 functions available at the start of a turn, and it labels that a soft suggestion. GitHub reached a similar result inside Copilot: it trimmed the default built-in tool set from 40 to 13 core tools and expanded the rest on demand. In GitHub's own benchmarks, that approach improved success rates by 2 to 5 percentage points and cut average latency by 400 milliseconds.
So the practical limit sits below the printed cap. A server with 12 sharply named tools often beats a server with 90 tools that all sound alike.
Test Selection With a Chat Model
You can measure tool confusion yourself in about ten minutes, without touching your editor config.
- Export the tool list (names plus one line descriptions) from each server with the MCP Inspector and paste all of it into a single chat.
- Write ten realistic tasks, such as "find open bugs assigned to me" or "resize this screenshot."
- Ask the model to name the one tool it would call for each task, then count the wrong picks.
- Delete half the tools and run it again.
Repeat the same experiment in a few models side by side. Claude Sonnet 5, GPT 5.4 and Gemini 3.1 Pro are all available as chat models on PicassoIA, so you can compare how each one handles a crowded list. If the wrong picks drop sharply when the list shrinks, you have your answer on how many tools your workflow can really carry.
๐ก Caveat: this tests the model's choice from plain text. Real clients add their own routing, so treat the result as a signal, not a guarantee.
What You See at 41 Tools

Cursor shows a warning when the tools from your enabled servers pass 40. The wording reported on the Cursor forum runs along the lines of: you have 50 tools from enabled servers, the limit is 40, and some tools may not be available to the agent. In older releases the extras were simply ignored. Connect servers that expose 45 tools in total and the agent used 40, with no detailed diagnostics about which five vanished.
Some recent write-ups say newer Cursor builds register tools more flexibly and load some on demand. The MCP documentation page we checked does not publish a tool count, so 40 is still the safe planning number.
๐ก Watch the math: the cap is total, not per server. Three servers with 15 tools each already put you at 45, five over the line, even though no single server looks big.
Cursor Workarounds That Work
Four moves, ordered from least to most effort:
- Toggle servers per task. Cursor's docs describe switching a server on or off from the Customize sidebar without removing it. Keep the database server off while you write front end code.
- Split your config. Put project specific servers in
.cursor/mcp.json and leave only your everyday servers in the global ~/.cursor/mcp.json.
- Choose lean servers. Count a server's tools before you install it. Six focused tools are a better deal than sixty broad ones.
- Route through a proxy. Forum users mention aggregator proxies that expose a handful of meta tools and forward calls to hundreds behind them. It works, but you add one more moving part to debug.
Per tool toggles are a repeated feature request on Cursor's forum, so if you need that granularity today, you will be pruning at the server level.
VS Code and the 128 Ceiling
The "More Than 128" Error

VS Code is the most explicit of the three. Its documentation states that a chat request can have a maximum of 128 tools enabled at a time. Cross it and Copilot Chat fails with: "Cannot have more than 128 tools per request."
The count applies to everything enabled in the tools picker, not just MCP tools, so it fills up faster than you would expect. Two or three large servers can be enough.
The quick fix lives in the tools picker of the Chat view. Uncheck individual tools or whole MCP servers until the count drops under 128, then send the request again.
Virtual Tools Beat Manual Pruning
If you want to go past 128, VS Code offers virtual tools. Set github.copilot.chat.virtualTools.threshold and VS Code groups similar tools under a single virtual tool. The model sees the groups first and expands a group only when the prompt needs it, which lets a chat request go beyond the 128 limit.
Think of it as a table of contents: the model reads chapter titles, then opens one chapter. The trade is an extra expansion step the first time a group is needed, in exchange for a much shorter starting list.
GitHub's tests of embedding based routing show why grouping pays off. The needed tool was ready in 94.5 percent of cases, against 87.5 percent for LLM-based selection and 69.0 percent for a static tool list.
Claude Code Defers Tools by Default

Claude Code takes a different route. Instead of capping the count, it avoids paying for every definition up front. With tool search, MCP tool definitions are withheld from the request and the model gets a single search tool to load the ones it needs. Anthropic's MCP docs list the exceptions: tool search is off with a custom ANTHROPIC_BASE_URL, with ENABLE_TOOL_SEARCH=false, and with models older than the Claude 4.5 generation.
Community write-ups describe more values for the same variable, such as auto, which defers only when definitions pass 10 percent of the context window, and auto:5 for a 5 percent threshold. Check the docs for your installed version before relying on those. One write-up measured 147 tool definitions at roughly 90,000 tokens when loaded upfront, against about 15,000 with tool search on.
Output size matters too. Claude Code warns when a single MCP tool result passes 10,000 tokens and caps output at 25,000 by default. Raise it with MAX_MCP_OUTPUT_TOKENS only when a job truly needs it.
What Claude Desktop Does Instead
We could not find a fixed tool count for Claude Desktop in Anthropic's documentation. Third party write-ups describe a simpler mechanism: every tool exposed by every connected server is injected into the conversation context for every message, with no dynamic retrieval. Their practical advice is to stay between 30 and 50 active tools before saturation and confusion set in. Treat that range as community experience, not an official limit.
The takeaway for Claude users is that the limit is a budget, not a number. Enable only the connectors the current conversation needs, and switch the rest off.
The Real Cost in Tokens
Measured Costs for Real Servers

Published numbers vary widely because schema detail varies. These figures come from public token cost listings and the write-up mentioned above:
| Server or setup | Tool definitions | Upfront tokens | Per tool (approx.) |
|---|
| GitHub MCP server | 86 | 14,406 | about 170 |
| Smaller GitHub server variant | 36 | 11,430 | about 320 |
| GitHub Projects server | 19 | 2,316 | about 120 |
| Mixed setup from one write-up | 147 | about 90,000 | about 610 |
The spread is the lesson. A tool costs anywhere from roughly 120 to 610 tokens depending on how much description and schema it carries. Count tools, but weigh them too.
A Quick Budget Formula
Use this to sanity check any setup:
tools ร tokens per tool รท context window = share spent before you type
| Tools | Tokens each | Total | Share of a 200,000 window |
|---|
| 10 | 300 | 3,000 | 1.5% |
| 40 | 300 | 12,000 | 6% |
| 86 | 170 | 14,620 | 7.3% |
| 128 | 600 | 76,800 | 38.4% |
The 200,000 window is only an example, and your model may offer more or less. The pattern holds anyway: 40 lean tools cost about 6 percent, while 128 heavy ones can eat more than a third. The weight of the schemas matters more than the number printed on the cap.
Fixes That Keep You Under
Group Tools by Job

Sort every tool into three or four jobs, for example code, data, docs and media. Then enable one job at a time. In Cursor that means per project config, in VS Code it means the tools picker, and in Claude it means the connector toggles.
Grouping also fixes naming collisions. When the docs tools and the code tools never share a conversation, search can only mean one thing.
Trim Servers Before Tools
Most setups carry dead weight. Run through this list once:
- Count the tools each server exposes with the MCP Inspector.
- Rank servers by how often you actually call them this month.
- Remove any server you have not touched in 30 days.
- Merge duplicates, such as two servers that both offer file search.
- Pick toolsets where the server allows it. Some servers, GitHub's among them, let you choose which groups of tools to load at startup.
๐ก Rule of thumb: if a tool has not been called in a month, it is costing tokens and earning nothing.
Build Small Servers Instead

If you write your own MCP server, design for the budget: one job, short descriptions, names that never overlap, and tight argument schemas. The PicassoIA connector for Claude is a good picture of the lean approach. It exposes nine tools: generate_image, edit_image, two video generators, get_generation, list_generations, list_models, get_account and cancel_generation. All of them serve one job, making images and videos.
A server that size sits comfortably under Cursor's 40 and Windsurf's 100, and it leaves room for half a dozen other servers before you feel any pressure.
Try It With Picasso IA

The easiest way to feel what a tidy tool budget gives you is to run one real workflow through it. Ask for an image in plain language, check the result, adjust one detail, ask again. A focused setup keeps that loop fast because the model is not wading through sixty unrelated tools to find the one it needs.
The photos in this article came from P-Image and short, specific prompts. You can run the same kind of prompts on Picasso IA right now:
Picasso IA also turns a still image into a short video, so the same photo can become a clip with one more step. Browse every model at picassoia.com/en/all-models, pick one, write a prompt about your own project, and see what you get. Change one word, run it again, and compare. That small habit of testing one variable at a time is the same habit that keeps an MCP setup healthy.