Large Language ModelsGenerate imagesGenerate videos
Reduce MCP Token Usage in Claude Code: Fix the Context Window Problem
MCP servers can eat tens of thousands of tokens before you type a single prompt. This article shows how to measure that cost with /context, remove idle servers, turn on tool search, cap tool output, and use subagents so Claude Code keeps its context window for your actual work.
You open Claude Code, type one short prompt, and the context meter already looks half full. Nothing is wrong with your prompt. The culprit is usually the pile of MCP servers you connected last month and forgot about. Every server hands Claude a list of tool definitions the moment a session starts, and every definition is paid for in tokens before any real work begins. Add a few bloated tool responses on top and the context window is gone long before the task is done. The good news is that this is one of the most fixable problems in the whole workflow. This article walks through the order that works: measure first, then prune servers, turn on tool search, cap output, and reset sessions at the right moment.
If you recognize the symptoms, you are in good company. Sessions that feel sluggish, answers that forget instructions from ten minutes ago, and an auto-compact that fires in the middle of a refactor all point to the same thing: too much fixed overhead and too little room for the work itself.
Why MCP Servers Eat Your Context
The Model Context Protocol (MCP) lets Claude Code talk to databases, browsers, issue trackers, design tools and image generators. Each connection is genuinely useful. The price is that Claude has to know what every tool does, so each tool's name, description and full JSON schema gets loaded into the prompt.
Where the Tokens Actually Go
MCP overhead comes from four places, and only some of them are obvious.
Source
When it loads
Typical impact
Who controls it
Tool definitions
Session start
Hundreds of tokens per tool, thousands per server
You, by choosing servers
Tool responses
Every call
A few hundred to tens of thousands of tokens
The server and your output limit
Memory files like CLAUDE.md
Session start
Grows with every rule you add
You
Conversation history
Builds up all session
Grows with every turn
Compacting and clearing
Anthropic's engineering write-up on tool search describes a setup with 58 tools across five servers that consumed roughly 55K tokens before the conversation even started. That works out to about 950 tokens per tool. A server with 35 tools, such as a full-featured code hosting integration, can single-handedly eat a double-digit slice of a standard window.
💡 Quick math: if your tools average 900 tokens each and you connect 40 of them, you have spent about 36,000 tokens on a menu Claude may never order from.
The Real Cost of Idle Tools
Idle tools hurt in two ways. The first is plain space: a window that starts 30% full has 30% less room for code, logs and reasoning. Prompt caching softens the price of that fixed overhead on repeat turns, but it does nothing for the room it occupies.
The second is decision noise. When five servers expose overlapping search or fetch tools, the model has more near-duplicates to choose between. Anthropic's own testing of on-demand tool loading found that exposing fewer tools up front also improved how reliably the right one got picked. Fewer tools is not only cheaper, it is often more accurate.
Measure the Damage First
Do not start deleting servers on a hunch. Measure, change one thing, measure again.
Run /context Before Touching Anything
Open a fresh session in your project and run /context before sending any prompt. Claude Code prints a breakdown of the window by category: system prompt, system tools, MCP tools, memory files, messages and free space. Because you have not typed anything yet, the messages line sits near zero and everything else is pure overhead.
Then run /mcp to list connected servers and their status. From a regular terminal, claude mcp list gives you the same inventory. Write down the MCP tools number from /context. That is your baseline.
Read the Numbers Like a Budget
There is no official threshold, but this rule of thumb works well in practice:
MCP tools share of window
Verdict
What to do
Under 5%
Healthy
Leave it alone
5% to 15%
Worth a look
Remove servers you touch less than weekly
Over 15%
A real problem
Prune hard and enable tool search
Check the memory files line too. A CLAUDE.md that has grown into a wall of rules costs tokens on every single session, just like a chatty MCP server does.
Cut and Scope Your Servers
The cheapest token is the one never loaded. Server hygiene beats every clever setting.
Disable First, Delete Later
Open the /mcp menu and switch off anything you do not need for today's task. Depending on your version, you can toggle a server there without losing its configuration. When you are sure a server is dead weight, remove it for good with claude mcp remove <name>.
Ask three questions about every server:
Did I call a tool from it in the last week?
Does a command line tool already do the same job?
Does it duplicate tools another server already offers?
One honest "no" is enough reason to disable it.
Scope Servers to Projects
MCP servers can be added at three scopes: local, project and user. A project-scoped server lives in a .mcp.json file at the repo root, so it only loads where it matters:
claude mcp add --scope project my-db -- npx my-db-mcp-server
Keep your user scope almost empty. Put the database server in the repo that has a database, and the design server in the repo that has designs. For one-off sessions, you can also start Claude Code with only the servers named in a config file:
claude --strict-mcp-config --mcp-config ./mcp/docs-only.json
That session ignores every other configured server, which makes it ideal for a focused task or a clean before and after comparison.
Swap Servers for Plain CLIs
If a tool already exists as a command line program, such as git, gh, docker, psql or aws, Claude can run it through the shell. That costs zero tool definition tokens, because the model already knows how those commands work.
Anthropic pushed this idea further in its post on code execution with MCP. Presenting servers as code APIs that the agent calls from a script, instead of as individual tools, cut one example workflow from about 150,000 tokens to about 2,000, a 98.7% reduction. You do not need to rebuild your stack to benefit from the lesson.
MCP still wins in some cases:
Authentication flows that a CLI cannot handle cleanly
Remote services with no command line equivalent
Structured results you want typed and validated
For everything else, try the CLI first.
Let Tool Search Load on Demand
Sometimes you really do need dozens of tools available. That is where deferred loading earns its place.
How Deferred Loading Works
Instead of pasting every tool definition into the prompt, the client loads a small search tool plus a list of tool names. When Claude decides it needs something, it searches, and only the matching definitions get pulled in. In recent Claude Code releases this switches on automatically once MCP tool definitions would take up a large share of the window, around 10% at the time of writing.
Anthropic reported roughly an 85% drop in tool definition tokens in its own example. The idea is a library card catalog: you do not carry every book to your desk, you carry the index and fetch what you need.
Tune It for Your Setup
Behavior is controlled with the ENABLE_TOOL_SEARCH environment variable:
# Default: only kicks in when MCP tools get large
export ENABLE_TOOL_SEARCH=auto
# Lower the trigger point (percent of the window)
export ENABLE_TOOL_SEARCH=auto:5
Names and thresholds have shifted between releases, so check the documentation for the version you have installed.
💡 Always re-run /context after a change. If the MCP tools line did not shrink, the setting is not doing what you think.
There is a trade-off. The first use of a deferred tool costs one extra search step. If you call the same three tools in every session, keep that small server loaded directly and let tool search handle the long tail.
Stop Oversized Tool Responses
Definitions are a fixed cost. Responses are the variable cost that surprises people.
Cap the Output Size
Claude Code warns when a single MCP tool result passes about 10,000 tokens and truncates at 25,000 by default. You can lower the ceiling with an environment variable:
export MAX_MCP_OUTPUT_TOKENS=10000
A cap is a safety net, not a design. The better fix is a server that returns less in the first place. If you build or configure servers, push for these patterns:
Pattern
Problem
Better approach
Return a whole table
Tens of thousands of rows in context
Filter, sort and limit on the server
Embed a base64 image
Thousands of tokens per image
Return a hosted URL
Return raw page HTML
Markup noise dominates
Extract text from one selector
Poll status with the full payload
The bulk repeats on every poll
Return only a status field until done
Return URLs, Not Image Blobs
Images are priced by size, roughly width times height divided by 750 in input tokens. A 1,000 by 1,000 pixel image costs around 1,300 tokens, and a full-resolution screenshot costs several times that. Servers that generate pictures or clips should hand back a short URL and a status, never the pixels themselves.
Image and video generation is where this shows up first. PicassoIA's MCP connector works the right way: a generation call returns a prediction ID straight away, and you poll for status until a hosted URL appears. If you are picking a generator to call from an agent, text-to-image models like P-Image and Flux 2 Pro pair well with this pattern, and a video model like Seedance 2.5 Lite fits the same create, poll, fetch loop.
Isolate Work and Reset Often
Even a lean setup fills up over a long session. The last layer of defense is structure.
Give Subagents One Job
A subagent runs in its own context window and sends back only a summary. That makes it the perfect home for noisy work: browser automation, large searches, image generation loops. A subagent file in .claude/agents/ takes a name, a description, a tools allowlist and a model, so you can restrict it to exactly the two or three tools it needs.
Think of it as mise en place: each bowl holds one ingredient, and nothing else touches the counter. For mechanical jobs, a small model like Claude 4.5 Haiku is often plenty, and it keeps your main session free for the thinking that needs a bigger model. Newer releases also let a subagent declare its own MCP servers, so the heavy ones never touch the parent conversation.
When /compact Pays Off
/compact replaces the conversation so far with a summary, and you can steer it:
/compact keep the failing test names, file paths and the final design decision
Run it at natural breakpoints: a feature just landed, tests just went green, you are about to start a new phase. Do not wait for auto-compact. It fires when the window is nearly full, which can be right in the middle of a delicate edit.
When /clear Wins
If you are switching to an unrelated task, do not compact. Clear. Old context about a different bug is not an asset, it is noise that pulls answers sideways. After /clear, load only what the new task needs.
💡 Durable facts belong in CLAUDE.md, but keep it short. It loads every session, so every extra paragraph is a tax you pay forever. Review it monthly and delete rules the model already follows without being told.
How to Use Sonnet 5 on PicassoIA
Here is a low-stakes way to shrink the input side of your sessions before it ever reaches Claude Code. Claude Sonnet 5 on PicassoIA is built for coding and tool-use tasks, which makes it a handy sandbox for condensing a bloated CLAUDE.md, summarizing a long log, or drafting a tighter prompt. Paste the result into Claude Code instead of the raw mess.
Fill the required Prompt field. Paste the text you want condensed and tell the model what to keep, for example: "Reduce this log to the ten lines that explain the failure."
Set the effort level. The default is low, which turns thinking off for the fastest and cheapest replies. Raise it only for genuinely tricky reasoning.
Limit the output.max_tokens defaults to 8,192. For a summary, a much lower number keeps the reply tight.
Add a system prompt if you reuse the task. One line such as "Answer in bullets, under 150 words" locks the format across runs.
Attach an image only when needed. The max_image_resolution option defaults to 0.5 megapixels and scales pictures down before they reach the model, which saves time and money.
Generate and copy the result into your Claude Code session.
Settings That Save Tokens
Setting
Default
Token-saving move
effort
low
Leave low for summaries and rewrites
max_tokens
8192
Lower it to match the size of the answer you want
max_image_resolution
0.5 megapixels
Keep it low unless fine detail matters
system_prompt
empty
Set a short format rule once and reuse it
For harder reasoning jobs, such as planning a refactor across many files, Claude Fable 5 and Claude Opus 4.7 are on the same platform, so you can compare how each one handles the same trimmed-down input.
Put It to Work on PicassoIA
Here is the whole plan on one page, ordered by effort versus payoff:
Fix
Effort
Typical payoff
Run /context and record a baseline
2 minutes
Shows where the tokens go
Disable or remove idle servers
5 minutes
Often the biggest single win
Scope servers with .mcp.json
10 minutes
Stops global servers from loading everywhere
Replace thin servers with CLIs
15 minutes
Removes definitions entirely
Enable tool search
2 minutes
Large cut for big server stacks
Lower MAX_MCP_OUTPUT_TOKENS
1 minute
Protects against runaway responses
Subagents for noisy work
20 minutes
Keeps the main window clean
/compact at breakpoints, /clear between tasks
Ongoing
Prevents mid-task overflow
Start with the first two rows today. They take under ten minutes and usually return a surprising amount of room.
Once your window has space again, put it to work on something creative. Open PicassoIA, pick an image model such as GPT Image 2 or P-Image, and generate your first picture from a plain sentence. Then animate it with Seedance 2.5 Lite or PicassoIA Video. Try your own idea, change one detail in the prompt, and watch how much the result moves. That is the fastest way to see what these models can do, and every experiment is a good excuse to keep your context lean.