Large Language ModelsGenerate imagesGenerate videos

Reduce MCP Token Usage in Claude Code: Fix the Context Window Problem

MCP servers can eat tens of thousands of tokens before you type a single prompt. This article shows how to measure that cost with /context, remove idle servers, turn on tool search, cap tool output, and use subagents so Claude Code keeps its context window for your actual work.

Reduce MCP Token Usage in Claude Code: Fix the Context Window Problem
Cristian Da Conceicao
Founder of Picasso IA

You open Claude Code, type one short prompt, and the context meter already looks half full. Nothing is wrong with your prompt. The culprit is usually the pile of MCP servers you connected last month and forgot about. Every server hands Claude a list of tool definitions the moment a session starts, and every definition is paid for in tokens before any real work begins. Add a few bloated tool responses on top and the context window is gone long before the task is done. The good news is that this is one of the most fixable problems in the whole workflow. This article walks through the order that works: measure first, then prune servers, turn on tool search, cap output, and reset sessions at the right moment.

Overhead view of a cluttered oak desk with a dark laptop, coffee and tangled charging cables at dawn

If you recognize the symptoms, you are in good company. Sessions that feel sluggish, answers that forget instructions from ten minutes ago, and an auto-compact that fires in the middle of a refactor all point to the same thing: too much fixed overhead and too little room for the work itself.

Why MCP Servers Eat Your Context

The Model Context Protocol (MCP) lets Claude Code talk to databases, browsers, issue trackers, design tools and image generators. Each connection is genuinely useful. The price is that Claude has to know what every tool does, so each tool's name, description and full JSON schema gets loaded into the prompt.

Where the Tokens Actually Go

MCP overhead comes from four places, and only some of them are obvious.

SourceWhen it loadsTypical impactWho controls it
Tool definitionsSession startHundreds of tokens per tool, thousands per serverYou, by choosing servers
Tool responsesEvery callA few hundred to tens of thousands of tokensThe server and your output limit
Memory files like CLAUDE.mdSession startGrows with every rule you addYou
Conversation historyBuilds up all sessionGrows with every turnCompacting and clearing

Low-angle view of tall stacks of heavy archive boxes and ring binders on a metal shelf

Anthropic's engineering write-up on tool search describes a setup with 58 tools across five servers that consumed roughly 55K tokens before the conversation even started. That works out to about 950 tokens per tool. A server with 35 tools, such as a full-featured code hosting integration, can single-handedly eat a double-digit slice of a standard window.

💡 Quick math: if your tools average 900 tokens each and you connect 40 of them, you have spent about 36,000 tokens on a menu Claude may never order from.

The Real Cost of Idle Tools

Idle tools hurt in two ways. The first is plain space: a window that starts 30% full has 30% less room for code, logs and reasoning. Prompt caching softens the price of that fixed overhead on repeat turns, but it does nothing for the room it occupies.

The second is decision noise. When five servers expose overlapping search or fetch tools, the model has more near-duplicates to choose between. Anthropic's own testing of on-demand tool loading found that exposing fewer tools up front also improved how reliably the right one got picked. Fewer tools is not only cheaper, it is often more accurate.

Measure the Damage First

Do not start deleting servers on a hunch. Measure, change one thing, measure again.

Run /context Before Touching Anything

Open a fresh session in your project and run /context before sending any prompt. Claude Code prints a breakdown of the window by category: system prompt, system tools, MCP tools, memory files, messages and free space. Because you have not typed anything yet, the messages line sits near zero and everything else is pure overhead.

Over-the-shoulder view of a developer studying a dark terminal with blurred horizontal bars on a large monitor

Then run /mcp to list connected servers and their status. From a regular terminal, claude mcp list gives you the same inventory. Write down the MCP tools number from /context. That is your baseline.

Read the Numbers Like a Budget

There is no official threshold, but this rule of thumb works well in practice:

MCP tools share of windowVerdictWhat to do
Under 5%HealthyLeave it alone
5% to 15%Worth a lookRemove servers you touch less than weekly
Over 15%A real problemPrune hard and enable tool search

Check the memory files line too. A CLAUDE.md that has grown into a wall of rules costs tokens on every single session, just like a chatty MCP server does.

Cut and Scope Your Servers

The cheapest token is the one never loaded. Server hygiene beats every clever setting.

Disable First, Delete Later

Open the /mcp menu and switch off anything you do not need for today's task. Depending on your version, you can toggle a server there without losing its configuration. When you are sure a server is dead weight, remove it for good with claude mcp remove <name>.

Close-up of a hand flipping one brass toggle switch in a row of switches on an aged wooden panel

Ask three questions about every server:

  • Did I call a tool from it in the last week?
  • Does a command line tool already do the same job?
  • Does it duplicate tools another server already offers?

One honest "no" is enough reason to disable it.

Scope Servers to Projects

MCP servers can be added at three scopes: local, project and user. A project-scoped server lives in a .mcp.json file at the repo root, so it only loads where it matters:

claude mcp add --scope project my-db -- npx my-db-mcp-server

Keep your user scope almost empty. Put the database server in the repo that has a database, and the design server in the repo that has designs. For one-off sessions, you can also start Claude Code with only the servers named in a config file:

claude --strict-mcp-config --mcp-config ./mcp/docs-only.json

That session ignores every other configured server, which makes it ideal for a focused task or a clean before and after comparison.

Swap Servers for Plain CLIs

If a tool already exists as a command line program, such as git, gh, docker, psql or aws, Claude can run it through the shell. That costs zero tool definition tokens, because the model already knows how those commands work.

Close-up of two hands typing in a dim workspace with a small plant in the background

Anthropic pushed this idea further in its post on code execution with MCP. Presenting servers as code APIs that the agent calls from a script, instead of as individual tools, cut one example workflow from about 150,000 tokens to about 2,000, a 98.7% reduction. You do not need to rebuild your stack to benefit from the lesson.

MCP still wins in some cases:

  • Authentication flows that a CLI cannot handle cleanly
  • Remote services with no command line equivalent
  • Structured results you want typed and validated

For everything else, try the CLI first.

Let Tool Search Load on Demand

Sometimes you really do need dozens of tools available. That is where deferred loading earns its place.

How Deferred Loading Works

Instead of pasting every tool definition into the prompt, the client loads a small search tool plus a list of tool names. When Claude decides it needs something, it searches, and only the matching definitions get pulled in. In recent Claude Code releases this switches on automatically once MCP tool definitions would take up a large share of the window, around 10% at the time of writing.

Low-angle view of an antique wooden card catalog with one drawer open and a hand lifting a single blank index card

Anthropic reported roughly an 85% drop in tool definition tokens in its own example. The idea is a library card catalog: you do not carry every book to your desk, you carry the index and fetch what you need.

Tune It for Your Setup

Behavior is controlled with the ENABLE_TOOL_SEARCH environment variable:

# Default: only kicks in when MCP tools get large
export ENABLE_TOOL_SEARCH=auto

# Lower the trigger point (percent of the window)
export ENABLE_TOOL_SEARCH=auto:5

Names and thresholds have shifted between releases, so check the documentation for the version you have installed.

💡 Always re-run /context after a change. If the MCP tools line did not shrink, the setting is not doing what you think.

There is a trade-off. The first use of a deferred tool costs one extra search step. If you call the same three tools in every session, keep that small server loaded directly and let tool search handle the long tail.

Stop Oversized Tool Responses

Definitions are a fixed cost. Responses are the variable cost that surprises people.

Cap the Output Size

Claude Code warns when a single MCP tool result passes about 10,000 tokens and truncates at 25,000 by default. You can lower the ceiling with an environment variable:

export MAX_MCP_OUTPUT_TOKENS=10000

Side-view macro of clear water pouring from a glass pitcher into a measuring jug just starting to overflow

A cap is a safety net, not a design. The better fix is a server that returns less in the first place. If you build or configure servers, push for these patterns:

PatternProblemBetter approach
Return a whole tableTens of thousands of rows in contextFilter, sort and limit on the server
Embed a base64 imageThousands of tokens per imageReturn a hosted URL
Return raw page HTMLMarkup noise dominatesExtract text from one selector
Poll status with the full payloadThe bulk repeats on every pollReturn only a status field until done

Return URLs, Not Image Blobs

Images are priced by size, roughly width times height divided by 750 in input tokens. A 1,000 by 1,000 pixel image costs around 1,300 tokens, and a full-resolution screenshot costs several times that. Servers that generate pictures or clips should hand back a short URL and a status, never the pixels themselves.

Image and video generation is where this shows up first. PicassoIA's MCP connector works the right way: a generation call returns a prediction ID straight away, and you poll for status until a hosted URL appears. If you are picking a generator to call from an agent, text-to-image models like P-Image and Flux 2 Pro pair well with this pattern, and a video model like Seedance 2.5 Lite fits the same create, poll, fetch loop.

Isolate Work and Reset Often

Even a lean setup fills up over a long session. The last layer of defense is structure.

Give Subagents One Job

A subagent runs in its own context window and sends back only a summary. That makes it the perfect home for noisy work: browser automation, large searches, image generation loops. A subagent file in .claude/agents/ takes a name, a description, a tools allowlist and a model, so you can restrict it to exactly the two or three tools it needs.

Top-down view of a chef's mise en place with small ceramic bowls arranged in a neat grid on dark slate

Think of it as mise en place: each bowl holds one ingredient, and nothing else touches the counter. For mechanical jobs, a small model like Claude 4.5 Haiku is often plenty, and it keeps your main session free for the thinking that needs a bigger model. Newer releases also let a subagent declare its own MCP servers, so the heavy ones never touch the parent conversation.

When /compact Pays Off

/compact replaces the conversation so far with a summary, and you can steer it:

/compact keep the failing test names, file paths and the final design decision

Run it at natural breakpoints: a feature just landed, tests just went green, you are about to start a new phase. Do not wait for auto-compact. It fires when the window is nearly full, which can be right in the middle of a delicate edit.

Medium shot of a person wiping a large white whiteboard clean, with faint marker streaks left on half of it

When /clear Wins

If you are switching to an unrelated task, do not compact. Clear. Old context about a different bug is not an asset, it is noise that pulls answers sideways. After /clear, load only what the new task needs.

💡 Durable facts belong in CLAUDE.md, but keep it short. It loads every session, so every extra paragraph is a tax you pay forever. Review it monthly and delete rules the model already follows without being told.

How to Use Sonnet 5 on PicassoIA

Here is a low-stakes way to shrink the input side of your sessions before it ever reaches Claude Code. Claude Sonnet 5 on PicassoIA is built for coding and tool-use tasks, which makes it a handy sandbox for condensing a bloated CLAUDE.md, summarizing a long log, or drafting a tighter prompt. Paste the result into Claude Code instead of the raw mess.

  1. Open the model page. Go to Claude Sonnet 5 on PicassoIA.
  2. Fill the required Prompt field. Paste the text you want condensed and tell the model what to keep, for example: "Reduce this log to the ten lines that explain the failure."
  3. Set the effort level. The default is low, which turns thinking off for the fastest and cheapest replies. Raise it only for genuinely tricky reasoning.
  4. Limit the output. max_tokens defaults to 8,192. For a summary, a much lower number keeps the reply tight.
  5. Add a system prompt if you reuse the task. One line such as "Answer in bullets, under 150 words" locks the format across runs.
  6. Attach an image only when needed. The max_image_resolution option defaults to 0.5 megapixels and scales pictures down before they reach the model, which saves time and money.
  7. Generate and copy the result into your Claude Code session.

Settings That Save Tokens

SettingDefaultToken-saving move
effortlowLeave low for summaries and rewrites
max_tokens8192Lower it to match the size of the answer you want
max_image_resolution0.5 megapixelsKeep it low unless fine detail matters
system_promptemptySet a short format rule once and reuse it

For harder reasoning jobs, such as planning a refactor across many files, Claude Fable 5 and Claude Opus 4.7 are on the same platform, so you can compare how each one handles the same trimmed-down input.

Put It to Work on PicassoIA

Here is the whole plan on one page, ordered by effort versus payoff:

FixEffortTypical payoff
Run /context and record a baseline2 minutesShows where the tokens go
Disable or remove idle servers5 minutesOften the biggest single win
Scope servers with .mcp.json10 minutesStops global servers from loading everywhere
Replace thin servers with CLIs15 minutesRemoves definitions entirely
Enable tool search2 minutesLarge cut for big server stacks
Lower MAX_MCP_OUTPUT_TOKENS1 minuteProtects against runaway responses
Subagents for noisy work20 minutesKeeps the main window clean
/compact at breakpoints, /clear between tasksOngoingPrevents mid-task overflow

Start with the first two rows today. They take under ten minutes and usually return a surprising amount of room.

Once your window has space again, put it to work on something creative. Open PicassoIA, pick an image model such as GPT Image 2 or P-Image, and generate your first picture from a plain sentence. Then animate it with Seedance 2.5 Lite or PicassoIA Video. Try your own idea, change one detail in the prompt, and watch how much the result moves. That is the fastest way to see what these models can do, and every experiment is a good excuse to keep your context lean.

Share this article