Large Language ModelsGenerate imagesGenerate videos

Skills vs MCP vs Plugins in Claude: Differences and Token Usage

Skills load in stages at about 100 tokens each until used. MCP servers connect Claude to live tools and can flood the context window with definitions and results. Plugins bundle both into one install. See what each loads, when it loads and what it costs per turn.

Skills vs MCP vs Plugins in Claude: Differences and Token Usage
Cristian Da Conceicao
Founder of Picasso IA

Every Claude setup eventually hits the same wall: the context window fills up with things you installed and rarely use. The cause is usually a mix-up between three features that sound alike. Skills, MCP servers and plugins all extend what Claude can do, yet each one works differently, loads at a different moment and bills tokens in a different way. Pick the wrong one and you pay for it on every turn. Pick the right one and Claude gains a new ability for almost nothing.

This article sets out what each one is, when it enters the context window, what it costs, and how to choose between them. The figures come from Anthropic's own documentation, and anything estimated is labeled as an estimate.

The Short Answer

Here is the whole comparison in one table. The sections after it explain every row.

SkillMCP serverPlugin
What it isA folder with a SKILL.md file, plus optional scripts and reference filesA process or remote service that exposes tools to ClaudeA package that bundles skills, subagents, hooks and MCP servers
Its jobTeaches Claude how to do somethingGives Claude access to somethingDistributes a ready-made setup
Always in contextName and description, about 100 tokens per skillLittle with tool search, every tool definition without itName and description of each skill, agent and command Claude can invoke itself
Loaded when usedSKILL.md body (under 5k tokens), then only the files it readsThe tool call and its resultWhatever its bundled parts load
Where it runsClaude Code, claude.ai, Claude APIClaude Code, plus connectors on claude.aiClaude Code, plus claude.ai and Cowork (different component set)

A hiker packing a canvas backpack beside a luggage scale in a wooden cabin

Think of the context window as a backpack. Everything you pack gets carried on every single turn, whether you use it or not. Skills are folded flat and weigh almost nothing until you open them. MCP servers can be heavy, depending on how many tools they ship. Plugins are a box that holds both, so their weight is the sum of what is inside.

💡 Rule of thumb: a skill is a recipe, an MCP server is the supplier who delivers the ingredients, and a plugin is the meal kit box that ships both together.

How Skills Work

A skill is the cheapest way to change how Claude behaves. It is a directory containing instructions, optional code and reference material, written the way you would brief a new teammate. Claude reads it only when the task calls for it. Unlike a prompt, which lives inside one conversation, a skill sits on the filesystem and is available every time.

Three Loading Levels

Anthropic calls the design progressive disclosure. Content arrives in stages instead of all at once:

LevelWhen it loadsToken costWhat loads
1. MetadataAlways, at startupAbout 100 tokens per skillname and description from the frontmatter
2. InstructionsWhen the skill triggersUnder 5k tokensThe body of SKILL.md
3. ResourcesOnly when neededZero until accessedReference files, templates and scripts

A hand pulling one index card halfway out of an oak library card catalog drawer

The card catalog in the photo is the right mental model. Hundreds of cards sit in the drawer, but you only pull the one you need. A skill can hold dozens of reference files, and if the task only touches the sales schema, that single file is the one that loads. The rest cost nothing.

Scripts get even better treatment. When Claude runs validate_form.py, the code never enters the context window. Only the output does, something like "Validation passed". That makes a bundled script far cheaper than asking Claude to write equivalent code from scratch each time.

What a SKILL.md Looks Like

The file needs two frontmatter fields, which sit between two triple-hyphen lines at the top of the file. Here are the fields for a release notes skill:

name: release-notes
description: Draft release notes from merged pull requests in our house style. Use when the user asks for a changelog, release notes or a version summary.

Below the frontmatter comes the body, plain markdown that Claude reads when the skill triggers:

Steps:
1. List merged pull requests since the last tag.
2. Group them under Added, Changed and Fixed.
3. Write one plain sentence per item, no marketing language.

For the exact tone rules, see STYLE.md.

The limits are strict. A name allows 64 characters of lowercase letters, numbers and hyphens, and cannot include the reserved words "anthropic" or "claude". A description allows 1,024 characters and has to say both what the skill does and when to use it, because that sentence is what Claude matches your request against.

Two hands opening a ring binder with tab dividers and instruction sheets on a workbench

💡 Why the description matters twice: a vague one means the skill never triggers. A bloated one burns always-on tokens on every turn. Write it like a search snippet: what it does, when to use it, nothing else.

Where Skills Run

The same format works on three surfaces, with different rules on each:

  • Claude Code: drop a folder into ~/.claude/skills/ for personal use or .claude/skills/ for a project. Skills have the same network access as any program on your machine.
  • claude.ai: upload a zip under Settings, then Features. Custom skills are per user and need code execution enabled on a Pro, Max, Team or Enterprise plan.
  • Claude API: reference a skill_id in the container parameter next to the code execution tool. Skills run sandboxed with no network access and no runtime package installs, and they are shared across the workspace.

One catch: custom skills do not sync across surfaces. A skill uploaded to claude.ai is not visible to the API, and neither one sees your Claude Code folder.

How MCP Works

The Model Context Protocol is an open standard for connecting Claude to outside systems. An MCP server exposes tools, and Claude calls them: query a database, open a ticket, read a design file, drive a browser. Where a skill says "here is how we do it", MCP says "here is the system, go and touch it".

Tools Are the Product

A hand plugging a braided cable into a numbered patch panel in a small network closet

Each tool arrives as a definition: a name, a description and a JSON schema describing its inputs. Claude reads those definitions to decide which tool fits the request. You register a server with one command, for example:

claude mcp add --transport http stripe https://mcp.stripe.com

The scope you pick decides who gets it. Local (the default) loads only in the current project and stays private to you. Project writes a .mcp.json at the repo root so the whole team shares it. User makes the server available in all your projects.

Why Tool Definitions Cost So Much

A single definition usually runs a few hundred tokens once the schema is included. That is an estimate, not a documented figure, so measure your own setup with /context in Claude Code. A server with 40 tools at roughly 400 tokens each would add about 16,000 tokens before you type a word, if every definition loaded upfront.

Results cost tokens too, and they can be huge. Claude Code warns when a tool result passes 10,000 tokens and caps output at 25,000 tokens by default. You can raise the cap with MAX_MCP_OUTPUT_TOKENS. Oversized text results are saved to a file, and Claude gets a reference instead of the raw text.

Tool Search Changes the Math

On Claude Code v2.1.232 and later, MCP tool search is on by default. Instead of loading every tool at startup, Claude runs a ToolSearch call to find the tools a request needs. That makes it cheap to keep many servers connected.

Tool search does not apply everywhere. It is off with a custom ANTHROPIC_BASE_URL, when ENABLE_TOOL_SEARCH=false is set, and for models earlier than the Claude 4.5 generation on Google Cloud's Agent Platform. In those setups the old rule holds: every connected server taxes every turn.

How Plugins Work

A plugin adds no new kind of capability. It is a directory of components that Claude Code installs and loads as one unit, usually with a manifest at .claude-plugin/plugin.json. The components can be skills, subagents, hooks, MCP servers or a hooks module that draws panes and adds commands.

A Package, Not a Feature

A craftsperson filling labeled compartments of a wooden toolbox beside a closed shipping box

A skill inside a plugin runs under the plugin's namespace. A plugin named my-plugin with a review skill gives you /my-plugin:review, so two plugins can ship skills with the same name without clashing.

Most plugins arrive through a marketplace, which is a catalog that lists plugins and where to fetch them. Run /plugin to browse, then install by name, such as commit-commands@claude-plugins-official. Install scopes mirror MCP: user for every project on your machine, project for everyone who works in the repository, and local for you in one repository only.

Skills, subagents, hooks and MCP servers all work alone. Reach for a plugin when you want to hand several of them to teammates with one command, or publish versioned releases. A cloud session at claude.ai/code does not load the plugins in your local settings, so plan for that if your team works there.

What an Enabled Plugin Costs

An enabled plugin is part of every session, not only the ones where you use it. For each skill, agent and command that Claude can invoke on its own, the name and description sit in context on every turn. Those tokens count toward your usage and leave less room in the window even when nothing from the plugin runs. The full text loads only on use, and its MCP servers follow the tool search rules above.

The side effects go beyond tokens: its MCP servers run alongside each session, its hooks fire on their events, and whatever it runs, it runs as you. In the /plugin view, plugins from Anthropic's official marketplace show a Context cost estimate before you install, and the Installed tab lists plugins marked "Not used recently" that you could switch off with claude plugin disable.

Token Usage Side by Side

A brass balance scale with a thin slip of paper outweighed by a stack of leather ledgers

The Per-Turn Baseline

Here is what each option costs when idle versus when it fires:

Idle cost, every turnCost when usedWorst case
SkillAbout 100 tokens eachUnder 5k for the body, plus any file readA body near the 5k ceiling that triggers too often
MCP serverTool list via search, or all definitions without itTool call plus result, capped at 25k by defaultMany verbose tools with search off
PluginName and description per invocable partSame as its partsDozens of plugins enabled and unused

A quick worked example using the documented figure: 40 skills at about 100 tokens each add roughly 4,000 tokens. An MCP server with 40 tools at the estimated 400 tokens each adds about 16,000 when tool search is off. Same count, four times the weight, and the skills give more per token because the heavy content stays on disk until needed.

Where the Savings Come From

  • Keep SKILL.md lean. Move rarely needed detail into reference files so only the relevant one loads.
  • Ship scripts, not code snippets. Only a script's output enters context.
  • Trim descriptions. What and when, in one or two sentences.
  • Filter at the server. Return ten rows, not ten thousand.
  • Disable unused plugins. Every enabled one adds descriptions on every turn.
  • Use project scope for MCP. Load a server only where it earns its place.

Which One to Pick

Two colleagues reviewing a printed decision flowchart with sticky notes on a meeting table

Start from the job, not the feature name:

SituationBest fitWhy
Claude keeps getting your format wrongSkillInstructions, no live access needed
Claude has to read your database or ticket systemMCPNeeds a live connection
Your team needs the same setup on day onePluginOne install, versioned updates
A repeatable chore with a scriptSkillScript output is nearly free
Several steps mixing rules and live dataAll threeEach does one job

Pick a Skill for Repeatable Work

If you have ever pasted the same instructions twice, that is a skill. Brand voice rules, a pull request checklist, a report layout, a data cleanup script. Nothing leaves your machine, nothing stays running, and the idle cost is about 100 tokens.

Pick MCP for Live Access

Choose MCP when the data lives somewhere else and changes: a CRM, an issue tracker, a production database, a browser session. A skill cannot fetch fresh data on its own. It can only tell Claude how to handle the data once Claude has it.

Pick a Plugin to Share a Setup

When three teammates ask "how did you set that up?", package it. A plugin turns a skill, an agent and an MCP server into one install. For a purely personal setup it is overhead, so skip it.

Using All Three Together

An overhead view of an open meal kit box with a recipe card, vegetables and spice jars

Say your support team wants Claude to triage tickets. An MCP server for the helpdesk gives Claude access to open tickets. A skill holds the triage rubric, the escalation rules and the reply tone. A plugin ships both to twenty teammates with a single command. Short version: MCP to reach it, a skill to do it right, a plugin to ship it.

Mistakes That Waste Tokens

A kitchen drawer overflowing with tangled cables and chargers

  1. Wrapping a plain instruction in an MCP server. If no live system is involved, a skill does the same job for a fraction of the weight.
  2. Cramming everything into SKILL.md. Past the 5k mark you pay for text Claude rarely needs. Split it into reference files.
  3. Writing paragraph-long descriptions. The 1,024 character limit is a ceiling, not a target.
  4. Leaving unused servers connected with tool search off. The drawer of tangled cables in the photo is your context window.
  5. Returning whole tables from tools. One careless query can hit the 25,000 token cap in a single turn.
  6. Installing plugins "just in case". Each one adds descriptions to every turn, even in sessions where it never runs.
  7. Skipping the review of third-party code. Skills and plugins can run code with your privileges, so read what you install.

💡 Quick audit: run /context to see what is eating your window, then /plugin to switch off anything under "Not used recently".

Try It Yourself on Picasso IA

Skills, MCP and plugins live inside Claude's own apps and Claude Code. Picasso IA hosts Claude as a text model you can use right in the browser, which is handy for drafting a SKILL.md, tightening a description or sketching a tool schema before you commit it.

Use Claude Sonnet 5 on PicassoIA

  1. Open Claude Sonnet 5 on Picasso IA.
  2. Paste your draft into the Prompt field, for example "Rewrite this skill description in under 40 words."
  3. Set Effort. The default, low, turns off thinking for the fastest and cheapest reply. Raise it to high or max for a messy multi-file problem.
  4. Add a System Prompt to fix the role once, such as "You review SKILL.md files for token waste."
  5. Attach an Image if you want it to read a screenshot of an error or a diagram, and use Max Tokens (default 8,192) to cap the reply length.
  6. Run it, copy the result, and test it in Claude Code.

Documentation needs visuals too. For a README banner, a tutorial header or a diagram illustration, try Seedream 4.5, GPT Image 2 or Flux 2 Pro on Picasso IA, and browse every other model at picassoia.com/en/all-models. Describe your scene, pick a model and start creating your own images today.

Share this article