Large Language ModelsGenerate imagesGenerate videos
MCP vs CLI for AI Agents: Token Usage and Which Is Better
Real benchmark numbers for MCP vs CLI in AI agents: 1,365 tokens against 44,026 for the same GitHub task, and a Playwright test where the gap nearly vanished. See where tokens go, when MCP earns its overhead, and how to cut the bill without dropping the tools you need.
Your agent has not written a single word of its answer yet, and it has already burned thousands of tokens. That is the real argument behind MCP vs CLI for AI agents: not which protocol is tidier, but which one leaves more room in the context window for actual work. One benchmark measured 1,365 tokens for a GitHub lookup run through the command line and 44,026 for the same lookup through MCP. Another, built on Playwright, found almost no gap at all. Both results are real, and the difference between them says more than either number alone. Below you will see where the tokens go, what the measurements show, when MCP still earns its weight, and how to pick without guessing.
What MCP and CLI Mean for Agents
Both approaches give a model hands. They differ in how the model finds out what those hands can do, and when it pays for that knowledge.
How MCP Delivers Tools
The Model Context Protocol is a JSON-RPC standard. When a session starts, the client asks each connected server which tools it offers. Every tool arrives as a name, a description, and an input schema, and all of it sits in the model's context so it can decide what to call. Each call and each result then travels as structured JSON.
The upside is real: typed inputs, authentication handled on the server side, and tools that any compliant client can find on its own. The downside is that the whole catalog is paid for up front, whether the task needs one tool or none.
How CLI Delivers Tools
With the CLI approach, the agent writes a shell command, the runtime executes it, and standard output comes back as plain text. There is no catalog and no handshake. The model already knows git, grep, curl, and jq from its training data, and if it forgets a flag it runs --help and reads a short page.
The cost shows up only for the commands actually used. That single design difference explains most of what follows.
💡 Quick definition: a token is a chunk of text the model reads or writes. Tool definitions, commands, and results all occupy the context window, and all of them count as input on the next turn.
Where the Tokens Actually Go
Token cost in an agent loop comes from three places: what the model is told about its tools, what it sends, and what comes back. MCP and CLI differ mostly in the first and the third.
Schema Overhead Before Any Work
A typical MCP server injects the definition of every tool into the conversation before the first question is asked. GitHub's official server exposes 43 tools, according to the Scalekit benchmark, so even a trivial query carries 43 schemas. Checkly measured the Playwright server's definitions alone at 5.9k tokens. Anthropic states the problem bluntly in its engineering write-up on code execution with MCP: tool definitions overload the context window.
Results Flowing Back Through Context
The second cost is easier to miss. With default MCP calls, every intermediate result passes through the model. Anthropic's example is a meeting transcript fetched from one service and written into another: the text flows through the context twice, which can add more than 50,000 tokens for a long recording. A shell pipeline can filter, count, or truncate that data before the model ever sees it.
Cost source
MCP, default setup
CLI
Tool catalog
All definitions loaded at session start
None, --help on demand
Call format
JSON envelope with name and arguments
One shell string
Results
Full response returned to the model
Piped, filtered, or written to a file
Tool lookup
Server declares its tools
Model recalls commands or reads help
What the Benchmarks Show
Two public tests give the clearest picture, and they point in different directions.
Scalekit's GitHub Test
Scalekit ran five read-only GitHub tasks with Claude Sonnet 4, 25 runs per approach, against the official GitHub MCP server. They compared plain CLI, CLI with short instruction files called skills, and MCP, which makes 75 runs in total. Tokens per task:
Task
CLI
CLI + skills
MCP
MCP vs CLI
Repo language and license
1,365
4,724
44,026
32x
PR details and review status
1,648
2,816
32,279
20x
Repo metadata and install
9,386
12,210
82,835
9x
Merged PRs by contributor
5,010
6,107
33,712
7x
Latest release and dependencies
8,750
6,860
37,402
4x
Averaged across the five tasks, that is about 5,200 tokens for CLI against about 46,000 for MCP, roughly 9x. Scalekit's estimate at 10,000 operations a month is about $3.20 for CLI against $55.20 for direct MCP, a 17x difference. Reliability moved the same way: CLI finished 25 of 25 runs, MCP finished 18 of 25 (72%), and all seven failures were TCP-level timeouts.
Look at the skills column too. On the last task the skills version used fewer tokens than plain CLI (6,860 against 8,750), likely because a short instruction file saved the agent some trial and error. On the simple lookups it cost more, since the instructions load every time.
⚠️ Read the limits: one model, one service, read-only tasks. Timeouts are connection problems, not reasoning errors, so the reliability gap says more about that server setup than about the protocol itself.
Playwright: The Gap That Closed
Checkly ran a different test: open a demo shop, search for a product, click through, add items to the cart, and validate the contents. The MCP session used 48k to 50k tokens of context. The CLI session, with skills installed, used 45k to 48k. After three runs the difference was negligible, and the explanation is simple: the Playwright CLI and the MCP server share a backend and write the same snapshot files to disk.
Checkly also concedes that the criticism of MCP was fair in 2025. Two things drove it: servers loaded every definition up front, and every action returned a full page snapshot inline. Both have since been softened, because modern agent harnesses defer loading MCP tools until they are needed, and servers can save snapshots to disk. Checkly's warning is worth repeating: AI advice comes with an expiry date measured in months.
The lesson is not that one side got it wrong. The 32x figure describes one implementation with every schema loaded. The near tie describes another that avoids the overhead. Token cost is a property of how a server is built, not of the protocol label.
Why CLI Often Wins on Cost
When both options run out of the box, CLI wins more often than not. Two reasons carry most of the weight.
Models Already Speak Shell
Decades of shell scripts, README files, and forum answers sit in the training data of every large model. The agent does not need a schema to use git log or grep -r, and a one-line command replaces a JSON call with a name, an arguments object, and a surrounding envelope. When it is unsure, --help costs one short page, not a full catalog at the start of every session.
Pipes Trim Output First
The most underrated CLI advantage is composition. Here is one way to list the titles of recent merged pull requests:
gh pr list --state merged --limit 10 --json title --jq '.[].title'
The model sees ten lines of titles. A default MCP call to a pull request tool often returns the full payload for each item: authors, labels, URLs, timestamps, review state. Most of it is ignored, but all of it is read, and all of it is billed.
A CLI session is also easier to debug. You can paste the same command into your own terminal and watch exactly what the agent saw.
Where MCP Earns Its Overhead
Raw token count is one axis. Some jobs need what only a protocol can give.
Auth and Permissions
A shell command runs with the permissions of whoever launched it. That is fine on your own laptop and a problem in a product where many users each connect their own accounts. An MCP server can hold scoped credentials per user and expose only the actions you intend, such as "read issues" without "delete repository." It also lets non-terminal clients, from desktop assistants to IDE panels, find tools without anyone installing a binary.
Long Sessions and Media Jobs
Stateful tools favor MCP. A browser session that must survive across dozens of turns, or a database connection with an open transaction, is awkward to rebuild from one-shot shell commands.
Media generation is a good example. Image and video models run as asynchronous jobs: you submit a request, get a job ID, and poll until the result is ready. PicassoIA's connector wraps that flow in nine tools at the time of writing: image generation, image editing, two video generators, status lookup, cancel, a list of past jobs, a model list, and an account check. The generation tools return a job ID together with a suggested wait before the next poll, so the agent knows when to check back instead of looping blindly. The models behind it include PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video, and Seedance 2.5 Lite, with up to five predictions running at once per account.
You can reach the same models through the REST API at https://api.picassoia.com/v1 with plain curl. That works well, but the agent has to remember the endpoint, attach credentials, and write its own polling loop. The MCP tools package those steps. Notice also how catalog size matters: nine small tools weigh far less than GitHub's 43, which is why a lean server hurts less than a sprawling one.
How to Cut MCP Token Cost
If MCP is the right fit, you do not have to accept the default bill.
Load Tools on Demand
Anthropic calls this progressive disclosure: let the model read tool definitions when it needs them instead of all at once. Two forms work. One is a filesystem layout where each tool is a small file the agent opens on demand. The other is a search function that finds and loads only the relevant definitions. Whatever your harness calls it, check whether deferred loading is switched on, and connect only the servers the current project needs.
Write Code Against Tools
Anthropic's bigger move is to present MCP tools as code the agent can call from a sandbox. The agent writes a short script, the script talks to the servers, and only a summary returns to the model. In their Google Drive to Salesforce example, token use dropped from 150,000 to 2,000, a saving of 98.7%.
Look at what that really is: CLI-style behavior, where data is filtered before the model sees it, layered on top of MCP's authentication and typed interfaces. The two camps end up borrowing from each other.
Quick wins you can apply today:
Disconnect idle servers. Every connected tool costs tokens, used or not.
Prefer small, focused servers over one server with dozens of tools.
Cap results. Use limit and field-selection parameters wherever the tool offers them.
Write bulky output to disk and return a path, as Playwright now does with snapshots.
Measure. Check your provider's usage numbers before and after each change.
Which Is Better for You
On raw token cost with default settings, CLI wins most of the time, by 4x to 32x in the best-documented benchmark. On access control, long-lived state, and hosted services, MCP wins. And when MCP tools load on demand and large results stay out of the context, the gap can shrink to almost nothing, as the Playwright test showed.
Situation
Better pick
Why
Local work with git, files, and builds
CLI
Familiar commands, filtered output
Short scripted or CI tasks
CLI
No handshake, small footprint
Server with 40+ tools at default settings
CLI or code execution
Schema overhead dominates
Many users, each with their own accounts
MCP
Scoped credentials per user
Long browser sessions
MCP
State survives between turns
Asynchronous media generation
MCP
Job IDs and polling hints
Pick CLI when the tool already exists as a command, the agent runs on a machine you control, and every token counts.
Pick MCP when you need per-user permissions, persistent state, or a hosted service with no good command-line tool.
Mix both when in doubt. Most real setups do: shell for local work, MCP for a few hosted services, each trimmed to what the task needs.
Try It on PicassoIA
Use Claude Sonnet 5 on PicassoIA
You can test these trade-offs on your own setup with Claude Sonnet 5, a model built for multi-step coding and tool use. Here is a quick audit workflow:
Open the model page and find the Prompt field.
Paste the names and descriptions of the tools your MCP servers expose, then ask which ones a typical coding task would never call.
Set Effort. low turns thinking off and answers fastest, which is enough for triage. Move to high when you want trade-offs reasoned out across several servers.
Leave Max Tokens at 8192 for long comparisons, or lower it when you want a terse verdict.
Add a System Prompt such as "You are a cost reviewer. Answer with a table and one line of advice" to keep replies short.
Ask for a shell equivalent of your most-called tool. Compare the two with wc -c for a rough size check, and use your provider's usage numbers for exact token counts.
💡 Tip: run the same audit with Kimi K2.6 or GPT 5.6 Sol and compare which tools each model would cut.
Then Make Your Own Images
Every photo in this article follows a simple pattern: one clear subject, one light source, one camera and lens choice. Try the same recipe yourself. Describe a workbench at nine in the morning, a hiker at a trail fork, or your own version of a tool-heavy desk, and run it through PicassoIA Image. Refine the result with PicassoIA Image Editor Pro, and when a still deserves motion, send it to PicassoIA Video. Pick a prompt from your next project and see what comes back.