Large Language ModelsGenerate imagesGenerate videos

Is MCP Dead? The MCP vs CLI Debate Explained

MCP was declared dead after a wave of posts claiming CLIs are cheaper for AI agents. This article breaks down the real token numbers, the security trade-offs and the cases where each approach wins, so you can choose the right one for your own project.

Is MCP Dead? The MCP vs CLI Debate Explained
Cristian Da Conceicao
Founder of Picasso IA

In March 2026, Perplexity's CTO, Denis Yarats, said his company was moving away from MCP internally and leaning on plain APIs and command line tools instead. A post about it went viral, and within days the verdict was everywhere: MCP is dead. Developers shared screenshots of tool lists eating their context window, and the replies were full of one-line shell commands doing the same job in a fraction of the space.

Then something odd happened. The protocol did not die. It kept shipping new servers, kept winning support from the biggest names in AI, and now sits under the stewardship of the Linux Foundation. So which is it?

This article sorts the noise from the numbers. You will see what MCP actually does, where the token complaints are fair, why CLIs feel so natural to coding agents, and where a terminal simply cannot do the job. Along the way there is a decision table you can use today, and a short section on trying both approaches with PicassoIA's own MCP connector and REST API.

💡 Short answer: MCP is not dead, but the lazy way of using it is. Loading a hundred tool definitions up front is a design problem, not a protocol problem.

Developer hands typing on a worn mechanical board beside a plain terminal window

Why Everyone Is Asking This

The post that started it

The spark was a short statement with a long tail. Perplexity's CTO said the company was moving off MCP for its internal tooling and going back to direct APIs and CLIs. Blog posts with titles like "MCP is dead" and "MCP vs CLI" followed within weeks, and the argument hardened into two camps.

None of those posts were an official verdict. They were opinions, and several were backed by real benchmarks showing large token savings when an agent runs a command instead of loading a protocol server. The complaints were fair. The claim that the whole protocol is finished was too big for the evidence.

The numbers that circulated were striking. One browser automation comparison reported about 52,000 tokens to read a product page through an MCP snapshot, against roughly 1,200 tokens for a few targeted command line queries. Another benchmark, listing devices in a Microsoft admin tool, reported around 35 times fewer tokens with a CLI. Treat figures like these as directional, since they depend on the server, the task and the client. But the direction is hard to argue with: a bloated server is expensive.

What the critics actually say

Strip away the hot takes and four complaints remain:

  • Context bloat: every tool's name, description and JSON schema is loaded before the agent does anything useful.
  • Data detours: large results flow through the model even when the agent needs only one field.
  • Setup friction: each server needs its own install, config and credentials.
  • Built-in familiarity: models have seen git, curl, grep and docker in countless examples, so they call them well with no extra instructions.

Each point is true in some setups. None of them is a flaw in the idea of a shared protocol.

Context is not free. Every token spent on tool descriptions is a token not spent on your code, your documents or the conversation itself. Long prompts also cost money on every call, and models tend to lose track of details buried in the middle of a huge context. That is why the token argument lands so hard with people who run agents all day.

What MCP Actually Does

A connector, not a brain

Anthropic introduced the Model Context Protocol in November 2024 as a shared way for AI applications to reach outside tools and data. Think of it as USB-C for agents. Without a standard, every AI app needs a custom plugin for every service, and the number of integrations multiplies fast.

Aluminum docking hub with several different cables plugged into its ports

A client (the AI app) talks to a server (the integration) over a local pipe or over HTTP. The server advertises tools, resources and prompts, and the client lets the model call them. That is the whole idea: one plug shape, many devices.

The three building blocks map to everyday needs. A tool performs an action, like creating an issue or generating an image. A resource exposes data to read, like a file or a database row. A prompt is a reusable template the user can trigger, like "review this pull request". Most real servers lean almost entirely on tools, which is also where the token costs pile up.

Who backs it now

On December 9, 2025, Anthropic donated MCP to the Linux Foundation as a founding project of the new Agentic AI Foundation, co-founded with Block and OpenAI and supported by Google, Microsoft, Amazon Web Services, Cloudflare and Bloomberg. Protocols that are finished do not usually get a neutral home and that many competitors around one table.

So the useful question is not whether MCP survives. It is how people should call it.

The Token Cost Problem

Tool definitions eat context

Here the critics have a strong case. Reports put the official GitHub MCP server at roughly 50,000 tokens of tool descriptions before the agent has done a single thing. A short skill file that teaches the same workflow through the command line has been reported at around 200 tokens. That is a gap of two orders of magnitude, and it comes out of the space your model needs for the actual task.

Overhead view of a tall stack of printed pages beside one small index card

Results pass through the model

The second cost hides in the middle of a workflow. If an agent fetches a long transcript from one system and pastes it into another, every word crosses the model twice. A one-hour meeting transcript can run to ten thousand tokens or more, so passing it through twice means paying for it twice, even though the model never needed to read it.

In its November 4, 2025 engineering post on code execution with MCP, Anthropic described a Google Drive to Salesforce workflow that fell from about 150,000 tokens to about 2,000, a saving of roughly 98.7%. The agent moved the text in code, so the model saw only a short confirmation.

Read that carefully. The fix was still MCP. The agent just called it through code instead of one tool call at a time.

ApproachUpfront contextBest atWeak spot
Direct MCP tool callsHigh with many toolsSmall, curated toolsetsBloat and data detours
MCP with code executionLowBig datasets, many stepsNeeds a sandbox
CLI in a shellNear zeroDeveloper toolingNeeds a terminal
CLI plus a skill fileVery lowRepeatable workflowsNeeds upkeep

Why CLIs Feel So Good

Models already know the commands

Command line tools carry decades of documentation, forum answers and shell history. A model asked to list open pull requests by one author will write the right line on the first try:

gh pr list --state open --json number,title,author \
  | jq '.[] | select(.author.login == "maria") | .title'

One line, and only the final titles reach the model. The raw JSON never enters the context. Pipes, filters and redirects give agents composition for free.

Over-the-shoulder view of an engineer typing in a plain terminal window at a standing desk

Help text arrives when needed

A CLI does not announce its whole surface up front. The agent runs --help for the one subcommand it needs, reads a few lines, and moves on. That is progressive disclosure, and it is exactly what large tool schemas lack.

There is a human bonus too. Anything the agent runs, you can run yourself in a terminal to reproduce a bug. No hidden protocol, no special inspector.

💡 Rule of thumb: if a mature CLI already exists for the job (git, gh, aws, kubectl, docker), let the agent use it.

Where CLIs Break Down

No shell, no CLI

One Google DeepMind engineer made the cleanest counterpoint: if your agent lives inside a notes app on a phone, "just use the CLI" is not an option. The same goes for browser chat apps, locked-down enterprise assistants and any sandbox with no process access. Most people who use AI every day are not sitting in a terminal.

Permissions, identity and audit

A shell is a blunt instrument. An agent with a terminal inherits the filesystem, the environment variables and every logged-in session on that machine. You can fence it in, but the default is wide open.

To be fair, neither route is safe by default. A web page, an email or a ticket can carry hidden instructions aimed at the model, and that risk exists whether the text arrives through a tool call or through the output of curl. Sandboxes, read-only credentials and human approval for risky actions matter in both worlds. The difference is that MCP gives you a natural place to enforce those limits, while a raw shell leaves it up to you.

Close-up of a brass padlock on the steel latch of a server rack door

An MCP server can expose only the tools you choose. A remote server can use OAuth, so each person acts under their own identity, and a central gateway can log every call. For a team of fifty people, that beats fifty laptops each holding a long-lived token in a dotfile.

Wide view down a tidy data center aisle between rows of server racks

How MCP Is Adapting

The protocol is responding to the same criticism the critics raised, and quickly:

  • Code execution: the agent writes a small script that calls MCP tools, filters the data in the sandbox, and returns only the result. Cloudflare's Code Mode applies the same idea by turning tools into a typed API the model writes code against.
  • On-demand tool loading: clients such as Claude Code now load tool definitions only when the agent needs them through tool search, instead of dumping every schema at the start.
  • Curated servers: the better servers ship five to ten well-named tools, not a mechanical wrap of ninety REST endpoints.
  • Remote servers with OAuth: one hosted server, many users, proper sign-in and central logs.

Aerial view of a large low-rise data center campus at golden hour

A practical summary of the current mood: use a CLI when you have a terminal, use a curated MCP server when you need a protocol, and never auto-convert an entire API into dozens of tools.

A Simple Way to Choose

Three questions to ask

  1. Does the agent have a shell? If not, MCP is probably your only route.
  2. Whose identity does it act under? If many users need separate sign-in and audit, lean on MCP.
  3. How big are the intermediate results? If they are huge, run the work in code, whether through a CLI pipeline or MCP with code execution.

Two colleagues reviewing a laptop in a glass meeting room with a whiteboard behind them

SituationBetter pickWhy
Local coding agent with a terminalCLILowest overhead, models know the tools
Tool already has a great CLICLIPipes keep data out of the context
Browser or phone chat appMCPNo shell available
Many users, shared tools, audit needsMCPOAuth, scoped tools, central logs
SaaS with no CLI and OAuth onlyMCPStandard sign-in and tool schema
Long workflow moving big filesCode executionData stays out of the model

Most teams end up with both. MCP handles the few systems that need sign-in and governance, and the CLI handles everything developers already do in a terminal.

Here is how that looks in practice. A solo developer working in a terminal agent will reach for git, gh and docker all day and add one or two MCP servers for things with no CLI, like a design tool or a ticket tracker. A support team using a chat assistant in the browser has no shell at all, so a hosted MCP server with sign-in is the only workable path. A platform team running dozens of agents wants central logs and scoped tools, so it puts a gateway in front of a few curated servers.

Pick a model that handles tools

Either route depends on a model that calls tools reliably. On PicassoIA you can compare several large language models in one place: Claude Sonnet 5 for automating coding tasks, GPT 5.6 Sol for complex coding work, Kimi K2.6 for building agents, and Gemini 3.5 Flash for fast chat and code. Give the same tool task to two of them and watch how each handles a failed call.

Try Both With PicassoIA

Image and video generation is a fine test bed for this debate. Outputs are large, jobs take time, and the tool has to report progress. PicassoIA exposes the same four models through an MCP connector and a REST API: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite for video with audio.

Connect through MCP

  1. Sign in at picassoia.com and open the MCP connections page in your account (picassoia.com/en/mcp/accounts).
  2. Add the connector to your AI client, following the instructions shown on that page.
  3. Ask for an image in plain language, such as "a 16:9 photo of a lighthouse at dawn".
  4. The tool starts a job and returns an ID. The assistant checks it until it finishes, then shows you the result.

Up to five jobs can run at once per account, shared across every connection.

Notice what the protocol does well here. The assistant does not need to know a URL, a header or a polling interval. The tool descriptions tell it how to start a job and how to check on it, and the result comes back in the chat. That convenience is exactly what MCP promises, and for a non-technical user in a browser it is the difference between "works" and "not possible".

Creative director studying printed landscape photographs on a studio wall beside an open laptop

Call the REST API From a Terminal

The same models are available at https://api.picassoia.com/v1, with a bearer token that starts with pia_sk_. The shape follows the Replicate style: create a prediction, poll it, fetch the output.

curl -s https://api.picassoia.com/v1/models/picassoia/picassoia-image/predictions \
  -H "Authorization: Bearer $PICASSOIA_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"input": {"prompt": "A lighthouse at dawn, natural light, 16:9"}}'

The response carries a prediction ID. Poll GET /v1/predictions/{id} until the status reads succeeded, then download the file. Check the API pages on picassoia.com for the exact input fields of each model.

💡 Try it yourself: run the same prompt once through the MCP connector and once through curl. Time both, count the steps, and decide which one fits your workflow.

MCP is not dead and the CLI is not a fad. They are two tools with different jobs, and the smart move is to match each one to where your agent actually lives. Open picassoia.com, pick a model, write your first prompt and see both paths in action with your own images.

Share this article