An MCP server is a door into your systems that an AI model opens on its own. Once a client such as a chat assistant or a code editor connects, the model can read files, query databases, call internal APIs and send messages, usually with the permissions of whoever installed the server. That is why MCP security best practices belong in the first sprint, not the postmortem. This article maps the risks that actually show up in production, lays out a 20 point checklist you can paste into a ticket, and walks through an audit your team can finish in one afternoon.
Short version, in case you are in a hurry: treat every tool description as untrusted input, give each server the smallest permission set that works, keep secrets out of the model's reach, and log every tool call so you can replay what happened.
Why MCP Security Is Different

Classic API security assumes a developer wrote the code that calls your endpoint. MCP breaks that assumption. The Model Context Protocol lets a language model pick tools at run time, based on text it reads. Text can be forged, and the model cannot reliably tell a legitimate instruction from a planted one. That makes MCP server security a different problem from locking down a REST API.
The Model Is the New Caller
In a normal integration, a human reviews the code path before it ships. With MCP, the caller is a probabilistic system that reads tool names, descriptions and results as part of its prompt. A sentence hidden in a tool description carries nearly the same weight as a sentence in your own system prompt. That single fact explains most MCP incidents.
It also changes who counts as an attacker. You do not need access to the server. Anyone who can put text in front of the model can try to steer it: the author of a web page, a customer filing a support ticket, a stranger opening an issue on a public repository. Good AI agent security starts by admitting that the input channel is open to the world.
Where the Trust Boundaries Sit
Draw four boundaries before you write a line of config:
| Boundary | What crosses it | Who you trust |
|---|
| User to client | Prompts and approvals | The signed-in user |
| Client to server | Tool calls and results | Only servers you vetted |
| Server to backend | Queries, files, API calls | A scoped service identity |
| Server to internet | Fetched pages, webhooks | Nobody |
Transport matters too. A local server launched over stdio is a process on the user's machine running with the user's permissions, so a malicious package is effectively arbitrary code execution. A remote server over HTTP adds network exposure, authentication and session handling. Each transport needs its own threat model, and a server that offers both needs two reviews.
💡 Tip: The riskiest setup is one agent that can read private data, read untrusted content and send data outward. Security researchers call this combination the "lethal trifecta". Remove at least one leg of it and most exfiltration paths close.
The Seven Risks That Matter
Not every threat deserves equal attention. These seven show up again and again in public MCP research and incident reports, and each one maps to checklist items further down.
| # | Risk | What happens | Typical impact |
|---|
| 1 | Prompt injection | The model follows instructions hidden in a page, ticket or file | Data leak, unwanted actions |
| 2 | Tool poisoning | Malicious instructions sit inside tool descriptions | Silent data theft |
| 3 | Rug pull | A server changes its tools after you approved it | Approved is not the same as current |
| 4 | Leaked credentials | Tokens end up in configs, logs or model context | Account takeover |
| 5 | Excessive permissions | A server runs with admin scope | One mistake becomes a breach |
| 6 | Token passthrough | A server forwards tokens that were not issued for it | Access beyond intent |
| 7 | Supply chain | Look-alike or backdoored server packages | Code runs on your machine |
Prompt Injection Through Tool Output

Prompt injection is the headline risk because it needs no exploit code. In 2025, researchers showed that a GitHub integration could be steered by a malicious issue in a public repository into reading private repositories and posting the contents back. The model did exactly what it was asked. The request just came from an attacker.
Defenses that work:
- Treat tool results as data, never as instructions. Wrap results in clear delimiters and say so in the system prompt.
- Split duties. The agent that reads untrusted pages should not hold write access to anything valuable.
- Gate outbound actions. Sending, posting and deleting need a human click.
- Filter both directions. Run a safety classifier over inputs and outputs, as shown in the tutorial below.
Tool Poisoning and Rug Pulls
Tool poisoning hides instructions inside a tool's description or schema. You see a harmless "add two numbers" tool. The model sees an extra paragraph telling it to read a credentials file and pass the contents as a parameter. A rug pull is the slow version: the server behaves during review, then quietly changes its tool definitions after you approved it.
The fix is boring and effective. Hash the full tool manifest (names, descriptions and schemas) at approval time, compare it on every connection and refuse to run when the hash changes. Show users the full description, not a shortened summary.
Supply chain deserves its own line. In 2025 a look-alike email server published on npm was reported to copy outgoing messages to an outside address after a routine update. Pinning versions and keeping an allow-list (checklist items 9 and 10) are the defense.
Leaked Credentials and Tokens

Secrets leak in unglamorous ways: a token pasted into a config file that gets committed, an environment variable echoed into a debug log, a database password returned inside a tool result and now sitting in the model's context for the rest of the session. Once a secret enters context, any later injection can ask the model to repeat it.
Store secrets in a vault or your operating system's credential store, hand them to the server process at start-up and never return them in output. Prefer short-lived tokens that expire in minutes, because a stolen token that dies at lunchtime is a small problem.
Excessive Permissions Add Up

A server that "just needs to read tickets" often ships with an admin token because that was the fastest path to a demo. Then a single injected instruction can close tickets, export customer lists or change settings. Permissions are the blast radius, and you choose its size on day one.
Apply least privilege to the narrowest verb: read, not write; one project, not the whole workspace; one table, not the database. When a vendor only offers an all-or-nothing token, put a thin proxy server in front that exposes only the calls you accept.
A 20 Point Security Checklist

Paste this into your tracker and score each server against it. Anything you cannot tick today becomes a ticket with an owner and a date.
Identity and Access
- Use OAuth 2.1 with short-lived access tokens for remote servers. Avoid static shared tokens.
- Check that every token was issued for your server (audience validation). Never forward a client token to a downstream API. Request a separate one.
- Give each server its own identity so you can revoke one without touching the rest.
- Default to read-only scopes. Write scopes need a written reason.
- Require human approval for destructive or outbound actions: deleting, sending, paying, publishing.
- Bind sessions to the signed-in user and never treat a session ID as proof of identity.
- Remove access for departing teammates on their last day, through an automated offboarding step.
Server Hardening

- Run local servers in a container or sandbox with no access to the home directory by default.
- Pin versions and checksums, and review the diff before every update.
- Keep an allow-list of servers your team may install. Block the rest.
- Re-approve a server whenever its tool list or descriptions change.
- Validate every tool argument on the server: paths, SQL, URLs and shell strings. Never pass model output to a shell unescaped.
- Bind local HTTP servers to 127.0.0.1 and check the Origin header to stop DNS rebinding.
- Keep debugging tools such as protocol inspectors off shared networks and behind authentication.
Data and Network
- Store secrets in a vault or OS credential store, never in prompts, repo files or tool results.
- Redact secrets and personal data from tool results before they reach the model.
- Restrict outbound traffic with an egress allow-list so a hijacked agent cannot post data to arbitrary hosts.
- Enforce rate limits and spending caps per server and per user.
- Log every tool call with user, server, argument hash, result size and timestamp. Ship logs to storage the agent cannot edit.
- Maintain a kill switch that disables any server in under five minutes, and rehearse it.
💡 Tip: Items 5, 11 and 17 are the cheapest to ship and they close the most dangerous paths: unapproved actions, silent tool changes and data leaving the network. If your week is short, start there.
How to Run an MCP Audit

An audit answers one question: what can the model do right now, and who approved it? Block out an afternoon, invite one engineer and one security reviewer, and work in three passes. Before the first pass, write a one page threat model with four answers: which data the agent can reach, which untrusted content it reads, where it can send data, and which actions are irreversible.
Inventory Every Server
You cannot protect what you cannot list. Collect servers from client config files, IDE settings, shared team repositories, CI pipelines and any "temporary" setups that never got removed. Record this for each one:
| Field | Example |
|---|
| Name and source | Vendor package, internal repo or community project |
| Transport | Local stdio or remote HTTP |
| Identity used | Service account, personal token or none |
| Scopes granted | Read tickets, write files, admin |
| Data it can reach | Customer records, source code, public web |
| Owner | A named person, not a team alias |
Expect surprises. Teams routinely find servers nobody remembers installing, running under a personal admin token.
Replay and Review the Logs
Pull thirty days of tool calls and look for patterns that should not exist: calls at odd hours, unusually large results, tools invoked right after the agent read an external page, and arguments that contain file paths or URLs the user never mentioned. For triage at scale, a model such as Claude Sonnet 5 can group thousands of calls into a short list of oddities, provided you redact secrets and personal data before sending anything.
Score and Fix Findings
Give each finding a severity and a deadline. Keep the scale small so people actually use it:
| Severity | Example finding | Fix window |
|---|
| Critical | Admin token in a shared config file | 24 hours |
| High | No approval step for outbound email | 1 week |
| Medium | Unpinned server version | 30 days |
| Low | Missing owner on a test server | Next review |
Re-run the audit after every major change and at least once a quarter. A server list that was clean in January is rarely clean in June.
Monitoring and Incident Response

Prevention fails eventually, so plan for the day it does. The goal is to notice within minutes and contain within an hour.
What to Log
Log enough to reconstruct a session without storing secrets. Capture user, client, server, tool name, a hash of the arguments, result size, latency and the approval decision. Alert on three signals first: a burst of calls to one tool, any call to a tool that was never used before, and large results shortly after the agent fetched external content.
Kill Switches and Rollbacks
Every server needs an off switch that one person can flip without a deploy. Revoke its tokens, remove it from the allow-list and rotate any secret it touched. Then restore the last known-good manifest hash. Practice this once a quarter so the first real attempt is not also the first rehearsal.
💡 A platform-side ceiling helps as well. PicassoIA, for instance, caps each account at 5 concurrent predictions across every connection, so a runaway agent cannot flood the queue. Ask your own vendors what their ceilings are.
How to Use Llama Guard on PicassoIA
A safety classifier will not replace permissions or approvals, but it adds a useful filter. Llama Guard 4 12B reads text or images and returns a safe or unsafe verdict with the matching harm category, spanning areas such as violence, hate and dangerous instructions. It is not a dedicated prompt injection detector, so use it as one layer next to the controls above, for example to screen what users submit and what your agent is about to send back.
- Open the model page. Go to the Llama Guard 4 12B page on PicassoIA.
- Fill the required fields. Paste the content to check into Prompt. In System Prompt, state your policy and the output you want, such as "Reply only with safe or unsafe, then the category."
- Tune the optional settings. Set Temperature to 0 for repeatable verdicts, lower Max Completion Tokens from the default 512 to something short, and attach screenshots through Image Input when you need to check an image.
- Run it. Click generate and read the verdict. Anything marked unsafe goes to a human reviewer instead of the agent.
- Save and compare. Keep a small set of test inputs and rerun them whenever you change the system prompt.
| Parameter | Suggested value | Why |
|---|
| Temperature | 0 | Same input, same verdict |
| Max Completion Tokens | 64 to 128 | A verdict is short |
| System Prompt | Policy plus a strict output format | Easy to parse in code |
| Top P | 1 (default) | Leave alone when temperature is 0 |
| Image Input | Screenshots to check | Optional |
Mistakes Teams Keep Repeating
- Approving once and never looking again. Tool descriptions change. Re-approval on every change is the cheapest control you have.
- Trusting a server because it is popular. Popularity is not a review. Check maintainers, release history and what the server can reach.
- Putting secrets in the prompt "just for testing". Test secrets become production secrets within a week.
- Giving the agent every tool. Load only the servers a task needs. Fewer tools means fewer chances to be steered.
- Skipping logs because nothing has happened yet. Without them you will not notice a quiet leak.
Create Your Own Images on PicassoIA

Security work is mostly invisible, which makes it hard to explain. Good visuals help: a photo header for a runbook, a scenario picture for a training slide, a data center aisle for the internal wiki. PicassoIA turns a text prompt into photorealistic images and short videos, with no design software and no stock photo license to track.
Try it with the topic you just read about. Describe a scene such as "a security engineer reviewing a printed checklist in a quiet server room, soft window light, 35mm lens", run it, then adjust the angle, lighting and lens until it fits your document. Open PicassoIA, browse the full model list and experiment with your own prompts today.