Large Language ModelsGenerate imagesGenerate videos
Replicate MCP Server: Run Models From Claude, Claude Code and Cursor
Add the official Replicate MCP server to Claude Desktop, Claude Code and Cursor, then search models and run image, video and text predictions without leaving your chat or editor. Includes exact commands, JSON config, token safety, costs and a zero-setup alternative on PicassoIA.
You can now ask Claude to find an image model on Replicate, run it with your prompt, and hand back the finished file without leaving the chat. That is the whole point of the Replicate MCP server. It plugs Replicate's HTTP API into any client that speaks the Model Context Protocol, so Claude Desktop, Claude Code and Cursor can search thousands of models and create predictions on your behalf.
The setup takes about five minutes. What slows people down are the small details: which URL to paste, which transport to pick, where the token goes, and what shows up on the bill afterward. Below you get the exact commands for each client, a token safety checklist, a look at the experimental code mode, and a zero-setup route for readers who only want finished images and videos.
What the Replicate MCP Server Does
Replicate hosts a large catalog of image, video, audio and text models behind a single API. The official MCP server wraps that API as tools, so an assistant can call it directly instead of you juggling curl commands and browser tabs. According to Replicate's docs, it supports all of the operations in Replicate's HTTP API: searching, listing and inspecting models, then creating and fetching predictions.
How a Prompt Becomes a Prediction
Say you type: make a 16:9 photo of a lighthouse at dawn. The client forwards your message to the language model, the model decides it needs an image tool, and it calls the server. The server relays the request to Replicate, which runs the chosen model on its GPUs and returns an output URL. The sequence usually looks like this:
The assistant searches for a model that fits the task.
It reads that model's input options, such as the prompt field and the aspect ratio.
It creates a prediction with your prompt and settings.
It fetches the result and shows you the output link.
💡 Tip: If you already know the model you want, name it in the prompt. That skips the search step and keeps results repeatable from one session to the next.
The same loop works for video and audio models. Ask for a five second clip and the assistant picks a video model, fills in the duration field, and polls until the run finishes. Video runs take longer than images, so expect a wait of a minute or more instead of a few seconds.
The Tools You Get
Replicate's MCP page names three tools, and the docs list the underlying operations:
Tool or operation
What it does
search_models
Finds models by task or description
create_predictions
Starts a run with the inputs you supply
list_hardware
Shows the hardware available for running models
models.list, models.get
Lists models and returns one model's metadata
predictions.get
Fetches the status and output of a run
Hardware matters because it sets both speed and price, so list_hardware is worth a try before a big batch.
Hosted or Local Setup
You can run the server in two ways, and the choice decides how authentication works.
Hosted server
Local server
Address or command
https://mcp.replicate.com
npx -y replicate-mcp
Sign in
Browser approval flow
REPLICATE_API_TOKEN environment variable
Runs where
On Replicate's side
On your machine, needs Node
Docs verdict
Recommended for most users
Alternative with manual token setup
Choose the hosted server unless you have a reason to run a process on your own machine. It removes the token file from your setup entirely, and it is the path every example in the next sections uses first.
⚠️ Heads up: Older tutorials tell you to install mcp-replicate. That was a community experiment, and its README says it is no longer in active development. The official server replaced it, so skip the old package.
Hosted Server at mcp.replicate.com
The hosted server is the one Replicate recommends. Clients connect to the SSE endpoint, https://mcp.replicate.com/sse, and you finish with a web-based approval step where you paste a Replicate API token. Create one at replicate.com/account/api-tokens before you start and keep it on your clipboard.
Local Server With npx
The local option runs the same server as a process on your machine. Set REPLICATE_API_TOKEN, start it with npx -y replicate-mcp, and stop it with Ctrl+C. Most clients launch it for you from a config file like this:
VS Code with GitHub Copilot reads .vscode/mcp.json and uses a top-level servers object instead of mcpServers, so copy the block and rename that one field. If the server never shows up, check that Node is installed and that npx runs from a normal terminal.
Connect It to Claude Desktop
Claude Desktop is the quickest route because the hosted server needs no config file at all.
Add the Custom Connector
Open Settings and go to Connectors.
Click Add Custom Connector.
Enter the server URL: https://mcp.replicate.com/sse.
Click Connect and paste your Replicate API token.
Click Log in and Approve.
Restart Claude Desktop. The tools appear under the slider icon below the chat input.
Prefer the local server? Put the JSON from the previous section into claude_desktop_config.json and restart the app.
First Prompt to Try
Start with something that forces the full search, inspect and run loop:
Search Replicate for a fast photorealistic text-to-image model, show me its input options, then generate a 16:9 photo of a lighthouse at dawn.
Fast FLUX variants such as FLUX Schnell are typical candidates for a request like that, though the search decides what comes up first. If the answer arrives as a link instead of an inline picture, that is normal: the server returns the output URL, and the client decides how to display it.
Once that works, get specific. Name the aspect ratio, the lens, the light direction and the number of variations you want. Vague prompts waste runs, and every run is billed.
Connect It to Claude Code
Claude Code takes one command, and it keeps the connection across projects.
One Command Install
Install the CLI if you do not have it yet, then add the server:
npm install -g @anthropic-ai/claude-code
claude mcp add replicate https://mcp.replicate.com/sse --transport sse --scope user
The --scope user flag stores the server in your user configuration, so every project you open gets access to it.
Authenticate With /mcp
Launch claude, type /mcp, choose the Replicate entry, and a browser window opens for the same approval flow Claude Desktop uses. Once it reports as connected, ask for a run in plain language.
Because Claude Code works inside your repository, it can do more than show a link. Try this: Generate three hero image options for the landing page, then download each output into public/images. The download step matters more than it looks, and the costs section below explains why.
If the tools do not appear, run /mcp again and check the status line for the server. A server that shows as disconnected usually needs a fresh authentication, not a reinstall.
Connect It to Cursor
Cursor reads MCP servers from a JSON file, so this setup is a few lines of config.
Edit the mcp.json File
Open Cursor Settings, go to Tools & Integrations and click New MCP Server. Cursor opens ~/.cursor/mcp.json, where you add the hosted server through the mcp-remote bridge:
For a single repository, create .cursor/mcp.json in the project instead, and use the local-server JSON from earlier with your token in env. Save the file and the server should appear in the same settings panel, along with its tools.
From there, Cursor's agent can generate placeholder art, product shots or icons while it edits the component that needs them. That is the real payoff: the asset and the code change land in one conversation.
Other Clients That Work
The same server works beyond these three:
Claude.ai in the browser
Codex CLI
Google Gemini CLI
Windsurf
Chorus
VS Code with GitHub Copilot, using .vscode/mcp.json
The pattern never changes: point the client at the hosted URL or launch replicate-mcp locally, then authenticate once.
Tokens, Costs and Code Mode
Three things decide whether this setup stays pleasant: how you handle the token, what each run costs, and how far you push automation.
Protect Your Token
Create the token at replicate.com/account/api-tokens and treat it like a password.
Prefer the hosted flow when you can. The token goes through the approval screen instead of sitting in a file in your repository.
With the local server, keep the token in an environment variable, or keep the config file outside version control. Add .cursor/mcp.json or .vscode/mcp.json to .gitignore whenever they hold a token.
Rotate the token immediately if it shows up in a commit, a screenshot or a shared chat.
What You Pay For
The MCP server is a front door to the same API, so every prediction is billed to your Replicate account like any other API call. The price depends on the model and the hardware it runs on, which is why list_hardware and a quick look at model pages help before a large batch.
A simple habit keeps the bill predictable. Draft with a fast, cheap model such as FLUX Schnell, generate four options, pick the best composition, and rerun only that one prompt on a higher-quality model. Tell the assistant this rule once at the start of a session and it will follow it for every request after that.
Item
What happens
A prediction you run
Billed to your Replicate account
API output files
Removed after an hour by default
API input files and logs
Removed after an hour by default
Predictions made on the website
Kept indefinitely
💡 Tip: Output URLs from API predictions disappear after about an hour. Ask the assistant to download or upload every result right away, especially in a long session.
Code Mode in a Sandbox
Replicate also ships an experimental code mode. Instead of calling one tool at a time, the model writes TypeScript and runs it inside a Deno sandbox. That suits jobs that chain several predictions, such as generating six variations and keeping the best two. It needs Deno installed. Add it to Claude Code with:
claude mcp add "replicate-code-mode" --scope user --transport stdio -- npx -y replicate-mcp@alpha --tools=code
The @alpha tag tells you what to expect. Treat it as a testing ground, keep the regular server installed next to it, and check its output before you rely on it for production work.
A Zero-Setup Alternative on PicassoIA
If you only want the finished image or clip and not an afternoon of configuration, PicassoIA keeps hundreds of models on browsable pages with no tokens to paste. Pair it with an LLM such as Claude Sonnet 5 to draft prompts, then generate in the browser.
How to Use PicassoIA Image
PicassoIA Image is the platform's own text-to-image model, and it is a fast place to test a prompt before you automate anything.
Write a prompt with a subject, a setting, a light direction and a lens, for example a lighthouse at dawn, low-angle, soft golden light from the left, 35mm lens, film grain.
Choose the aspect ratio you need, such as 16:9 for blog headers.
Generate, review the result, and tweak one detail at a time.
💡 Tip: Change one variable per run. When the result improves, you know which edit caused it.
PicassoIA API and MCP
PicassoIA also exposes a Replicate-style REST API, so the create, poll and fetch pattern you just set up will feel familiar:
Base URL: https://api.picassoia.com/v1, authenticated with a Bearer token from your account
Create a run: POST /v1/models/{owner}/{name}/predictions
Check a run: GET /v1/predictions/{id}
Cancel a run: POST /v1/predictions/{id}/cancel
Four models are available through the API and the MCP connection: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite. Each account gets 5 concurrent predictions, shared across tokens and MCP connections, and prompts can run up to 4,000 characters. API and MCP access depend on your plan, so check the pricing page before you build on it, and read the docs at picassoia.com/en/api.
Replicate MCP server
PicassoIA API and MCP
Catalog
Thousands of models
Four models
Sign in
Replicate API token
Bearer token from your PicassoIA account
Billed to
Your Replicate account
Your PicassoIA plan
Best fit
Wide model choice and experiments
A fast, fixed set for images and video
Make Your First Image Today
Pick the route that matches your day. If you live in a terminal or an editor, connect the Replicate MCP server with the command for your client and run the lighthouse prompt. If you want a result in the next sixty seconds, open Picasso IA and test the same prompt there. Run both, compare the outputs, and keep whichever fits your workflow, or use each for what it does best.
Three prompts to try on PicassoIA Image, each short enough to paste:
A fishing boat at sunrise, low-angle, wet wood texture, 35mm lens, soft mist, film grain
A ceramic studio shelf in morning window light, 85mm lens, shallow depth of field
A cyclist crossing a stone bridge at dusk, aerial view, warm light, subtle motion blur
Browse every model, from image and video generators to language models, at picassoia.com/en/all-models, and start creating your own images on Picasso IA today.