Large Language ModelsGenerate imagesGenerate videos

Replicate MCP Server: Run Models From Claude, Claude Code and Cursor

Add the official Replicate MCP server to Claude Desktop, Claude Code and Cursor, then search models and run image, video and text predictions without leaving your chat or editor. Includes exact commands, JSON config, token safety, costs and a zero-setup alternative on PicassoIA.

Replicate MCP Server: Run Models From Claude, Claude Code and Cursor
Cristian Da Conceicao
Founder of Picasso IA

You can now ask Claude to find an image model on Replicate, run it with your prompt, and hand back the finished file without leaving the chat. That is the whole point of the Replicate MCP server. It plugs Replicate's HTTP API into any client that speaks the Model Context Protocol, so Claude Desktop, Claude Code and Cursor can search thousands of models and create predictions on your behalf.

The setup takes about five minutes. What slows people down are the small details: which URL to paste, which transport to pick, where the token goes, and what shows up on the bill afterward. Below you get the exact commands for each client, a token safety checklist, a look at the experimental code mode, and a zero-setup route for readers who only want finished images and videos.

What the Replicate MCP Server Does

Replicate hosts a large catalog of image, video, audio and text models behind a single API. The official MCP server wraps that API as tools, so an assistant can call it directly instead of you juggling curl commands and browser tabs. According to Replicate's docs, it supports all of the operations in Replicate's HTTP API: searching, listing and inspecting models, then creating and fetching predictions.

Overhead view of a tidy oak desk with a laptop, notebook and espresso cup

How a Prompt Becomes a Prediction

Say you type: make a 16:9 photo of a lighthouse at dawn. The client forwards your message to the language model, the model decides it needs an image tool, and it calls the server. The server relays the request to Replicate, which runs the chosen model on its GPUs and returns an output URL. The sequence usually looks like this:

  1. The assistant searches for a model that fits the task.
  2. It reads that model's input options, such as the prompt field and the aspect ratio.
  3. It creates a prediction with your prompt and settings.
  4. It fetches the result and shows you the output link.

💡 Tip: If you already know the model you want, name it in the prompt. That skips the search step and keeps results repeatable from one session to the next.

The same loop works for video and audio models. Ask for a five second clip and the assistant picks a video model, fills in the duration field, and polls until the run finishes. Video runs take longer than images, so expect a wait of a minute or more instead of a few seconds.

The Tools You Get

Replicate's MCP page names three tools, and the docs list the underlying operations:

Tool or operationWhat it does
search_modelsFinds models by task or description
create_predictionsStarts a run with the inputs you supply
list_hardwareShows the hardware available for running models
models.list, models.getLists models and returns one model's metadata
predictions.getFetches the status and output of a run

Hardware matters because it sets both speed and price, so list_hardware is worth a try before a big batch.

Close-up of a graphics card on a clean workbench with a screwdriver nearby

Hosted or Local Setup

You can run the server in two ways, and the choice decides how authentication works.

Hosted serverLocal server
Address or commandhttps://mcp.replicate.comnpx -y replicate-mcp
Sign inBrowser approval flowREPLICATE_API_TOKEN environment variable
Runs whereOn Replicate's sideOn your machine, needs Node
Docs verdictRecommended for most usersAlternative with manual token setup

Choose the hosted server unless you have a reason to run a process on your own machine. It removes the token file from your setup entirely, and it is the path every example in the next sections uses first.

⚠️ Heads up: Older tutorials tell you to install mcp-replicate. That was a community experiment, and its README says it is no longer in active development. The official server replaced it, so skip the old package.

Hosted Server at mcp.replicate.com

The hosted server is the one Replicate recommends. Clients connect to the SSE endpoint, https://mcp.replicate.com/sse, and you finish with a web-based approval step where you paste a Replicate API token. Create one at replicate.com/account/api-tokens before you start and keep it on your clipboard.

Local Server With npx

The local option runs the same server as a process on your machine. Set REPLICATE_API_TOKEN, start it with npx -y replicate-mcp, and stop it with Ctrl+C. Most clients launch it for you from a config file like this:

{
  "mcpServers": {
    "replicate": {
      "command": "npx",
      "args": ["-y", "replicate-mcp"],
      "env": {
        "REPLICATE_API_TOKEN": "your-token-here"
      }
    }
  }
}

VS Code with GitHub Copilot reads .vscode/mcp.json and uses a top-level servers object instead of mcpServers, so copy the block and rename that one field. If the server never shows up, check that Node is installed and that npx runs from a normal terminal.

Connect It to Claude Desktop

Claude Desktop is the quickest route because the hosted server needs no config file at all.

Woman typing a message on a laptop in a sunlit cafe next to a latte

Add the Custom Connector

  1. Open Settings and go to Connectors.
  2. Click Add Custom Connector.
  3. Enter the server URL: https://mcp.replicate.com/sse.
  4. Click Connect and paste your Replicate API token.
  5. Click Log in and Approve.
  6. Restart Claude Desktop. The tools appear under the slider icon below the chat input.

Prefer the local server? Put the JSON from the previous section into claude_desktop_config.json and restart the app.

First Prompt to Try

Start with something that forces the full search, inspect and run loop:

Search Replicate for a fast photorealistic text-to-image model, show me its input options, then generate a 16:9 photo of a lighthouse at dawn.

Fast FLUX variants such as FLUX Schnell are typical candidates for a request like that, though the search decides what comes up first. If the answer arrives as a link instead of an inline picture, that is normal: the server returns the output URL, and the client decides how to display it.

Once that works, get specific. Name the aspect ratio, the lens, the light direction and the number of variations you want. Vague prompts waste runs, and every run is billed.

Connect It to Claude Code

Claude Code takes one command, and it keeps the connection across projects.

Low-angle view of a developer workstation with a dark terminal window and a warm desk lamp

One Command Install

Install the CLI if you do not have it yet, then add the server:

npm install -g @anthropic-ai/claude-code
claude mcp add replicate https://mcp.replicate.com/sse --transport sse --scope user

The --scope user flag stores the server in your user configuration, so every project you open gets access to it.

Authenticate With /mcp

Launch claude, type /mcp, choose the Replicate entry, and a browser window opens for the same approval flow Claude Desktop uses. Once it reports as connected, ask for a run in plain language.

Because Claude Code works inside your repository, it can do more than show a link. Try this: Generate three hero image options for the landing page, then download each output into public/images. The download step matters more than it looks, and the costs section below explains why.

If the tools do not appear, run /mcp again and check the status line for the server. A server that shows as disconnected usually needs a fresh authentication, not a reinstall.

Connect It to Cursor

Cursor reads MCP servers from a JSON file, so this setup is a few lines of config.

Two developers collaborating at a standing desk in a brick-walled studio

Edit the mcp.json File

Open Cursor Settings, go to Tools & Integrations and click New MCP Server. Cursor opens ~/.cursor/mcp.json, where you add the hosted server through the mcp-remote bridge:

{
  "mcpServers": {
    "replicate": {
      "command": "npx",
      "args": ["-y", "mcp-remote@latest", "https://mcp.replicate.com/sse"]
    }
  }
}

For a single repository, create .cursor/mcp.json in the project instead, and use the local-server JSON from earlier with your token in env. Save the file and the server should appear in the same settings panel, along with its tools.

From there, Cursor's agent can generate placeholder art, product shots or icons while it edits the component that needs them. That is the real payoff: the asset and the code change land in one conversation.

Other Clients That Work

The same server works beyond these three:

  • Claude.ai in the browser
  • Codex CLI
  • Google Gemini CLI
  • Windsurf
  • Chorus
  • VS Code with GitHub Copilot, using .vscode/mcp.json

The pattern never changes: point the client at the hosted URL or launch replicate-mcp locally, then authenticate once.

Tokens, Costs and Code Mode

Three things decide whether this setup stays pleasant: how you handle the token, what each run costs, and how far you push automation.

Brass padlock on a black leather notebook beside a slim laptop

Protect Your Token

  • Create the token at replicate.com/account/api-tokens and treat it like a password.
  • Prefer the hosted flow when you can. The token goes through the approval screen instead of sitting in a file in your repository.
  • With the local server, keep the token in an environment variable, or keep the config file outside version control. Add .cursor/mcp.json or .vscode/mcp.json to .gitignore whenever they hold a token.
  • Rotate the token immediately if it shows up in a commit, a screenshot or a shared chat.

What You Pay For

The MCP server is a front door to the same API, so every prediction is billed to your Replicate account like any other API call. The price depends on the model and the hardware it runs on, which is why list_hardware and a quick look at model pages help before a large batch.

A simple habit keeps the bill predictable. Draft with a fast, cheap model such as FLUX Schnell, generate four options, pick the best composition, and rerun only that one prompt on a higher-quality model. Tell the assistant this rule once at the start of a session and it will follow it for every request after that.

Overhead view of a budget notebook, calculator and coins beside a laptop

ItemWhat happens
A prediction you runBilled to your Replicate account
API output filesRemoved after an hour by default
API input files and logsRemoved after an hour by default
Predictions made on the websiteKept indefinitely

💡 Tip: Output URLs from API predictions disappear after about an hour. Ask the assistant to download or upload every result right away, especially in a long session.

Code Mode in a Sandbox

Replicate also ships an experimental code mode. Instead of calling one tool at a time, the model writes TypeScript and runs it inside a Deno sandbox. That suits jobs that chain several predictions, such as generating six variations and keeping the best two. It needs Deno installed. Add it to Claude Code with:

claude mcp add "replicate-code-mode" --scope user --transport stdio -- npx -y replicate-mcp@alpha --tools=code

Developer at a whiteboard filled with hand-drawn flowchart boxes and arrows

The @alpha tag tells you what to expect. Treat it as a testing ground, keep the regular server installed next to it, and check its output before you rely on it for production work.

A Zero-Setup Alternative on PicassoIA

If you only want the finished image or clip and not an afternoon of configuration, PicassoIA keeps hundreds of models on browsable pages with no tokens to paste. Pair it with an LLM such as Claude Sonnet 5 to draft prompts, then generate in the browser.

Photographer sorting printed test shots at a long table in front of a wall of prints

How to Use PicassoIA Image

PicassoIA Image is the platform's own text-to-image model, and it is a fast place to test a prompt before you automate anything.

  1. Open the PicassoIA Image model page.
  2. Write a prompt with a subject, a setting, a light direction and a lens, for example a lighthouse at dawn, low-angle, soft golden light from the left, 35mm lens, film grain.
  3. Choose the aspect ratio you need, such as 16:9 for blog headers.
  4. Generate, review the result, and tweak one detail at a time.
  5. Send the winner to PicassoIA Image Editor Pro for edits such as replacing an object or fixing a detail.
  6. Animate it with PicassoIA Video, which takes text or an image, or try Seedance 2.5 Lite for clips up to 10 seconds.

💡 Tip: Change one variable per run. When the result improves, you know which edit caused it.

PicassoIA API and MCP

PicassoIA also exposes a Replicate-style REST API, so the create, poll and fetch pattern you just set up will feel familiar:

  • Base URL: https://api.picassoia.com/v1, authenticated with a Bearer token from your account
  • Create a run: POST /v1/models/{owner}/{name}/predictions
  • Check a run: GET /v1/predictions/{id}
  • Cancel a run: POST /v1/predictions/{id}/cancel

Four models are available through the API and the MCP connection: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite. Each account gets 5 concurrent predictions, shared across tokens and MCP connections, and prompts can run up to 4,000 characters. API and MCP access depend on your plan, so check the pricing page before you build on it, and read the docs at picassoia.com/en/api.

Replicate MCP serverPicassoIA API and MCP
CatalogThousands of modelsFour models
Sign inReplicate API tokenBearer token from your PicassoIA account
Billed toYour Replicate accountYour PicassoIA plan
Best fitWide model choice and experimentsA fast, fixed set for images and video

Make Your First Image Today

Pick the route that matches your day. If you live in a terminal or an editor, connect the Replicate MCP server with the command for your client and run the lighthouse prompt. If you want a result in the next sixty seconds, open Picasso IA and test the same prompt there. Run both, compare the outputs, and keep whichever fits your workflow, or use each for what it does best.

Three prompts to try on PicassoIA Image, each short enough to paste:

  • A fishing boat at sunrise, low-angle, wet wood texture, 35mm lens, soft mist, film grain
  • A ceramic studio shelf in morning window light, 85mm lens, shallow depth of field
  • A cyclist crossing a stone bridge at dusk, aerial view, warm light, subtle motion blur

Browse every model, from image and video generators to language models, at picassoia.com/en/all-models, and start creating your own images on Picasso IA today.

Share this article