Large Language ModelsGenerate imagesGenerate videos
Flux MCP Server: Flux 2 and Kontext in Claude via Replicate, Setup and Real Prompts
Connect Claude to Replicate's hosted MCP server and run Flux 2 for generation and Flux Kontext for edits without leaving the chat. Setup steps for Claude.ai, Claude Desktop and Claude Code, a model comparison table, working prompts and the mistakes that waste credits.
If you have ever copied a prompt out of Claude, pasted it into a separate image tool, downloaded the result and then gone back to the chat to ask for a tweak, you know the friction. The chat knows what you want. The image model lives in another tab. A Flux MCP server setup closes that gap: Claude writes the prompt, calls the model on Replicate, and returns the finished image URL inside the same conversation.
One clarification first, because search results blur it. There is no standalone product called "the Flux MCP server." What people mean is Replicate's official hosted MCP server, which opens Replicate's model library to Claude, combined with the Black Forest Labs models that run there: the Flux 2 family for generating images and the Flux Kontext family for editing them. This article shows how to connect the two, which model to call for which job, and the prompts that behave well in practice.
What the Flux MCP Server Really Is
One Hosted Server, Thousands of Models
Replicate runs a hosted MCP server at https://mcp.replicate.com/sse. Its setup page says it lets apps like Claude, Cursor, Codex, Gemini and VS Code reach Replicate's HTTP API, so they can find, compare and run models. The tools it lists include search_models, create_predictions and list_hardware. In plain terms, Claude can search the catalog, read what a model accepts as input, and start a prediction with your prompt.
Flux is one corner of that catalog. Black Forest Labs publishes its models on Replicate, so asking Claude for a Flux model uses the same mechanism as asking for any other image model. That is handy: the habits you build here carry over to every other model on the platform.
Why Not Just Write a Script?
You can call Replicate from a Python or Node script, and plenty of people do. The MCP route earns its place in a few situations:
Claude picks the parameters. It reads the model's input schema and fills in aspect ratio, resolution and output format instead of you hunting through docs.
Prompts and results stay together. Every attempt, seed and URL sits in one thread you can scroll back through.
Iteration is conversational. "Same shot, warmer light" is a sentence, not a code change.
Nothing to maintain. The server is hosted, so Claude.ai and Claude Desktop need no local install.
The trade-off is cost and control. Every run is billed to your Replicate account, and Claude decides how many calls to make unless you say otherwise. More on that in the mistakes section.
Connect Replicate to Claude
Start by signing in to Replicate and creating an API token from your account settings. Treat it like a password. The token bills to your account, so never paste it into a prompt or commit it to a repository. The connector flows below ask for it in a dedicated field.
Claude.ai and Claude Desktop
Open Settings and go to Connectors. On the web, that is claude.ai/settings/connectors.
Choose Add Custom Connector.
Enter the server URL https://mcp.replicate.com/sse.
Click Connect and paste your Replicate API token when asked.
In Claude Desktop, restart the app so the new tools load.
Claude Code in One Command
If you work in a terminal, one command registers the server for your user account:
claude mcp add replicate https://mcp.replicate.com/sse --transport sse --scope user
--scope user makes the server available in every project instead of only the folder where you ran the command. You authenticate with the same Replicate API token.
Cursor, Windsurf, Codex and Gemini CLI
Other clients work too, with slightly different plumbing:
Client
How it connects
Cursor
Edit ~/.cursor/mcp.json with an npx -y mcp-remote@latest entry
Windsurf
Edit ~/.codeium/windsurf/mcp_config.json with the same kind of npx entry
Codex CLI
Run codex mcp add replicate, then open /mcp inside codex
Gemini CLI
Add the server URL to ~/.gemini/settings.json, then run /mcp auth replicate
For Cursor and Windsurf, copy the exact JSON block from the Replicate MCP page instead of retyping it from memory. One misplaced bracket and the server silently fails to load.
Check That It Works
Open a fresh chat and ask: "Search Replicate for black-forest-labs/flux-2-pro and list the inputs it accepts." If the connection is live, Claude calls the search tool and reports fields such as prompt, input_images, aspect_ratio, resolution, output_format, seed and safety_tolerance. If Claude says it has no tools for that, restart the client and confirm the connector shows as connected.
Pick the Right Flux Model
Replicate's Black Forest Labs page lists a long run of Flux 2 and Kontext models. You do not need all of them. This table lists the ones worth knowing, with each description taken from Replicate's own wording.
Flux 2 for generating.Flux 2 Pro is the workhorse. It makes images from a text prompt alone, or from a prompt plus up to eight reference images that steer style, subject or composition. Output goes up to 4 MP, with a 2048x2048 ceiling, and you can pick WebP, JPEG or PNG. The Flux 2 Klein 4B variant trades some polish for speed, which makes it the right model for the first rounds of a prompt.
Kontext for editing.Flux Kontext Pro takes an input image and an instruction written in plain language. No masks, no selection tools. It can also generate from scratch if you leave the image field empty, though Flux 2 is the better first pick for that job.
Quick Decision Rules
No source image and you want the best result: Flux 2 Pro, or Flux 2 Max for a final.
The change involves words on a sign, label or poster: Flux Kontext Max.
Ten rounds of prompt tweaking ahead: Flux 2 Klein 4B first, then move up.
Prompts That Work in Claude
Claude does not need special syntax. It needs the model name, the settings you care about and a clear description. Everything below is a pattern, not a rule.
Plain Text-to-Image
A request that tends to work on the first try:
Use Replicate to run black-forest-labs/flux-2-pro with aspect_ratio 16:9, resolution 2 MP and output_format jpg. Prompt: a ceramic teapot on a linen cloth beside a window, soft morning light from the left, 85mm lens, shallow depth of field, fine film grain. Return the image URL.
Naming the model removes guesswork. Stating the aspect ratio and resolution stops Claude from falling back on defaults, which for Flux 2 Pro means a square 1 MP image.
💡 The Flux 2 Pro schema says up to 4 MP is possible but 2 MP or below is recommended, and that very high resolutions may not be honored when the aspect ratio is not 1:1. Start at 1 or 2 MP.
Multi-Reference Prompts
Reference images are where Flux 2 pulls ahead of older models. The trick is to refer to them by number. A prompt from the model's own example gallery:
The person from image 1 is petting the cat from image 2, the bird from image 3 is next to them.
Hand Claude three image URLs, list them in order, and use "image 1", "image 2" and "image 3" in the prompt. For a brand look, pass three to five examples of the target style and describe the new scene in one sentence.
Edits With Kontext
Edit prompts work best when they state the change and what must stay put. This one comes from the Flux Kontext Pro gallery:
Change the background to a beach while keeping the person in the exact same position, scale, and pose. Maintain identical subject placement, camera angle, framing, and perspective. Only replace the environment around them.
Short instructions work as well for small jobs. "Change the car to blue" is a real example on the Flux 2 Pro page, and "make this a photo" appears in the Kontext Pro gallery. Both are one change, clearly named.
Text in the frame. Put the exact words in quotes and keep them short. The Kontext Pro gallery shows a pure replacement prompt: Replace 'joy' with 'Pro'. For longer lettering, Flux Kontext Max is the model Replicate describes as having improved typography. The Flux 2 Pro gallery even includes a labeled infographic of a landmark with measurements, so dense text is possible there too, but check every character before you publish.
Parameters Worth Setting
Claude will fill these in if you leave them out. Setting them yourself saves a retry. The values below come from the Flux 2 Pro and Flux Kontext Pro input schemas.
Defaults to 1:1; also supports 16:9, 3:2, 4:5, 9:16 and more, plus custom
Defaults to match_input_image; 14 presets
resolution
0.5, 1 (default), 2 or 4 MP
Not offered, follows the input
output_format
webp (default), jpg or png
png (default) or jpg
output_quality
0 to 100, default 80, ignored for png
Not offered
seed
Fixed value reproduces a result
Fixed value reproduces a result
safety_tolerance
1 strict to 5 permissive, default 2
Default 2, which is also the maximum when input images are used
prompt_upsampling
Not listed in the schema
Off by default, improves the prompt automatically
Seeds make edits comparable. If you change the prompt and the seed at the same time, you cannot tell which one caused the difference. Ask Claude to reuse the seed from the previous run and change only the wording. It is the closest thing to a controlled experiment you get in image generation.
Let Claude draft, but read the prompt. Claude is good at expanding "a teapot" into a lighting-and-lens description. It is less good at remembering that you wanted no text in the frame. Ask it to show the final prompt before it runs the prediction, especially on the larger models.
A Product Photo Workflow That Holds Up
Here is a sequence that suits a small shop or a freelancer building a catalog. Keep each step in the same chat so the history does the record-keeping for you.
Draft with Klein. Ask for six variations of the product scene on Flux 2 Klein 4B. Pick the composition you like.
Rebuild on Pro. Run the winning prompt on Flux 2 Pro at 2 MP and note the seed.
Edit with Kontext. Use Flux Kontext Pro for one change per round: background, color, lighting.
Log everything. Ask Claude to keep a running table of prompt, model, seed and URL. Your next session will start from facts instead of memory.
When the job is "put this product next to this person," a dedicated tool helps. Multi Image Kontext Max merges two photos into one scene, and it is a useful option when you do not need the full flexibility of eight references.
Because the whole run lives in one thread, a colleague can open it, read the prompts, and repeat or tweak any step. Compare that with a folder of unnamed downloads.
Mistakes That Waste Time
1. Stacking Five Edits in One Prompt
"Change the background, make the shirt red, add glasses, soften the light and crop tighter" asks the model to guess priorities. Split it into five short edits and review each. You will spend more messages but fewer credits on bad results.
2. Running the Biggest Model for Drafts
There is no reason to spend Flux 2 Max money on a prompt you have not settled yet. Use a fast model to find the idea, then spend on the final.
3. Unbudgeted Loops and Loose Tokens
Tell Claude how many generations it may run per request, for example "no more than four." Every prediction bills your Replicate account, and each model's page shows its current price. Check it before you ask for thirty variations.
Keep the token out of the chat too. It belongs in the connector's authentication field, not in a message. If it ever lands in a conversation or a screenshot, revoke it in your Replicate settings and make a new one.
Try Flux on PicassoIA
The Replicate MCP route rewards people who like terminals, schemas and logs. If you would rather skip the setup and look at results first, open a model page and type a prompt. Start with Flux 2 Pro for a fresh image, then take the result into Flux Kontext Pro and change one thing about it: the background, the color of a jacket, the time of day. Two or three rounds are enough to feel how differently the two families think. The model pages also show example prompts and inputs, which is a fast way to see how each one reads instructions.
Prompt quality decides most outcomes, and you do not need a connector to improve it. Claude Sonnet 5 is available on PicassoIA, so you can draft a prompt there and paste it straight into a Flux model page.
A strong still can also become the first frame of a short clip. Image-to-video models such as Wan 2.7 I2V take a photo and animate it, so a Flux render can turn into a product loop or a social teaser without starting over. PicassoIA also has its own developer API and MCP connection for its image and video models, documented at picassoia.com/en/api, if you want one account behind both the browser and your tools.
Pick a scene you actually need, whether that is a product shot, a portrait or a blog header, and make your first image today. Run the same prompt on a fast model and a premium one, and keep whichever result you would be happy to publish.