Large Language ModelsGenerate imagesGenerate videos
fal.ai MCP Server: Generate Images in Claude Code and Codex
Connect the fal.ai MCP server to Claude Code and Codex in minutes. See the commands for OAuth and token setups, the eleven tools your agent gets, a cost-control prompt, fixes for common errors, and a second MCP option for images and video.
You are halfway through a feature in Claude Code, the landing page needs a hero image, and the usual routine is a slog: open a browser tab, pick a model, wait for the render, download the file, rename it, drag it into the repo. The fal.ai MCP server removes that whole detour. Once it is connected, your coding agent can search fal's catalog of 1,000+ generative models, read a model's input schema, check the price, run the job and hand back an image URL without leaving the terminal.
This article shows how to connect the server to both Claude Code and Codex, which commands to run, which tools you get, how to keep spending predictable, and what breaks most often. It also shows where PicassoIA's own MCP connector fits if you want a second back end for the same workflow. The commands follow fal's published setup pages and each client's MCP documentation. Where fal's own pages disagree with each other, I flag it.
What the fal.ai MCP Server Does
MCP, the Model Context Protocol, is the open standard that lets an AI client call outside tools. The fal server is a hosted endpoint, so there is nothing to install, build or keep running on your machine. Your agent calls the tools, and fal runs the models on its own GPUs. fal's docs state the billing plainly: you only pay for the model runs you trigger, at the same pricing as direct API calls.
That matters for image work in particular. Instead of hard-coding one model into a script, you let the agent pick from the catalog, ask what inputs the model accepts and check what it will cost, all in plain English.
The Tools Your Agent Gets
Tool
What it does
search_models
Searches the catalog by topic or category
get_model_schema
Reads a model's input and output parameters
get_pricing
Checks the price before a run
search_docs
Searches fal's documentation
recommend_model
Suggests models for a specific task
run_model
Runs a model and waits, 45 seconds by default
submit_job
Starts a long job without waiting
check_job
Reports the status of a job
get_job_result
Fetches the output of a finished job
cancel_job
Stops a queued or running job
upload_file
Uploads a file to fal's CDN for use as a model input
💡 Tip: fal's blog announcement lists nine tools, while the current docs page lists eleven. Expect the list to keep shifting, and ask your agent "which fal tools can you see?" right after connecting.
Why MCP Beats Raw API Calls
Schema first: the agent reads each model's parameters before it sends a request, so invalid inputs get caught before they cost anything.
Price first:get_pricing turns "how much will this run?" into a question the agent answers before it acts.
No glue code: no SDK install, no script, no JSON payload to write by hand.
Chaining: one session can write a prompt, render the image, then refine the prompt based on the result.
One catalog: images, video, audio, 3D and upscaling sit behind the same handful of tools.
Two Ways to Connect
fal documents two routes, and they use different URLs. Pick one per machine rather than adding both under the same name.
The OAuth Relay Route
The docs point to https://mcp.fal.ai/mcp-relay, which speaks Streamable HTTP and signs you in through the browser. You never paste a token into a config file or a chat. This is the better choice for a laptop where you can open a browser window.
The Bearer Token Route
fal's blog post describes https://mcp.fal.ai/mcp, where you send your fal API token in an Authorization: Bearer header. It suits servers, containers and CI, where no browser is available. fal says the token is never stored on its side, but treat it like a password anyway: keep it in an environment variable and never in a repository.
Route
URL
Sign-in
Best for
OAuth relay
https://mcp.fal.ai/mcp-relay
Browser sign-in
Laptops and desktops
Bearer token
https://mcp.fal.ai/mcp
Authorization header
Servers, CI, headless machines
💡 Tip: after connecting either route, send a harmless test prompt: "Use fal to search for image generation models. Do not run a model." A list of search results proves that sign-in and tools both work, and it costs nothing.
Set Up in Claude Code and Codex
Claude Code Commands
Claude Code adds remote servers with claude mcp add. Both routes take a single command. For the OAuth route:
claude mcp add --transport http fal https://mcp.fal.ai/mcp-relay
Open Claude Code and run /mcp. Select fal and finish the browser sign-in. This matches fal's own wording: add the remote server, then authenticate with /mcp.
The variable name FAL_TOKEN is only a label, so use any name you like. Add --scope user to make the server available in every project, or --scope project to write it into a shared .mcp.json. If a team shares that file, reference the variable there instead of pasting the token.
To check the connection, list your servers:
claude mcp list
Inside a session, /mcp shows every server and its status. If fal appears as connected, ask for the tool list. If it asks to authenticate, run the sign-in step again.
Codex Commands
Codex keeps MCP settings in ~/.codex/config.toml, and trusted projects can add their own .codex/config.toml. You can edit the file by hand or use the codex mcp command family. The one-command route looks like this:
The login step is the one fal's docs call out for Codex: add the server, then run codex mcp login fal. Codex changes quickly, so if a flag gets rejected, run codex mcp add --help for the current syntax.
Export FAL_TOKEN in your shell before you start Codex. After that, codex mcp list should show the server.
Claude Code vs Codex at a Glance
Step
Claude Code
Codex
Add the server
claude mcp add --transport http
codex mcp add --url
Sign in
/mcp inside a session
codex mcp login fal
Config location
.claude.json or .mcp.json
~/.codex/config.toml
Project rules file
CLAUDE.md
AGENTS.md
List servers
claude mcp list
codex mcp list
Your First Image Request
A Prompt That Works
Start specific, and make the agent show its work before it spends anything:
Use fal to find a fast photorealistic text-to-image model. Show me its price and input schema, wait for my OK, then generate one 16:9 image of a quiet harbor at dawn.
A well-behaved session runs search_models, then get_model_schema and get_pricing, pauses for your approval, and only then calls run_model. fal's own docs tell assistants to show the estimated cost and ask for approval before generating, so this flow matches the intended design.
Save the result in your repo. The tool returns a URL. Ask for the next step in the same breath: "Download the image to public/images/harbor.jpg with curl and reference it in the hero component." A local file means your page does not depend on a remote link staying alive.
Short Jobs and Long Jobs
run_model waits up to 45 seconds by default, which suits most image models. Slower work, such as video or heavy upscaling, belongs in the queue:
submit_job starts the work and returns at once.
check_job reports the status.
get_job_result fetches the output when the job finishes.
cancel_job stops a job you started by mistake.
Tell the agent which mode you want. "Submit this as a job and check it every 20 seconds" works well for video.
Keep Spending Predictable
Price First, Run Second
Put the rule where the agent reads it every session: CLAUDE.md for Claude Code, AGENTS.md for Codex.
fal.ai rules:
- Call get_pricing before every run_model or submit_job.
- Show the estimated cost and wait for my approval when it is above $0.50.
- Never generate more than four images per request without asking.
Read the schema once.get_model_schema lists a model's inputs: aspect ratio, number of images, seed, guidance. When the agent reads it first, you avoid failed requests caused by a wrongly guessed parameter name. Ask the agent to save the working settings in your project notes so the next session skips the lookup.
Mind the Concurrency Limits
fal states that the MCP server respects the same concurrency limits as direct API calls. If you ask for twelve variations at once, expect some of them to queue. Batches of three or four finish faster and are easier to review.
Fixes for Common Errors
Symptom
Likely cause
Fix
No fal tools appear
The session started before the server was added
Restart Claude Code or Codex, then check /mcp or codex mcp list
Browser sign-in never finishes
The OAuth step was skipped
Run /mcp in Claude Code or codex mcp login fal in Codex
Unauthorized error on the token route
The variable is empty or the header is malformed
Export FAL_TOKEN again and confirm the header starts with Bearer
A job times out
run_model stops waiting after 45 seconds
Switch to submit_job, then poll with check_job
Tools list logs and apps instead of models
The Platform MCP was added by mistake
Remove it for image work and add the main fal server
Image comes back with the wrong shape
The aspect ratio was left at the default
Ask the agent to read get_model_schema and set the ratio explicitly
Alternatives to the Hosted Server
The hosted server is not the only way to reach fal from an agent, and fal is not the only back end worth wiring in.
Community Servers on GitHub
Server
Tools
Where it runs
Notable
raveenb/fal-mcp-server
18
Your machine or Docker
MIT license, STDIO and HTTP/SSE, Claude Code plugin install
wynandw87/claude-code-fal_ai-mcp
22
Your machine with Node
Video, lipsync, face swap, 3D and music tools
The first one installs as a Claude Code plugin:
/plugin install fal-ai@raveenb/fal-mcp-server
Community servers run code on your computer with your fal token in the environment, so read the source before you add one. The hosted server avoids that risk because fal runs it.
The Read-Only Platform MCP
fal also ships a separate Platform MCP at https://api.fal.ai/v1/mcp/platform. It is strictly read-only and its 16 tools are aimed at running your account: serverless apps, request history, logs and analytics. It uses a different authorization scheme from the main server, so never reuse the header between them. It is not an image tool, but you can connect both at the same time.
PicassoIA MCP as a Second Option
If you want the same agent workflow on a different back end, PicassoIA runs its own MCP connector. Inside Claude it exposes nine tools: generate_image, edit_image, generate_video_picassoia, generate_video_seedance, get_generation, list_generations, cancel_generation, list_models and get_account.
Jobs are asynchronous. A generate call returns a predict_id as soon as a GPU accepts the job, and you poll get_generation after the suggested delay until the status reads succeeded or failed. The connector serves four models: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite, the last two producing video. An account allows five concurrent predictions, shared across every connection.
💡 Tip: the connector's own description says generations on PicassoIA's GPU models are free on the Infinite and Wonder plans. Check the pricing page for what your plan includes before you build a workflow around it.
Models Worth Calling by Task
Whichever server you use, the right model depends on the job. These are the ones I would try first, all available in the PicassoIA catalog:
The text model matters too, because it writes the prompt your image model receives. Claude Sonnet 5 and GPT 5.6 Sol both sit in the PicassoIA catalog, so you can test prompt writing side by side before you commit to one.
Which route fits you? This table puts the three options side by side:
Option
Hosting
Strength
Pick it when
Hosted fal MCP
fal
1,000+ models, price and schema checks
You want the widest catalog with nothing to install
Community fal server
Your machine
Extra tools such as lipsync and 3D
You want local control and can review the code
PicassoIA connector
PicassoIA
Four first-party models plus async polling
You want a small, focused image and video toolset
Your Turn to Generate
Connecting a server takes five minutes. Choosing a model well takes a few experiments, and that part is more fun in a browser. Open Picasso IA, browse the full model list, and run one prompt through two or three models side by side. Start with PicassoIA Image, listed as an unlimited text-to-image generator, then try Flux 2 Pro and Seedream 4.5 on the same wording.
Once you know which model gives you the look you want, bring that choice back to Claude Code or Codex and write it into your rules file. Your agent then stops guessing, and every hero image in your next project starts from a model you picked on purpose.