Large Language ModelsGenerate imagesGenerate videos

fal.ai MCP Server: Generate Images in Claude Code and Codex

Connect the fal.ai MCP server to Claude Code and Codex in minutes. See the commands for OAuth and token setups, the eleven tools your agent gets, a cost-control prompt, fixes for common errors, and a second MCP option for images and video.

fal.ai MCP Server: Generate Images in Claude Code and Codex
Cristian Da Conceicao
Founder of Picasso IA

You are halfway through a feature in Claude Code, the landing page needs a hero image, and the usual routine is a slog: open a browser tab, pick a model, wait for the render, download the file, rename it, drag it into the repo. The fal.ai MCP server removes that whole detour. Once it is connected, your coding agent can search fal's catalog of 1,000+ generative models, read a model's input schema, check the price, run the job and hand back an image URL without leaving the terminal.

This article shows how to connect the server to both Claude Code and Codex, which commands to run, which tools you get, how to keep spending predictable, and what breaks most often. It also shows where PicassoIA's own MCP connector fits if you want a second back end for the same workflow. The commands follow fal's published setup pages and each client's MCP documentation. Where fal's own pages disagree with each other, I flag it.

What the fal.ai MCP Server Does

MCP, the Model Context Protocol, is the open standard that lets an AI client call outside tools. The fal server is a hosted endpoint, so there is nothing to install, build or keep running on your machine. Your agent calls the tools, and fal runs the models on its own GPUs. fal's docs state the billing plainly: you only pay for the model runs you trigger, at the same pricing as direct API calls.

That matters for image work in particular. Instead of hard-coding one model into a script, you let the agent pick from the catalog, ask what inputs the model accepts and check what it will cost, all in plain English.

The Tools Your Agent Gets

ToolWhat it does
search_modelsSearches the catalog by topic or category
get_model_schemaReads a model's input and output parameters
get_pricingChecks the price before a run
search_docsSearches fal's documentation
recommend_modelSuggests models for a specific task
run_modelRuns a model and waits, 45 seconds by default
submit_jobStarts a long job without waiting
check_jobReports the status of a job
get_job_resultFetches the output of a finished job
cancel_jobStops a queued or running job
upload_fileUploads a file to fal's CDN for use as a model input

💡 Tip: fal's blog announcement lists nine tools, while the current docs page lists eleven. Expect the list to keep shifting, and ask your agent "which fal tools can you see?" right after connecting.

Why MCP Beats Raw API Calls

  • Schema first: the agent reads each model's parameters before it sends a request, so invalid inputs get caught before they cost anything.
  • Price first: get_pricing turns "how much will this run?" into a question the agent answers before it acts.
  • No glue code: no SDK install, no script, no JSON payload to write by hand.
  • Chaining: one session can write a prompt, render the image, then refine the prompt based on the result.
  • One catalog: images, video, audio, 3D and upscaling sit behind the same handful of tools.

Top-down view of a walnut desk with a laptop, a notebook diagram of three connected boxes and a cup of coffee

Two Ways to Connect

fal documents two routes, and they use different URLs. Pick one per machine rather than adding both under the same name.

The OAuth Relay Route

The docs point to https://mcp.fal.ai/mcp-relay, which speaks Streamable HTTP and signs you in through the browser. You never paste a token into a config file or a chat. This is the better choice for a laptop where you can open a browser window.

The Bearer Token Route

fal's blog post describes https://mcp.fal.ai/mcp, where you send your fal API token in an Authorization: Bearer header. It suits servers, containers and CI, where no browser is available. fal says the token is never stored on its side, but treat it like a password anyway: keep it in an environment variable and never in a repository.

RouteURLSign-inBest for
OAuth relayhttps://mcp.fal.ai/mcp-relayBrowser sign-inLaptops and desktops
Bearer tokenhttps://mcp.fal.ai/mcpAuthorization headerServers, CI, headless machines

💡 Tip: after connecting either route, send a harmless test prompt: "Use fal to search for image generation models. Do not run a model." A list of search results proves that sign-in and tools both work, and it costs nothing.

Low-angle close-up of a developer's hands resting on a laptop in soft daylight

Set Up in Claude Code and Codex

Claude Code Commands

Claude Code adds remote servers with claude mcp add. Both routes take a single command. For the OAuth route:

claude mcp add --transport http fal https://mcp.fal.ai/mcp-relay

Open Claude Code and run /mcp. Select fal and finish the browser sign-in. This matches fal's own wording: add the remote server, then authenticate with /mcp.

For the bearer token route:

export FAL_TOKEN="paste-your-fal-token-here"

claude mcp add --transport http fal-ai https://mcp.fal.ai/mcp \
  --header "Authorization: Bearer $FAL_TOKEN"

The variable name FAL_TOKEN is only a label, so use any name you like. Add --scope user to make the server available in every project, or --scope project to write it into a shared .mcp.json. If a team shares that file, reference the variable there instead of pasting the token.

To check the connection, list your servers:

claude mcp list

Inside a session, /mcp shows every server and its status. If fal appears as connected, ask for the tool list. If it asks to authenticate, run the sign-in step again.

Man in a sunlit home office leaning toward a monitor that shows a plain terminal window

Codex Commands

Codex keeps MCP settings in ~/.codex/config.toml, and trusted projects can add their own .codex/config.toml. You can edit the file by hand or use the codex mcp command family. The one-command route looks like this:

codex mcp add fal --url https://mcp.fal.ai/mcp-relay
codex mcp login fal

The login step is the one fal's docs call out for Codex: add the server, then run codex mcp login fal. Codex changes quickly, so if a flag gets rejected, run codex mcp add --help for the current syntax.

The config.toml route does the same job by hand:

[mcp_servers.fal]
url = "https://mcp.fal.ai/mcp-relay"

For the bearer token route, point Codex at an environment variable instead of writing the token into the file:

[mcp_servers.fal_token]
url = "https://mcp.fal.ai/mcp"
bearer_token_env_var = "FAL_TOKEN"

Export FAL_TOKEN in your shell before you start Codex. After that, codex mcp list should show the server.

Claude Code vs Codex at a Glance

StepClaude CodeCodex
Add the serverclaude mcp add --transport httpcodex mcp add --url
Sign in/mcp inside a sessioncodex mcp login fal
Config location.claude.json or .mcp.json~/.codex/config.toml
Project rules fileCLAUDE.mdAGENTS.md
List serversclaude mcp listcodex mcp list

Two developers at a long birch table in a bright coworking space, one pointing at a laptop

Your First Image Request

A Prompt That Works

Start specific, and make the agent show its work before it spends anything:

Use fal to find a fast photorealistic text-to-image model. Show me its price and input schema, wait for my OK, then generate one 16:9 image of a quiet harbor at dawn.

A well-behaved session runs search_models, then get_model_schema and get_pricing, pauses for your approval, and only then calls run_model. fal's own docs tell assistants to show the estimated cost and ask for approval before generating, so this flow matches the intended design.

Save the result in your repo. The tool returns a URL. Ask for the next step in the same breath: "Download the image to public/images/harbor.jpg with curl and reference it in the hero component." A local file means your page does not depend on a remote link staying alive.

Hand holding a printed photograph of a misty mountain lake above an open laptop

Short Jobs and Long Jobs

run_model waits up to 45 seconds by default, which suits most image models. Slower work, such as video or heavy upscaling, belongs in the queue:

  1. submit_job starts the work and returns at once.
  2. check_job reports the status.
  3. get_job_result fetches the output when the job finishes.
  4. cancel_job stops a job you started by mistake.

Tell the agent which mode you want. "Submit this as a job and check it every 20 seconds" works well for video.

Brass pocket watch lying open on a walnut desk beside a laptop trackpad

Keep Spending Predictable

Price First, Run Second

Put the rule where the agent reads it every session: CLAUDE.md for Claude Code, AGENTS.md for Codex.

fal.ai rules:
- Call get_pricing before every run_model or submit_job.
- Show the estimated cost and wait for my approval when it is above $0.50.
- Never generate more than four images per request without asking.

Read the schema once. get_model_schema lists a model's inputs: aspect ratio, number of images, seed, guidance. When the agent reads it first, you avoid failed requests caused by a wrongly guessed parameter name. Ask the agent to save the working settings in your project notes so the next session skips the lookup.

Mind the Concurrency Limits

fal states that the MCP server respects the same concurrency limits as direct API calls. If you ask for twelve variations at once, expect some of them to queue. Batches of three or four finish faster and are easier to review.

Woman's hand writing figures in a notebook next to a pocket calculator and a folded receipt

Fixes for Common Errors

SymptomLikely causeFix
No fal tools appearThe session started before the server was addedRestart Claude Code or Codex, then check /mcp or codex mcp list
Browser sign-in never finishesThe OAuth step was skippedRun /mcp in Claude Code or codex mcp login fal in Codex
Unauthorized error on the token routeThe variable is empty or the header is malformedExport FAL_TOKEN again and confirm the header starts with Bearer
A job times outrun_model stops waiting after 45 secondsSwitch to submit_job, then poll with check_job
Tools list logs and apps instead of modelsThe Platform MCP was added by mistakeRemove it for image work and add the main fal server
Image comes back with the wrong shapeThe aspect ratio was left at the defaultAsk the agent to read get_model_schema and set the ratio explicitly

Alternatives to the Hosted Server

The hosted server is not the only way to reach fal from an agent, and fal is not the only back end worth wiring in.

Community Servers on GitHub

ServerToolsWhere it runsNotable
raveenb/fal-mcp-server18Your machine or DockerMIT license, STDIO and HTTP/SSE, Claude Code plugin install
wynandw87/claude-code-fal_ai-mcp22Your machine with NodeVideo, lipsync, face swap, 3D and music tools

The first one installs as a Claude Code plugin:

/plugin install fal-ai@raveenb/fal-mcp-server

Community servers run code on your computer with your fal token in the environment, so read the source before you add one. The hosted server avoids that risk because fal runs it.

The Read-Only Platform MCP

fal also ships a separate Platform MCP at https://api.fal.ai/v1/mcp/platform. It is strictly read-only and its 16 tools are aimed at running your account: serverless apps, request history, logs and analytics. It uses a different authorization scheme from the main server, so never reuse the header between them. It is not an image tool, but you can connect both at the same time.

Four cameras lined up on a pale wooden bench, from a vintage rangefinder to a medium format body

PicassoIA MCP as a Second Option

If you want the same agent workflow on a different back end, PicassoIA runs its own MCP connector. Inside Claude it exposes nine tools: generate_image, edit_image, generate_video_picassoia, generate_video_seedance, get_generation, list_generations, cancel_generation, list_models and get_account.

Jobs are asynchronous. A generate call returns a predict_id as soon as a GPU accepts the job, and you poll get_generation after the suggested delay until the status reads succeeded or failed. The connector serves four models: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite, the last two producing video. An account allows five concurrent predictions, shared across every connection.

💡 Tip: the connector's own description says generations on PicassoIA's GPU models are free on the Infinite and Wonder plans. Check the pricing page for what your plan includes before you build a workflow around it.

Models Worth Calling by Task

Whichever server you use, the right model depends on the job. These are the ones I would try first, all available in the PicassoIA catalog:

TaskModels to try
Fast draftsFlux Schnell, P Image
Photorealistic stillsFlux 2 Pro, Imagen 4 Ultra
Text inside imagesGPT Image 2
Sharp 4K outputNano Banana Pro, Seedream 4.5
Editing an existing imageFlux Kontext Pro, PicassoIA Image Editor Pro
Short video from a stillSeedance 2.5 Lite, Kling v3 Video

The text model matters too, because it writes the prompt your image model receives. Claude Sonnet 5 and GPT 5.6 Sol both sit in the PicassoIA catalog, so you can test prompt writing side by side before you commit to one.

Which route fits you? This table puts the three options side by side:

OptionHostingStrengthPick it when
Hosted fal MCPfal1,000+ models, price and schema checksYou want the widest catalog with nothing to install
Community fal serverYour machineExtra tools such as lipsync and 3DYou want local control and can review the code
PicassoIA connectorPicassoIAFour first-party models plus async pollingYou want a small, focused image and video toolset

Bright creator studio with a wall of pinned photographic prints and a woman reviewing a laptop

Your Turn to Generate

Connecting a server takes five minutes. Choosing a model well takes a few experiments, and that part is more fun in a browser. Open Picasso IA, browse the full model list, and run one prompt through two or three models side by side. Start with PicassoIA Image, listed as an unlimited text-to-image generator, then try Flux 2 Pro and Seedream 4.5 on the same wording.

Once you know which model gives you the look you want, bring that choice back to Claude Code or Codex and write it into your rules file. Your agent then stops guessing, and every hero image in your next project starts from a model you picked on purpose.

Share this article