Large Language ModelsGenerate imagesGenerate videos
fal.ai Pricing: API Costs, Free Credits and MCP
fal.ai bills by output: cents per image, dollars per GPU hour and up to $0.40 per second of video, all paid from prepaid credits. This article lists current rates, explains credit expiry and the 2-request concurrency limit, shows what the MCP server costs, and compares the math with a flat plan.
Most pricing pages tell you what one run costs. The questions that decide your budget come right after: how many runs fit into a month, what happens to free credits you never touch, and whether the MCP server adds a fee on top of the model. This article puts the published fal.ai pricing in one place, checked between September 23 and October 8, 2026, and does the arithmetic so you can price a real project before you spend a cent.
The short version: fal.ai is prepaid and billed per output. Images run from $0.003 per megapixel up to $0.15 per image, video from a few cents to $0.40 per second, GPUs rent by the hour, and the MCP server itself is free. You pay only for the runs it triggers, at the same rates as a direct API call.
💡 Prices move. Treat every figure here as a snapshot. A model page, or the get_pricing tool inside the MCP server, shows the live number before you commit a budget.
How fal.ai Bills You
Prepaid credits, billed per output
fal runs on a prepaid credit model. Its terms say you buy credits in advance, and every run, through the web interface or an API call, deducts its cost from your balance. What you are charged per depends on what the model makes:
Output
Billing unit
Example from the pricing page
Images
Per image or per megapixel
$0.0398 per image (Nano Banana)
Video
Per second, or per video
$0.14 per second (Kling Video v3)
Language and other models
Per request or per compute second
Varies by model
Raw compute
Per GPU hour
$2.49 to $4.50 for an H100
Think of it as a utility meter, not a subscription. Nothing runs, nothing is charged. fal's own wording is that you only pay for the computing power you consume, and the docs add that you pay for successful outputs, never for time spent waiting in the queue.
Concurrency and failed runs
Two rules shape real-world cost more than the headline rate does.
Concurrency starts at 2. A new account can have 2 requests in flight at once. The limit rises automatically as you buy credits, up to 40. Anything higher means talking to sales.
Errors are mostly free. Server errors (HTTP 500 and above) are never charged. A client error (HTTP 422) may be charged if the GPU had already started work before the problem was detected. Cold start time on Model APIs is not billed, so you pay for inference time only.
One more rule bites in production. When your balance falls below the account's lock threshold, the account locks and API requests are rejected until you add credits. If an app serves customers, set a top-up reminder well above that line.
Image Prices Per Run
Cheap models, priced per megapixel
Budget image models bill by resolution. A 1024 by 1024 image is about 1.05 megapixels (MP), and a 1920 by 1080 frame is about 2.07 MP.
The first two rows come from fal's pricing page. The last two come from a price comparison dated September 23, 2026. At $0.003 per MP, 1,000 square images cost roughly $3.
Premium models, priced per image
Higher-end models skip the megapixel math and charge a flat amount per image. Some add a quality tier on top.
💡 GPT Image 2 alone spans about 35 times from its low setting to its high one. Choose the tier per job: drafts on low, finals on high.
Video Prices Per Second
Per-second rates compared
Video is where bills grow, because you pay for every second of every attempt, including the ones you throw away. These are 720p rates from the September 23 comparison, with and without generated audio:
fal's own pricing page lists Kling Video v3 at $0.14 per second and Kling v2.5 at $0.07, which falls in the same range as the Standard and Pro tiers above. Seedance 2.0 is billed by tokens ($0.014 per 1,000 tokens), so every per-second figure for it is an estimate that moves with resolution and duration.
💡 Audio is not free. It doubles the Veo 3.1 rate from $0.20 to $0.40 per second, and raises Veo 3.1 Lite by two thirds. If the soundtrack comes from elsewhere, render silent.
Three monthly budgets
Here is what the rates above mean for real workloads:
100 clips of 5 seconds on Kling v3. At $0.14 per second, each clip is $0.70, so the month costs $70.
One 8 second clip with audio. On Veo 3.1 that is 8 x $0.40 = $3.20. On Veo 3.1 Lite it is 8 x $0.05 = $0.40, eight times cheaper.
1,000 clips of 5 seconds with audio on Veo 3.1 Fast. Each clip is $0.75, so the month costs $750.
Retries are the hidden multiplier. If one clip in three is unusable, your real cost per keeper is 50% higher than the table says.
GPU Hourly Rates
When no hosted model fits, fal rents GPUs by the hour. The pricing page shows a range for each card:
GPU
Memory
Hourly rate
RTX PRO 6000
96 GB
$1.99 to $4.00
H100
80 GB
$2.49 to $4.50
H200
141 GB
$2.99 to $6.00
B200
192 GB
$5.49 to $7.99
GB200
192 GB
$5.89 to $9.99
B300
288 GB
$5.99 to $12.99
An 8 hour job on one H100 costs between $19.92 and $36. Leave that same card running around the clock for a 730 hour month and the bill is $1,817.70 at the low end and $3,285 at the high end.
Compare that with per-run pricing. On price alone, an always-on H100 at $2.49 per hour only beats FLUX.1 schnell at $0.003 per image once you pass roughly 600,000 one-megapixel images a month. Below that, per-run billing wins, and it needs no one watching an idle machine.
Free Credits and Expiry
What a new account gets
fal gives new accounts starter credits for testing, but the pricing page and the documentation publish no amount. Third-party pages describe it as a limited number of generations before you add a payment method, and the figure changes with promotions. Plan on the offer being small.
💡 Spend trial credits on the model you would really ship, not the cheapest one. A test on a budget model tells you nothing about the cost of a premium one.
When credits expire
The expiry rules differ by credit type, and the two official pages do not agree on promotional credits.
Credit type
Expiry
Refundable
Purchased
365 days from purchase
No, per the terms
Free or promotional
90 days per the terms, 1 week to 1 year per the FAQ
Not applicable
The FAQ says free credits and coupons have variable expiration depending on the specific grant. The terms say promotional credits expire in 90 days. Read the date on your own grant in the billing dashboard, and do not buy a large balance you cannot spend inside a year, since purchased credits are consumption-only and non-refundable.
What the fal MCP Server Costs
Connecting it
Nothing. fal states it plainly: the MCP server is free, and you only pay for the model runs you trigger, at standard fal pricing. The same concurrency limits apply as for direct API calls, and the server is stateless, so it stores nothing between requests.
There are two documented ways in. The endpoint https://mcp.fal.ai/mcp takes a bearer token in the Authorization header. The endpoint https://mcp.fal.ai/mcp-relay uses a browser sign-in, so no token is needed. A typical client config looks like this:
It works with Claude Code, Claude Desktop, Cursor, Windsurf, ChatGPT and Codex CLI. The assistant behind it can be a model such as Claude Sonnet 5 or Gemini 3.5 Flash. The model you chat with costs separately from the fal runs it triggers.
Tools that cap your spend
The server exposes about nine tools. Four of them matter for cost control:
Tool
What it does for your budget
get_pricing
Shows a model's price before you generate
recommend_model
Suggests a model for the task you describe
run_model
Runs a model and waits up to 45 seconds for the result
submit_job and check_job
Queue long jobs and poll them without blocking
An agent can chain a dozen runs from one sentence, and each one is billed. Two habits keep that safe. Tell the assistant to call get_pricing first and ask before any run above a threshold you name. And keep a modest prepaid balance, because a prepaid account is a hard spending ceiling: when the money is gone, the account locks.
A Flat-Rate Alternative
Per-run billing rewards light use and punishes experimentation. The opposite model is a plan with generous or unlimited generations on certain models, and that is how Picasso IA approaches it. Its pricing page lists unlimited generations for PicassoIA Image across its paid plans, unlimited PicassoIA Video on the Elite and Infinite plans, and promotional unlimited use of Seedance 2.5 Lite on those same two.
Models and limits
PicassoIA also has a developer API and an MCP connector for Claude and ChatGPT. Four models are available through both:
An account can have 5 predictions queued or running at once, shared across its API tokens and MCP connections. Request bodies are capped at 10 MB and prompts at 4,000 characters. The API reference says API predictions are currently free and use no credits, but also that creating predictions needs an Infinite plan, while the pricing page lists API Access and MCP Connections on Elite as well. Check the plan table before you build on it.
Calling the API
The flow is create, poll, fetch. Replace the token with your own pia_sk_ value:
curl -s -X POST https://api.picassoia.com/v1/models/picassoia/picassoia-image/predictions \
-H "Authorization: Bearer $PICASSOIA_TOKEN" \
-H "Content-Type: application/json" \
-d '{"input": {"prompt": "a lighthouse at sunset, oil painting", "aspect_ratio": "16:9"}}'
The response returns an id that starts with api_, a status of starting, and an eta with a recommended wait. Then poll GET /v1/predictions/{id} until the status reads succeeded, failed or canceled. On success, output holds the image URLs. Optional inputs include num_outputs (1 or 2), output_format (webp, jpg or png) and seed for repeatable results.
Try It Without a Meter
The fastest way to compare is to run the same prompt twice. Take the prompt you were about to send to a metered API, paste it into PicassoIA Image, and see whether the result is good enough to skip the per-image bill. Then send the winner through PicassoIA Image Editor Pro for fixes, and animate it with Seedance 2.5 Lite if you need motion with sound.
Open Picasso IA, pick a model, and make your own images today. If your project needs an API or an MCP connection later, the same models are waiting behind it.