Large Language ModelsGenerate imagesGenerate videos

fal.ai Pricing: API Costs, Free Credits and MCP

fal.ai bills by output: cents per image, dollars per GPU hour and up to $0.40 per second of video, all paid from prepaid credits. This article lists current rates, explains credit expiry and the 2-request concurrency limit, shows what the MCP server costs, and compares the math with a flat plan.

fal.ai Pricing: API Costs, Free Credits and MCP
Cristian Da Conceicao
Founder of Picasso IA

Most pricing pages tell you what one run costs. The questions that decide your budget come right after: how many runs fit into a month, what happens to free credits you never touch, and whether the MCP server adds a fee on top of the model. This article puts the published fal.ai pricing in one place, checked between September 23 and October 8, 2026, and does the arithmetic so you can price a real project before you spend a cent.

The short version: fal.ai is prepaid and billed per output. Images run from $0.003 per megapixel up to $0.15 per image, video from a few cents to $0.40 per second, GPUs rent by the hour, and the MCP server itself is free. You pay only for the runs it triggers, at the same rates as a direct API call.

💡 Prices move. Treat every figure here as a snapshot. A model page, or the get_pricing tool inside the MCP server, shows the live number before you commit a budget.

How fal.ai Bills You

Prepaid credits, billed per output

fal runs on a prepaid credit model. Its terms say you buy credits in advance, and every run, through the web interface or an API call, deducts its cost from your balance. What you are charged per depends on what the model makes:

OutputBilling unitExample from the pricing page
ImagesPer image or per megapixel$0.0398 per image (Nano Banana)
VideoPer second, or per video$0.14 per second (Kling Video v3)
Language and other modelsPer request or per compute secondVaries by model
Raw computePer GPU hour$2.49 to $4.50 for an H100

Analog electricity meter on a concrete wall, a pay-per-use picture of fal.ai billing

Think of it as a utility meter, not a subscription. Nothing runs, nothing is charged. fal's own wording is that you only pay for the computing power you consume, and the docs add that you pay for successful outputs, never for time spent waiting in the queue.

Concurrency and failed runs

Two rules shape real-world cost more than the headline rate does.

  • Concurrency starts at 2. A new account can have 2 requests in flight at once. The limit rises automatically as you buy credits, up to 40. Anything higher means talking to sales.
  • Errors are mostly free. Server errors (HTTP 500 and above) are never charged. A client error (HTTP 422) may be charged if the GPU had already started work before the problem was detected. Cold start time on Model APIs is not billed, so you pay for inference time only.

Supermarket with two checkout lanes and only one open, a picture of a 2 request concurrency limit

One more rule bites in production. When your balance falls below the account's lock threshold, the account locks and API requests are rejected until you add credits. If an app serves customers, set a top-up reminder well above that line.

Image Prices Per Run

Cheap models, priced per megapixel

Budget image models bill by resolution. A 1024 by 1024 image is about 1.05 megapixels (MP), and a 1920 by 1080 frame is about 2.07 MP.

ModelPriceCost of a 1 MP image
FLUX.1 schnell$0.003 per MPabout $0.003
Z Image Turbo$0.005 per MPabout $0.005
FLUX.2 dev$0.012 per MPabout $0.012
FLUX.2 pro$0.03 for the first MP$0.03 or more

The first two rows come from fal's pricing page. The last two come from a price comparison dated September 23, 2026. At $0.003 per MP, 1,000 square images cost roughly $3.

Premium models, priced per image

Higher-end models skip the megapixel math and charge a flat amount per image. Some add a quality tier on top.

ModelPrice per image
Nano Banana$0.0398
Nano Banana Pro$0.15
GPT Image 2$0.006 low, $0.053 medium, $0.211 high
Seedream 5 Lite$0.035
Seedream 5 Pro$0.0675
Qwen Image 3$0.04 at 1K

Photographer holding a contact sheet of printed frames, a picture of image generation cost per image

Run 1,000 images a month and the spread is dramatic:

💡 GPT Image 2 alone spans about 35 times from its low setting to its high one. Choose the tier per job: drafts on low, finals on high.

Video Prices Per Second

Per-second rates compared

Video is where bills grow, because you pay for every second of every attempt, including the ones you throw away. These are 720p rates from the September 23 comparison, with and without generated audio:

ModelPer second, no audioPer second, with audio
Veo 3.1 Lite$0.03$0.05
Veo 3.1 Fast$0.10$0.15
Veo 3.1$0.20$0.40
Kling v3 Standard$0.084$0.126
Kling v3 Pro$0.112$0.168
Wan 3$0.10not listed
Grok Imagine Video 1.5$0.14not listed
Seedance 2.0 Fastabout $0.24not listed
Seedance 2.0about $0.30not listed

fal's own pricing page lists Kling Video v3 at $0.14 per second and Kling v2.5 at $0.07, which falls in the same range as the Standard and Pro tiers above. Seedance 2.0 is billed by tokens ($0.014 per 1,000 tokens), so every per-second figure for it is an estimate that moves with resolution and duration.

Video editor desk with a paused mountain frame and a stopwatch, a picture of video cost per second

💡 Audio is not free. It doubles the Veo 3.1 rate from $0.20 to $0.40 per second, and raises Veo 3.1 Lite by two thirds. If the soundtrack comes from elsewhere, render silent.

Three monthly budgets

Here is what the rates above mean for real workloads:

  1. 100 clips of 5 seconds on Kling v3. At $0.14 per second, each clip is $0.70, so the month costs $70.
  2. One 8 second clip with audio. On Veo 3.1 that is 8 x $0.40 = $3.20. On Veo 3.1 Lite it is 8 x $0.05 = $0.40, eight times cheaper.
  3. 1,000 clips of 5 seconds with audio on Veo 3.1 Fast. Each clip is $0.75, so the month costs $750.

Retries are the hidden multiplier. If one clip in three is unusable, your real cost per keeper is 50% higher than the table says.

GPU Hourly Rates

When no hosted model fits, fal rents GPUs by the hour. The pricing page shows a range for each card:

GPUMemoryHourly rate
RTX PRO 600096 GB$1.99 to $4.00
H10080 GB$2.49 to $4.50
H200141 GB$2.99 to $6.00
B200192 GB$5.49 to $7.99
GB200192 GB$5.89 to $9.99
B300288 GB$5.99 to $12.99

Quiet data center aisle with rows of server racks, a picture of GPU hourly rates

An 8 hour job on one H100 costs between $19.92 and $36. Leave that same card running around the clock for a 730 hour month and the bill is $1,817.70 at the low end and $3,285 at the high end.

Compare that with per-run pricing. On price alone, an always-on H100 at $2.49 per hour only beats FLUX.1 schnell at $0.003 per image once you pass roughly 600,000 one-megapixel images a month. Below that, per-run billing wins, and it needs no one watching an idle machine.

Free Credits and Expiry

What a new account gets

fal gives new accounts starter credits for testing, but the pricing page and the documentation publish no amount. Third-party pages describe it as a limited number of generations before you add a payment method, and the figure changes with promotions. Plan on the offer being small.

To see how far a dollar goes at the rates above:

  • About 333 one-megapixel images on FLUX.1 schnell
  • About 25 images on Nano Banana
  • 4 clips of 5 seconds with audio on Veo 3.1 Lite
  • About 1.4 clips of 5 seconds on Kling v3

Hand sliding a plain prepaid card across a cafe counter, a picture of fal.ai free credits

💡 Spend trial credits on the model you would really ship, not the cheapest one. A test on a budget model tells you nothing about the cost of a premium one.

When credits expire

The expiry rules differ by credit type, and the two official pages do not agree on promotional credits.

Credit typeExpiryRefundable
Purchased365 days from purchaseNo, per the terms
Free or promotional90 days per the terms, 1 week to 1 year per the FAQNot applicable

Calendar with circled days and an hourglass, a picture of fal.ai credit expiry

The FAQ says free credits and coupons have variable expiration depending on the specific grant. The terms say promotional credits expire in 90 days. Read the date on your own grant in the billing dashboard, and do not buy a large balance you cannot spend inside a year, since purchased credits are consumption-only and non-refundable.

What the fal MCP Server Costs

Connecting it

Nothing. fal states it plainly: the MCP server is free, and you only pay for the model runs you trigger, at standard fal pricing. The same concurrency limits apply as for direct API calls, and the server is stateless, so it stores nothing between requests.

There are two documented ways in. The endpoint https://mcp.fal.ai/mcp takes a bearer token in the Authorization header. The endpoint https://mcp.fal.ai/mcp-relay uses a browser sign-in, so no token is needed. A typical client config looks like this:

{
  "mcpServers": {
    "fal-ai": {
      "url": "https://mcp.fal.ai/mcp",
      "headers": { "Authorization": "Bearer YOUR_FAL_TOKEN" }
    }
  }
}

It works with Claude Code, Claude Desktop, Cursor, Windsurf, ChatGPT and Codex CLI. The assistant behind it can be a model such as Claude Sonnet 5 or Gemini 3.5 Flash. The model you chat with costs separately from the fal runs it triggers.

Woman typing in a sunlit cafe while an assistant answers beside photo thumbnails, a picture of fal.ai MCP use

Tools that cap your spend

The server exposes about nine tools. Four of them matter for cost control:

ToolWhat it does for your budget
get_pricingShows a model's price before you generate
recommend_modelSuggests a model for the task you describe
run_modelRuns a model and waits up to 45 seconds for the result
submit_job and check_jobQueue long jobs and poll them without blocking

An agent can chain a dozen runs from one sentence, and each one is billed. Two habits keep that safe. Tell the assistant to call get_pricing first and ask before any run above a threshold you name. And keep a modest prepaid balance, because a prepaid account is a hard spending ceiling: when the money is gone, the account locks.

A Flat-Rate Alternative

Per-run billing rewards light use and punishes experimentation. The opposite model is a plan with generous or unlimited generations on certain models, and that is how Picasso IA approaches it. Its pricing page lists unlimited generations for PicassoIA Image across its paid plans, unlimited PicassoIA Video on the Elite and Infinite plans, and promotional unlimited use of Seedance 2.5 Lite on those same two.

Designer holding a printed photograph up to window light, a picture of unmetered image generation

Models and limits

PicassoIA also has a developer API and an MCP connector for Claude and ChatGPT. Four models are available through both:

ModelJob
PicassoIA ImageText to image
PicassoIA Image Editor ProEdit or combine 1 to 4 images
PicassoIA VideoText or image to video
Seedance 2.5 LiteText or image to video with audio

An account can have 5 predictions queued or running at once, shared across its API tokens and MCP connections. Request bodies are capped at 10 MB and prompts at 4,000 characters. The API reference says API predictions are currently free and use no credits, but also that creating predictions needs an Infinite plan, while the pricing page lists API Access and MCP Connections on Elite as well. Check the plan table before you build on it.

Calling the API

The flow is create, poll, fetch. Replace the token with your own pia_sk_ value:

curl -s -X POST https://api.picassoia.com/v1/models/picassoia/picassoia-image/predictions \
  -H "Authorization: Bearer $PICASSOIA_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"input": {"prompt": "a lighthouse at sunset, oil painting", "aspect_ratio": "16:9"}}'

The response returns an id that starts with api_, a status of starting, and an eta with a recommended wait. Then poll GET /v1/predictions/{id} until the status reads succeeded, failed or canceled. On success, output holds the image URLs. Optional inputs include num_outputs (1 or 2), output_format (webp, jpg or png) and seed for repeatable results.

Try It Without a Meter

The fastest way to compare is to run the same prompt twice. Take the prompt you were about to send to a metered API, paste it into PicassoIA Image, and see whether the result is good enough to skip the per-image bill. Then send the winner through PicassoIA Image Editor Pro for fixes, and animate it with Seedance 2.5 Lite if you need motion with sound.

Open Picasso IA, pick a model, and make your own images today. If your project needs an API or an MCP connection later, the same models are waiting behind it.

Share this article