Large Language ModelsGenerate speechGenerate videosGenerate images

AI API Pricing Comparison 2027: Images, Video, Voice and LLMs

Real list prices for LLM tokens, image generations, video seconds and voice characters, side by side. See what a support bot, 1,000 images, 100 video clips and a ten-minute narration cost, plus the hidden fees, retries and discounts that change the total.

AI API Pricing Comparison 2027: Images, Video, Voice and LLMs
Cristian Da Conceicao
Founder of Picasso IA

Picking an AI API by its headline price is how budgets get blown. A low per-token rate means little when your workload burns ten times the tokens, and a cheap per-second video price changes fast once you add 1080p, audio and a few retries. This AI API pricing comparison puts real list prices for language models, image generators, video models and text-to-speech side by side, then turns them into monthly bills you can check against your own workload.

The short version: LLM prices span a 100x range from the cheapest tier to the priciest, one image model swings 35x between its low and high quality settings, video runs from $0.02 to $0.60 per second, and voice goes from half a cent to ten cents per 1,000 characters. Rates below were checked in October 2026 against provider rate cards and public pricing trackers. Treat every number as a snapshot and confirm it on the provider's own page before you commit a budget.

Developer's hands typing at a tidy desk beside a warm brass lamp

How API Billing Actually Works

Four modalities, four different meters. Mixing them up is the most common reason a "cheap" API turns expensive.

Four Units, Four Surprises

  • LLMs bill per million tokens. Output tokens cost roughly 5 to 6 times more than input tokens, so long answers drive the bill, not long prompts.
  • Images bill per picture, per megapixel or per token. Quality and resolution settings can multiply the price by 35x on the same model.
  • Video bills per second of output. Resolution tiers and audio change the rate, and some models bill per token, so any per-second figure is only an estimate.
  • Voice bills per 1,000 characters or per second of generated audio. It is the cheapest category, but the model tier still swings cost by 20x.

Discounts That Actually Exist

Four levers show up across almost every provider:

  1. Batch processing cuts token prices by 50% for work that can wait. Anthropic documents it plainly, and OpenAI offers the same halving.
  2. Prompt caching bills repeated context at about 10% of the normal input price. On Claude Opus 5.5 and Claude Sonnet 5.5 a cache read drops to 5%, and on Claude Fable 5.1 to 2.5%.
  3. Promotional windows are real but temporary. GPT 5.6 Sol sits at $4 input and $20 output through at least November 21, 2026, and ElevenLabs v4 carries a 72% discount that ends October 12. Budget at the regular price.
  4. Data residency works the other way. Pinning Claude inference to the US adds a 1.1x multiplier on every token category.

Top-down view of a calculator, receipt and notebook of handwritten sums on a walnut desk

LLM API Prices Side by Side

Language models have the widest price spread of any category here, and the middle of the market has quietly converged.

The Rate Card

Prices are USD per million tokens at standard, non-batch rates.

ModelInputOutputNotes
Claude Haiku 5.5$0.10$0.50Prompts up to 100k tokens
GPT 5.6 Luna$0.20$1.20Fast tier
Gemini 3.5 Flash$1.50$9.00Fast tier
Claude Sonnet 5$2.00$10.00Sonnet 5.5 costs the same
GPT 5.6 Terra$2.00$12.00Mid tier
Gemini 3.1 Pro$2.00$12.00$4 / $18 above 200k-token prompts
GPT 5.6 Sol$4.00$20.00Promo rate through Nov 21
Claude Opus 5.5$4.00$20.00Top reasoning tier
Claude Fable 5$10.00$50.00Most expensive here

Anthropic and Google figures come from their own pricing pages. OpenAI figures come from public trackers, because OpenAI's page was not readable when I checked, so verify them before you build a forecast.

💡 Compare output prices first. Three mid-tier models share the same $2 input rate, so the real gap sits in output: $10 for Sonnet 5, $12 for Terra and Gemini 3.1 Pro. On chat-heavy workloads that 20% difference matters more than the input rate.

What a Support Bot Costs

Take a small support assistant that sends 10 million input tokens and receives 2 million output tokens each month.

ModelMonthly cost
Claude Haiku 5.5$2.00
GPT 5.6 Luna$4.40
Gemini 3.5 Flash$33.00
Claude Sonnet 5.5$40.00
GPT 5.6 Terra$44.00
Gemini 3.1 Pro$44.00
GPT 5.6 Sol$80.00
Claude Opus 5.5$80.00
Claude Fable 5.1$200.00

The cheapest and priciest options differ by exactly 100x on the same traffic. Most support questions do not need a frontier model, so route routine tickets to a small tier and send only the hard ones upward.

Caching changes the picture again. If 80% of those input tokens are repeated instructions served from cache at Sonnet 5.5's $0.10 per million read rate, input spend falls from $20 to about $4.80 and the total from $40 to roughly $25, before one-time cache write costs. Batch pricing stacks on top, so a job that can wait overnight lands near $12.

Four colleagues gathered at a whiteboard of sticky notes in a bright office

Image Generation Costs Per Picture

Image APIs look cheap until you notice how many quality settings sit behind one model name.

Per-Image Price Table

ModelPrice per imageNotes
GPT Image 2$0.006 low, $0.053 medium, $0.211 high1024x1024 list prices
Nano Banana 2$0.0336 at 1K, $0.0504 at 2K, $0.113 at 4KGoogle lists the 2.1 update at these rates
Seedream 5 ProAbout $0.035Tracker estimate
FLUX.2 proAbout $0.03 per megapixelTracker estimate
Nano Banana 2 LiteAbout $0.01Tracker estimate

The Quality Tier Trap

Run 1,000 images a month and the same OpenAI model costs $6 at low quality, $53 at medium and $211 at high. Nano Banana 2 at 2K lands near $50, Seedream 5 Pro near $35, and Nano Banana 2 Lite near $10. The setting you leave on default decides the bill more than the model you choose.

Resellers add another layer. One gateway listed GPT Image 2 at exactly double OpenAI's rate for every tier, which turns a $211 job into a $422 job without any change in output.

💡 Draft low, finish high. Generate concepts at the cheapest tier, pick winners, then re-render only those at full quality. A 10% finalist rate at high quality still costs far less than rendering everything at high.

Close-up of a mirrorless camera lens with printed photographs blurred behind it

Video API Pricing Per Second

Video is where a wrong guess costs real money, because every second is billed and every retry repeats the full price.

Rate Card With Audio

ModelPrice per secondNotes
P Video$0.02 at 720p, $0.04 at 1080pAudio included
Veo 3.1 Lite$0.05 at 720p, $0.08 at 1080pAudio included
Grok Imagine Video 1.5$0.08Up to 1080p
Veo 3.1 Fast$0.10 at 720p, $0.12 at 1080p, $0.30 at 4KAudio included
Kling v3 VideoAbout $0.11Official Turbo tier, 720p
HappyHorse 1.1About $0.125 at 720p, $0.22 at 1080pJoint audio and video
LTX 2.3 Fast$0.06 to $0.241080p up to 4K
Veo 3.1$0.40 at 720p and 1080p, $0.60 at 4KTop quality tier

Veo, Grok and P Video rates are published by their makers. Kling and HappyHorse figures are converted from yuan, so currency rounding applies.

What 100 Clips Cost

Assume 100 clips of 8 seconds at 720p, or 800 seconds of footage.

  • P Video: $16
  • Veo 3.1 Lite: $40
  • Grok Imagine Video 1.5: $64
  • Veo 3.1 Fast: $80
  • Kling 3.0 Turbo: $88
  • HappyHorse: $100
  • Veo 3.1: $320

Those totals assume every clip is usable. At a 30% keep rate you need about 333 generations to end up with 100 good clips, so the Veo 3.1 Fast line grows from $80 to roughly $267.

Three traps sit behind those numbers. Seedance 2.0 and its Fast variant bill per token, so per-second prices are estimates, and resellers quote anywhere from about $0.09 to $0.25. Jumping to 4K multiplies cost by 1.5x to 3x. And resellers often run 2x to 3x above official rates.

💡 Check availability before you budget. OpenAI's Sora 2 API was scheduled to shut down on September 24, 2026. Sora 2 still appears in model catalogs, so confirm where it runs before you build a pipeline on it.

Low-angle view of a cinema camera on a tripod in a small studio beside a softbox

Voice and Speech Pricing

Speech is the cheapest modality, so quality and language support usually matter more than price.

Per 1,000 Characters Compared

ModelPriceNotes
Inworld TTS (earlier line)$0.005 standard, $0.01 MaxPer million: $5 and $10
Gemini Flash TTSAbout $0.00225 per 10 secondsGoogle's newest Flash tier, billed on audio tokens
ElevenLabs Flash and Turbo$0.04Low-latency tier
MiniMax Speech 2.8 Turbo$0.06Per million: $60
ElevenLabs v3$0.08Expressive tier
ElevenLabs v2 Multilingual$0.0830+ languages
MiniMax Speech 2.8 HD$0.10Per million: $100

A ten-minute narration runs about 9,000 characters. That costs roughly $0.09 on Inworld's Max tier, $0.14 on Gemini Flash TTS, $0.36 on ElevenLabs Flash, $0.54 on MiniMax Turbo, $0.72 on ElevenLabs v3 and $0.90 on MiniMax HD. A hundred such narrations on the priciest option total $90. Audio-token billing and per-character billing do not convert perfectly, so test with your own script length.

ElevenLabs v4 lists at $0.08 per 1,000 characters and v4 Turbo at $0.04, with a 72% launch discount that ends October 12. Inworld's figures come from its earlier TTS-1 line, so check the 1.5 models before relying on them.

Side profile of a woman speaking into a studio condenser microphone with headphones on

Hidden Costs That Break Budgets

List prices are the floor, not the forecast.

Retries, Limits and Long Outputs

  • Rejected outputs. If 20% of generations get thrown away, your effective price rises by 25%, because you pay for five runs to keep four.
  • Output length. Output tokens cost 5 to 6 times more than input, and reasoning tokens count as output. A "thinking" model can inflate a bill without a longer visible answer.
  • Concurrency caps. PicassoIA's API allows 5 predictions in progress at once per account, and many providers ramp limits by spend tier. Waiting in a queue costs engineering time.
  • Reseller markups. Gateways can charge 2x to 3x official rates in exchange for one bill and one integration.

Hands crossing items off a handwritten checklist beside a silver stopwatch

Where the Savings Are

  1. Route by difficulty. Send easy requests to the cheapest tier and escalate only on failure.
  2. Batch anything that can wait. Half price on tokens is the largest single discount available.
  3. Cache static context. System prompts, brand style sheets and reference documents are ideal.
  4. Lower the resolution while iterating. 720p video and low-quality images are enough for drafts.
  5. Cap output length. A maximum token limit prevents runaway generations.

These levers compound. A team that routes most tickets to a small model, caches its instructions and batches its nightly reports can pay a fraction of the list price for the same output, while a team that leaves every default on pays full rate for everything.

Aerial view of a tidy distribution warehouse with rows of wrapped pallets

Test Prices on PicassoIA

Reading a rate card tells you what a model costs. Running it tells you whether the output is worth that cost. PicassoIA hosts more than 120 video models, 75 language models and 24 speech models, plus a deep image catalog, so you can compare tiers side by side before choosing one for production.

Run the Same Prompt Twice

  1. Pick two models from the same category, such as Seedream 5 Pro and FLUX.2 pro.
  2. Open each model page and paste an identical prompt from your real workload.
  3. Match the settings: aspect ratio, resolution and quality.
  4. Compare the results at full size, then note which one you would actually publish.
  5. Divide the provider's per-unit price by your keep rate. The model that wins on cost per keeper is your real choice.

Automate With the API

PicassoIA also exposes a developer API at https://api.picassoia.com/v1, authenticated with a Bearer token that starts with pia_sk_. Jobs are asynchronous: you create a prediction, poll its status and fetch the result.

EndpointPurpose
POST /v1/models/{owner}/{name}/predictionsCreate a prediction
GET /v1/predictions/{id}Check status and fetch output
POST /v1/predictions/{id}/cancelCancel a running job
GET /v1/predictionsList recent predictions

Four models are available through the API: PicassoIA Image, PicassoIA Image Editor Pro, PicassoIA Video and Seedance 2.5 Lite. The documentation currently states that API predictions are free and use no credits, and that an Infinite plan is required. Confirm both on the API page before you plan around them, since plan terms can change.

Run Your Own Price Test

The best AI API is the one that gives you publishable results at a price your volume can carry, and only your own prompts can tell you which one that is. Take one real task from each category, run it through a cheap tier and a premium tier, and write down the cost per keeper. Five minutes of testing beats an afternoon of reading rate cards.

Picasso IA puts those models in one place, so you can generate images, video, voice and text without opening a separate account for each provider. Start with a prompt from your own project, try GPT Image 2 next to PicassoIA Image, and see which result you would actually ship. When you are ready to compare more, browse every option on the all models page.

Creative professional in a sunlit loft studio reviewing finished photographs on a large monitor

Share this article