Generate speechGenerate musicTranscribe audio

Realtime Voice API Compared: OpenAI, Gemini and Grok Pricing

Realtime voice APIs bill by tokens or by the minute, and the gap is large. This breakdown prices OpenAI gpt-realtime-2.1, Google Gemini 3.8 Live and xAI Grok Voice side by side, with a one-minute example, 10,000 and 100,000 minute budgets, and the hidden costs that inflate bills.

Realtime Voice API Compared: OpenAI, Gemini and Grok Pricing
Cristian Da Conceicao
Founder of Picasso IA

Three vendors now sell speech-to-speech models over a live socket, and they bill in three different ways. OpenAI meters audio tokens. Google meters audio tokens too, but also prints a per-minute equivalent. xAI skips tokens and charges a flat rate for every minute the connection stays open. That difference alone makes a sticker-price comparison misleading. A voice assistant that costs one cent per minute in a spreadsheet can cost ten times that once real conversations, growing context and tool calls enter the picture.

This breakdown puts the three realtime voice APIs on the same footing. You get the published rates as checked on 8 October 2026, a one-minute worked example, a monthly budget for 10,000 and 100,000 minutes, and the hidden costs that move a bill after launch. At the end there is a cheap way to prototype voices, music and transcripts on PicassoIA before you commit to any of the three.

A developer's hands resting beside a calculator and a printed receipt on a wooden desk

OpenAI Realtime Pricing

OpenAI's current lineup, released on 6 July 2026, is gpt-realtime-2.1 and gpt-realtime-2.1-mini. Both add configurable reasoning effort on top of speech-to-speech, plus better handling of alphanumeric strings, silence, background noise and interruptions. OpenAI also reported at least a 25% drop in p95 latency across its Realtime models, thanks to improved caching.

gpt-realtime-2.1 Rates

All prices are per 1 million tokens.

Token typeInputCached inputOutput
Text$4.00$0.40$24.00
Audio$32.00$0.40$64.00

Image input is billed at $5.00 per million tokens. The previous gpt-realtime-2 carries identical prices, while gpt-realtime-1.5 and the original gpt-realtime keep the same audio rates but charge $16.00 per million text output tokens. The jump to $24.00 arrived with the reasoning-capable 2.x line.

The Mini Alternative

gpt-realtime-2.1-mini costs roughly a third of the flagship on audio.

Token typeInputCached inputOutput
Text$0.60$0.06$2.40
Audio$10.00$0.30$20.00

Image input drops to $0.80 per million tokens. A published measurement of about 4,000 real sessions on the mini model found a typical assistant lands between $0.05 and $0.08 per minute, while heavy question-and-answer calls reached $0.12 to $0.15.

Cached Audio Discount

The cheapest line on OpenAI's price sheet is cached audio input: $0.40 instead of $32.00 on the flagship, a cut of about 99%. Caching rewards a stable prefix, meaning your system prompt and the earlier turns of the conversation. Anything that changes the start of the context breaks the discount. The same measurement found that injecting fresh retrieval results mid-call resets the cache and costs roughly $0.007 to $0.009 per event.

A customer support agent in a call center speaking to a caller through a headset

Gemini Live Pricing

Google sells its realtime voice models through the Live API, and its price sheet is the easiest to read because it prints both a token rate and a per-minute figure.

Gemini 3.8 Live Rates

gemini-3.8-live and gemini-3.8-live-extended-thinking share one price list on the paid tier.

TypePer 1M tokensPer minute
Text input$0.75n/a
Audio input$3.00$0.005
Image or video input$1.00$0.002
Text output$4.50n/a
Audio output$12.00$0.018

The two columns reconcile at 25 audio tokens per second, so you can budget by time or by tokens and land on the same number. A minute of the assistant talking costs $0.018, and a minute of the caller talking costs half a cent.

💡 Tip: Google is the only one of the three that publishes a per-minute equivalent for token billing. If your finance team wants a rate per call minute, Gemini hands you one without any conversion.

Two colleagues sharing a laptop and a phone at a bright workspace table

Free Tier and Translate Models

A free tier exists for the Live models, which makes Gemini the natural place to prototype at zero cost before moving to the paid tier. Two neighbors on the same page are worth knowing about:

  • Gemini 3.5 Live Translate (preview): speech-to-speech translation across 70+ languages, at $3.50 per million input tokens ($0.0053 per minute) and $21.00 per million output tokens ($0.0315 per minute).
  • Gemini 3.5 Transcribe Live: streaming speech-to-text, at $3.50 per million input tokens ($0.005 per minute) and $21.00 per million text output tokens ($0.004 per minute).

The older Gemini 2.5 Flash Native Audio preview is still listed, at $3.00 per million audio input tokens and $12.00 per million audio output tokens.

Grok Voice Pricing

xAI took the opposite approach. There are no audio tokens to count, only minutes.

Flat Rate Per Minute

The Voice Agent API runs grok-voice-think-fast-2.0, also reachable through the alias grok-voice-latest. The xAI pricing page lists it at $0.08 per minute, or $4.80 per hour. Early launch reports quoted $0.05 per minute, so confirm the live page before you build a forecast on either number.

Billing follows connection time. A minute of silence costs the same as a minute of chatter, and there is no context to re-bill on every turn. For a team that dreads surprise invoices, that predictability is the product. Related line items from the same page:

  • Speech to text: $0.10 per hour on REST, $0.20 per hour for streaming.
  • Text to speech: $15.00 per million characters.

Voices, Languages, Phone Routing

The API supports 20+ languages with native-quality accents, and the speech-to-speech and text-to-speech products share one voice roster. You can also create custom voices from a reference audio clip. Tool support includes function calling, file search, web search, X search and remote MCP servers.

For telephony, xAI documents SIP integration with native G.711, aimed at PSTN, contact-center and PBX traffic. Third-party write-ups add a few numbers that the pricing page did not confirm when I checked: sessions capped at 30 minutes, 100 concurrent sessions per team, a $5 per 1,000 charge for web and X search calls, and an extra $0.01 per minute for a provisioned phone number. Treat those as reported, not guaranteed.

A woman in a rain jacket speaking into her smartphone while walking down a city street

Side by Side Cost Table

Different billing units need one shared scenario. The one below is deliberately simple.

One Minute of Talk

Assumptions: in a 60 second window the caller speaks for 25 seconds and the assistant speaks for 25 seconds. OpenAI audio is counted at about 100 tokens per second, the rate seen in the real-session measurement mentioned earlier. Gemini audio is counted at its published 25 tokens per second. Grok is billed on connection time. Text, tools and growing context are left out on purpose.

API and modelBilling unitCost per minute
OpenAI gpt-realtime-2.1Audio tokensabout $0.240
OpenAI gpt-realtime-2.1-miniAudio tokensabout $0.075
Gemini 3.8 LiveAudio tokensabout $0.010
Grok grok-voice-think-fast-2.0Connection time$0.080

The arithmetic for the flagship is 2,500 input tokens at $32.00 per million ($0.080) plus 2,500 output tokens at $64.00 per million ($0.160). The mini model is the same sum at $10.00 and $20.00. Gemini works out to 0.4167 of a minute of caller audio at $0.005 plus 0.4167 of a minute of assistant audio at $0.018.

On audio alone, Gemini is about 25 times cheaper than gpt-realtime-2.1 and about 8 times cheaper than both the mini model and Grok.

💡 Read this table as a floor. It prices audio only. Real sessions add text tokens, reasoning, tool calls and conversation history, which the next section breaks down.

10,000 Minutes a Month

Scaling that floor to a real workload shows where the budget conversation changes.

API and model10,000 minutes100,000 minutes
OpenAI gpt-realtime-2.1$2,400$24,000
OpenAI gpt-realtime-2.1-mini$750$7,500
Gemini 3.8 Liveabout $96about $960
Grok grok-voice-think-fast-2.0$800$8,000

At 10,000 minutes the spread between the cheapest and the dearest option is more than $2,300 a month. At 100,000 minutes it passes $23,000. At the lower volume, most teams will choose on quality and features instead of price.

Three white ceramic cups beside stacks of coins of different heights on a kitchen counter

Hidden Costs That Inflate Bills

The floor in the table is where good news ends. Four things push real invoices above it.

Context Piles Up

With token billing, each turn can resubmit the conversation so far. At turn 20 you are paying again for 19 earlier exchanges unless caching or history pruning absorbs them. In the 4,000-session measurement, one unpruned call grew to 480,000 tokens and cost $2.05 for a single session. Audio output also dominates the bill, around 80% of spend on the mini model.

Three habits keep this under control:

  1. Cap response length. Shorter answers cut spend by 20% to 40%.
  2. Prune history. Summarize old turns instead of replaying them.
  3. Keep the prefix stable. On OpenAI that is the difference between $32.00 and $0.40 per million audio input tokens.

Grok's flat rate sidesteps the first two entirely, which is the strongest argument for it on long calls.

Tools and Reasoning Tokens

OpenAI's 2.1 models let you set reasoning effort. Higher effort improves hard tool-use turns but adds output tokens and latency, and text output now costs $24.00 per million. Use low effort for FAQ-style calls and spend reasoning only where a wrong answer costs more than a few cents.

Tool calls add their own lines. Retrieval injected mid-call can reset the cache, search calls are billed separately on Grok, and a phone number adds a per-minute fee on telephony setups. If you quote a customer a price for a voice agent, the measurement above suggests $0.15 to $0.20 per minute for the mini model to survive the worst-case sessions.

A hand holding a long paper receipt that curls over the edge of a wooden table

Which API Fits Your Product

Price is one input. The right pick depends on how long calls run, how many there are, and what a mistake costs.

Phone Agents

Contact-center traffic means long calls, SIP routing and a finance team that hates variance. Grok's connection-time billing and documented SIP support make the invoice easy to forecast. If the calls need heavy reasoning and tool use, such as rebooking, order changes or account checks, gpt-realtime-2.1 is the one built around configurable reasoning, and a $0.24 minute can still beat a failed call.

In-App Assistants

For consumer apps with huge volume and short turns, Gemini 3.8 Live wins on cost by a wide margin, and its free tier removes friction from the prototype stage. gpt-realtime-2.1-mini sits in the middle: three times cheaper than the flagship, with the same API shape.

SituationBest fitReason
Millions of short voice turnsGemini 3.8 LiveLowest audio rates, per-minute figures published
Long support calls over SIPGrok Voice AgentFlat $0.08 per minute, no context re-billing
Complex tool-heavy callsgpt-realtime-2.1Configurable reasoning effort
Balanced cost and qualitygpt-realtime-2.1-miniOne third of the flagship audio rate
Prototype on a zero budgetGemini free tierNo charge while you test

A bakery owner answering a landline phone while boxing pastries at the shop counter

Prototype Voices Before You Pay

A realtime API starts billing with the first second of audio, so settle the voice, the script and the tone somewhere cheaper first. PicassoIA lists 24 text to speech models, including voices from the same vendors you are comparing: Gemini 3.1 Flash TTS and Grok Text To Speech. For latency-minded tests there are also Inworld Realtime TTS 2, Realtime TTS 1.5 Mini with 120ms synthesis, and Realtime TTS 1.5 Max with sub-200ms output.

How to Use Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS offers 30 voices across 70+ languages, which is enough range to audition a brand voice in a few minutes.

  1. Open the model page and sign in to PicassoIA.
  2. Paste your script. Write the way people talk: short sentences, contractions, one idea per line.
  3. Pick a voice from the list of 30 and set the language of your script.
  4. Generate and listen. Check numbers, names and email addresses, the same strings that trip up live agents.
  5. Repeat with the same script in Grok Text To Speech and compare tone side by side.
  6. Download the audio you prefer and share it with your team before any code is written.

💡 Tip: Keep a fixed test script of 5 to 6 lines that includes a phone number, a price and a proper name. Using the same lines every time makes voice comparisons fair.

A podcaster in studio headphones seated at a desk with a condenser microphone

Music and Transcripts Round It Out

A voice product rarely needs speech alone. For hold music, an intro jingle or a branded sound bed, Lyria 3 and Music 2.6 generate original tracks from a text prompt. For the other direction, GPT 4o Transcribe, GPT 4o Mini Transcribe and Gemini 3 Pro turn recordings of test calls into text, so you can read what the agent actually said instead of replaying it.

A young musician playing a compact music controller in a small home studio

Build Your Own Voice Demo

The numbers above will shift, so rerun the one-minute math with your own call length and your own speaking ratio before you sign off on a vendor. Start the work on the creative side today: write a script, generate a voice, add a music bed, then transcribe the result and read it back. All of it runs on PicassoIA without writing a line of integration code.

Pick a voice model, paste a script, and listen to how each option sounds before the first invoice arrives. When you are ready for more, browse every model on PicassoIA, and keep experimenting with voices, music, transcripts and images until the demo sounds the way you want your product to sound.

Share this article