Three vendors now sell speech-to-speech models over a live socket, and they bill in three different ways. OpenAI meters audio tokens. Google meters audio tokens too, but also prints a per-minute equivalent. xAI skips tokens and charges a flat rate for every minute the connection stays open. That difference alone makes a sticker-price comparison misleading. A voice assistant that costs one cent per minute in a spreadsheet can cost ten times that once real conversations, growing context and tool calls enter the picture.
This breakdown puts the three realtime voice APIs on the same footing. You get the published rates as checked on 8 October 2026, a one-minute worked example, a monthly budget for 10,000 and 100,000 minutes, and the hidden costs that move a bill after launch. At the end there is a cheap way to prototype voices, music and transcripts on PicassoIA before you commit to any of the three.

OpenAI Realtime Pricing
OpenAI's current lineup, released on 6 July 2026, is gpt-realtime-2.1 and gpt-realtime-2.1-mini. Both add configurable reasoning effort on top of speech-to-speech, plus better handling of alphanumeric strings, silence, background noise and interruptions. OpenAI also reported at least a 25% drop in p95 latency across its Realtime models, thanks to improved caching.
gpt-realtime-2.1 Rates
All prices are per 1 million tokens.
| Token type | Input | Cached input | Output |
|---|
| Text | $4.00 | $0.40 | $24.00 |
| Audio | $32.00 | $0.40 | $64.00 |
Image input is billed at $5.00 per million tokens. The previous gpt-realtime-2 carries identical prices, while gpt-realtime-1.5 and the original gpt-realtime keep the same audio rates but charge $16.00 per million text output tokens. The jump to $24.00 arrived with the reasoning-capable 2.x line.
The Mini Alternative
gpt-realtime-2.1-mini costs roughly a third of the flagship on audio.
| Token type | Input | Cached input | Output |
|---|
| Text | $0.60 | $0.06 | $2.40 |
| Audio | $10.00 | $0.30 | $20.00 |
Image input drops to $0.80 per million tokens. A published measurement of about 4,000 real sessions on the mini model found a typical assistant lands between $0.05 and $0.08 per minute, while heavy question-and-answer calls reached $0.12 to $0.15.
Cached Audio Discount
The cheapest line on OpenAI's price sheet is cached audio input: $0.40 instead of $32.00 on the flagship, a cut of about 99%. Caching rewards a stable prefix, meaning your system prompt and the earlier turns of the conversation. Anything that changes the start of the context breaks the discount. The same measurement found that injecting fresh retrieval results mid-call resets the cache and costs roughly $0.007 to $0.009 per event.

Gemini Live Pricing
Google sells its realtime voice models through the Live API, and its price sheet is the easiest to read because it prints both a token rate and a per-minute figure.
Gemini 3.8 Live Rates
gemini-3.8-live and gemini-3.8-live-extended-thinking share one price list on the paid tier.
| Type | Per 1M tokens | Per minute |
|---|
| Text input | $0.75 | n/a |
| Audio input | $3.00 | $0.005 |
| Image or video input | $1.00 | $0.002 |
| Text output | $4.50 | n/a |
| Audio output | $12.00 | $0.018 |
The two columns reconcile at 25 audio tokens per second, so you can budget by time or by tokens and land on the same number. A minute of the assistant talking costs $0.018, and a minute of the caller talking costs half a cent.
💡 Tip: Google is the only one of the three that publishes a per-minute equivalent for token billing. If your finance team wants a rate per call minute, Gemini hands you one without any conversion.

Free Tier and Translate Models
A free tier exists for the Live models, which makes Gemini the natural place to prototype at zero cost before moving to the paid tier. Two neighbors on the same page are worth knowing about:
- Gemini 3.5 Live Translate (preview): speech-to-speech translation across 70+ languages, at $3.50 per million input tokens ($0.0053 per minute) and $21.00 per million output tokens ($0.0315 per minute).
- Gemini 3.5 Transcribe Live: streaming speech-to-text, at $3.50 per million input tokens ($0.005 per minute) and $21.00 per million text output tokens ($0.004 per minute).
The older Gemini 2.5 Flash Native Audio preview is still listed, at $3.00 per million audio input tokens and $12.00 per million audio output tokens.
Grok Voice Pricing
xAI took the opposite approach. There are no audio tokens to count, only minutes.
Flat Rate Per Minute
The Voice Agent API runs grok-voice-think-fast-2.0, also reachable through the alias grok-voice-latest. The xAI pricing page lists it at $0.08 per minute, or $4.80 per hour. Early launch reports quoted $0.05 per minute, so confirm the live page before you build a forecast on either number.
Billing follows connection time. A minute of silence costs the same as a minute of chatter, and there is no context to re-bill on every turn. For a team that dreads surprise invoices, that predictability is the product. Related line items from the same page:
- Speech to text: $0.10 per hour on REST, $0.20 per hour for streaming.
- Text to speech: $15.00 per million characters.
Voices, Languages, Phone Routing
The API supports 20+ languages with native-quality accents, and the speech-to-speech and text-to-speech products share one voice roster. You can also create custom voices from a reference audio clip. Tool support includes function calling, file search, web search, X search and remote MCP servers.
For telephony, xAI documents SIP integration with native G.711, aimed at PSTN, contact-center and PBX traffic. Third-party write-ups add a few numbers that the pricing page did not confirm when I checked: sessions capped at 30 minutes, 100 concurrent sessions per team, a $5 per 1,000 charge for web and X search calls, and an extra $0.01 per minute for a provisioned phone number. Treat those as reported, not guaranteed.

Side by Side Cost Table
Different billing units need one shared scenario. The one below is deliberately simple.
One Minute of Talk
Assumptions: in a 60 second window the caller speaks for 25 seconds and the assistant speaks for 25 seconds. OpenAI audio is counted at about 100 tokens per second, the rate seen in the real-session measurement mentioned earlier. Gemini audio is counted at its published 25 tokens per second. Grok is billed on connection time. Text, tools and growing context are left out on purpose.
| API and model | Billing unit | Cost per minute |
|---|
| OpenAI gpt-realtime-2.1 | Audio tokens | about $0.240 |
| OpenAI gpt-realtime-2.1-mini | Audio tokens | about $0.075 |
| Gemini 3.8 Live | Audio tokens | about $0.010 |
| Grok grok-voice-think-fast-2.0 | Connection time | $0.080 |
The arithmetic for the flagship is 2,500 input tokens at $32.00 per million ($0.080) plus 2,500 output tokens at $64.00 per million ($0.160). The mini model is the same sum at $10.00 and $20.00. Gemini works out to 0.4167 of a minute of caller audio at $0.005 plus 0.4167 of a minute of assistant audio at $0.018.
On audio alone, Gemini is about 25 times cheaper than gpt-realtime-2.1 and about 8 times cheaper than both the mini model and Grok.
💡 Read this table as a floor. It prices audio only. Real sessions add text tokens, reasoning, tool calls and conversation history, which the next section breaks down.
10,000 Minutes a Month
Scaling that floor to a real workload shows where the budget conversation changes.
| API and model | 10,000 minutes | 100,000 minutes |
|---|
| OpenAI gpt-realtime-2.1 | $2,400 | $24,000 |
| OpenAI gpt-realtime-2.1-mini | $750 | $7,500 |
| Gemini 3.8 Live | about $96 | about $960 |
| Grok grok-voice-think-fast-2.0 | $800 | $8,000 |
At 10,000 minutes the spread between the cheapest and the dearest option is more than $2,300 a month. At 100,000 minutes it passes $23,000. At the lower volume, most teams will choose on quality and features instead of price.

Hidden Costs That Inflate Bills
The floor in the table is where good news ends. Four things push real invoices above it.
Context Piles Up
With token billing, each turn can resubmit the conversation so far. At turn 20 you are paying again for 19 earlier exchanges unless caching or history pruning absorbs them. In the 4,000-session measurement, one unpruned call grew to 480,000 tokens and cost $2.05 for a single session. Audio output also dominates the bill, around 80% of spend on the mini model.
Three habits keep this under control:
- Cap response length. Shorter answers cut spend by 20% to 40%.
- Prune history. Summarize old turns instead of replaying them.
- Keep the prefix stable. On OpenAI that is the difference between $32.00 and $0.40 per million audio input tokens.
Grok's flat rate sidesteps the first two entirely, which is the strongest argument for it on long calls.
Tools and Reasoning Tokens
OpenAI's 2.1 models let you set reasoning effort. Higher effort improves hard tool-use turns but adds output tokens and latency, and text output now costs $24.00 per million. Use low effort for FAQ-style calls and spend reasoning only where a wrong answer costs more than a few cents.
Tool calls add their own lines. Retrieval injected mid-call can reset the cache, search calls are billed separately on Grok, and a phone number adds a per-minute fee on telephony setups. If you quote a customer a price for a voice agent, the measurement above suggests $0.15 to $0.20 per minute for the mini model to survive the worst-case sessions.

Which API Fits Your Product
Price is one input. The right pick depends on how long calls run, how many there are, and what a mistake costs.
Phone Agents
Contact-center traffic means long calls, SIP routing and a finance team that hates variance. Grok's connection-time billing and documented SIP support make the invoice easy to forecast. If the calls need heavy reasoning and tool use, such as rebooking, order changes or account checks, gpt-realtime-2.1 is the one built around configurable reasoning, and a $0.24 minute can still beat a failed call.
In-App Assistants
For consumer apps with huge volume and short turns, Gemini 3.8 Live wins on cost by a wide margin, and its free tier removes friction from the prototype stage. gpt-realtime-2.1-mini sits in the middle: three times cheaper than the flagship, with the same API shape.
| Situation | Best fit | Reason |
|---|
| Millions of short voice turns | Gemini 3.8 Live | Lowest audio rates, per-minute figures published |
| Long support calls over SIP | Grok Voice Agent | Flat $0.08 per minute, no context re-billing |
| Complex tool-heavy calls | gpt-realtime-2.1 | Configurable reasoning effort |
| Balanced cost and quality | gpt-realtime-2.1-mini | One third of the flagship audio rate |
| Prototype on a zero budget | Gemini free tier | No charge while you test |

Prototype Voices Before You Pay
A realtime API starts billing with the first second of audio, so settle the voice, the script and the tone somewhere cheaper first. PicassoIA lists 24 text to speech models, including voices from the same vendors you are comparing: Gemini 3.1 Flash TTS and Grok Text To Speech. For latency-minded tests there are also Inworld Realtime TTS 2, Realtime TTS 1.5 Mini with 120ms synthesis, and Realtime TTS 1.5 Max with sub-200ms output.
How to Use Gemini 3.1 Flash TTS
Gemini 3.1 Flash TTS offers 30 voices across 70+ languages, which is enough range to audition a brand voice in a few minutes.
- Open the model page and sign in to PicassoIA.
- Paste your script. Write the way people talk: short sentences, contractions, one idea per line.
- Pick a voice from the list of 30 and set the language of your script.
- Generate and listen. Check numbers, names and email addresses, the same strings that trip up live agents.
- Repeat with the same script in Grok Text To Speech and compare tone side by side.
- Download the audio you prefer and share it with your team before any code is written.
💡 Tip: Keep a fixed test script of 5 to 6 lines that includes a phone number, a price and a proper name. Using the same lines every time makes voice comparisons fair.

Music and Transcripts Round It Out
A voice product rarely needs speech alone. For hold music, an intro jingle or a branded sound bed, Lyria 3 and Music 2.6 generate original tracks from a text prompt. For the other direction, GPT 4o Transcribe, GPT 4o Mini Transcribe and Gemini 3 Pro turn recordings of test calls into text, so you can read what the agent actually said instead of replaying it.

Build Your Own Voice Demo
The numbers above will shift, so rerun the one-minute math with your own call length and your own speaking ratio before you sign off on a vendor. Start the work on the creative side today: write a script, generate a voice, add a music bed, then transcribe the result and read it back. All of it runs on PicassoIA without writing a line of integration code.
Pick a voice model, paste a script, and listen to how each option sounds before the first invoice arrives. When you are ready for more, browse every model on PicassoIA, and keep experimenting with voices, music, transcripts and images until the demo sounds the way you want your product to sound.