Gemini TTS: Voices, Pricing, API and Languages Broken Down
Gemini TTS turns a script into speech with 30 prebuilt voices, 70+ languages and expressive tags like whispering and laughing. This article breaks down every voice by tone, the real token math behind the price, a working Python API call and a browser route that needs no setup.
A one minute voiceover from Gemini TTS costs about three cents at the standard rate. That single number explains why developers, podcasters and course creators keep asking about it. Google's speech models turn a script into audio with 30 named voices, dozens of languages and a plain-English way to direct the performance. What the demo clips don't show is everything that decides whether it fits your project: which voice suits which job, how the token bill adds up, what the API call looks like and where the limits sit. Below you'll find each of those with real numbers, a working Python example and a no-code route you can try in your browser in under a minute.
What Gemini TTS Actually Does
Gemini TTS is the speech output side of Google's Gemini models. You send text plus an optional instruction, and the model returns audio instead of more text. Classic text to speech engines apply the same fixed delivery to every sentence. A Gemini TTS model reads the instruction and the script together, so "say this like a tired night shift nurse" changes pacing, breath and tone, not just pitch.
From Text to Spoken Audio
Every request has the same three ingredients: a script, a voice and a style instruction. The Gemini 3.1 Flash TTS page on Picasso IA exposes them as four fields: text, voice, prompt and language_code. The defaults are the voice Kore, the language en-US and the prompt "Say the following." Press generate and a finished audio file arrives in seconds. The two sample runs published on the model page took about 6.5 and 4.0 seconds. A loop that short changes how you work: rewrite a line, regenerate it and compare takes the way you would with a human narrator, without the booking fee or the studio hours.
Model Versions in Play
Google ships more than one TTS model, and the choice changes your bill:
Gemini 3.1 Flash TTS (preview): priced at $1.00 per million text input tokens and $20.00 per million audio output tokens.
Gemini 2.5 Flash Preview TTS: exactly half the price of 3.1 Flash on both sides.
Gemini 2.5 Pro Preview TTS: the same rates as 3.1 Flash, with no free tier.
Which one should you pick? Choose 3.1 Flash when you want the newest delivery control. Choose 2.5 Flash when you voice thousands of short strings, such as app notifications, where half the price adds up fast. Choose 2.5 Pro only if you prefer its output, since it costs the same as 3.1 Flash and has no free tier.
Google's speech documentation also lists newer Flash and Flash Lite TTS models. Their rates were not in the pricing table when this article was written, so check the table before you budget around them.
The 30 Voices, Sorted by Tone
Gemini TTS ships 30 prebuilt voices with names borrowed from astronomy and myth. Google attaches a one word label to each, such as "Bright" or "Firm", and those labels are the fastest shortcut for choosing. Treat them as a starting point, because a style prompt can push any voice warmer, faster or flatter.
Bright and Upbeat
Seven voices lean energetic: Zephyr (Bright), Puck (Upbeat), Autonoe (Bright), Laomedeia (Upbeat), Sadachbia (Lively), Fenrir (Excitable) and Leda (Youthful). Reach for them in ads, social clips, onboarding tours and kids' content, anywhere the listener should feel momentum in the first second.
Warm and Gentle
Six voices sit on the softer side: Sulafat (Warm), Achird (Friendly), Vindemiatrix (Gentle), Achernar (Soft), Enceladus (Breathy) and Aoede (Breezy). They suit meditation scripts, bedtime stories, wellness apps and customer support messages.
Firm and Clear
Ten voices project authority: Kore (Firm), Orus (Firm), Alnilam (Firm), Iapetus (Clear), Erinome (Clear), Charon (Informative), Rasalgethi (Informative), Sadaltager (Knowledgeable), Schedar (Even) and Pulcherrima (Forward). Kore is the default for a reason: it works for tutorials, explainers and product demos without sounding stiff.
The remaining seven fill the gaps. Algieba, Despina, Callirrhoe, Umbriel and Zubenelgenubi are smooth, easy-going or casual, which suits conversation. Algenib (Gravelly) and Gacrux (Mature) bring texture for characters and storytelling.
Job
Voices to audition first
Product demo
Kore, Charon, Sadaltager
Podcast intro
Puck, Fenrir, Algenib
Meditation
Vindemiatrix, Achernar, Enceladus
Social ad
Zephyr, Sadachbia, Leda
Audiobook narration
Gacrux, Schedar, Sulafat
💡 Audition before you commit. Paste one 30 word paragraph, render it with three voices and listen on the speakers your audience will use. The labels are Google's description. Your ears make the final call.
70+ Languages and Accents
The language field accepts more than 70 languages, and many come in regional flavors. That matters more than it sounds: Mexican and Castilian Spanish differ in vocabulary and rhythm, and listeners notice within a sentence.
Language
Codes available
English
en-US, en-GB, en-AU, en-IN
Spanish
es-ES, es-MX, es-419
French
fr-FR, fr-CA
Portuguese
pt-BR, pt-PT
Chinese
cmn-CN, cmn-tw
Arabic
ar-001, ar-EG
The workflow is simple. Write the script in the target language, set the matching code and keep the voice. In most cases one voice can read every language in the list, so you hold a single brand sound across markets. Other popular codes include de-DE, it-IT, ja-JP, ko-KR, hi-IN, tr-TR, pl-PL, nl-NL, sv-SE and vi-VN.
Rarer Languages in the List
Beyond the big markets, the list reaches Cebuano (ceb-PH), Haitian Creole (ht-HT), Javanese (jv-JV), Malagasy (mg-MG), Maithili (mai-IN), Konkani (kok-IN), Sindhi (sd-IN), Basque (eu-ES), Galician (gl-ES) and even Latin (la-VA). Quality will not be identical across all of them, so ask a native speaker to review anything customer facing. Before you localize a whole course, render a 20 second sample in the target language with two voices and send it to a native speaker with three questions: do the names sound right, does the rhythm feel natural, and does the accent match the region you target? If acronyms trip the model, spell them out phonetically in the script. If your source is a finished video rather than a script, ElevenLabs Dubbing translates it into 90+ languages on the same platform.
Tags and Style Prompts
Two levers shape the delivery. Tags act on single phrases, and the style prompt sets the mood for the whole take.
Inline Tags for Single Lines
Tags go in square brackets inside the script. The model page lists [sigh], [laughing], [whispering], [shouting] and [extremely fast], and the published samples show that descriptive tags work too. One sample opens with [like dracula] on the Algenib voice. Another runs [laughs] I did NOT expect that. [sigh] Can you believe it! on Callirrhoe.
Tag
What to expect
[whispering]
Hushed, close to the microphone
[laughing]
An audible laugh before or inside the line
[shouting]
Raised volume and urgency
[sigh]
A tired or relieved exhale
[extremely fast]
A rapid, compressed pace
Style Prompts for Whole Takes
The prompt field takes plain language up to 4,000 bytes. Short instructions like "Say this in a calm, professional tone" or "Speak with excitement and energy" already move the result. Longer ones build a scene: "You are having a casual conversation with a friend. Say the following in a friendly and amused way." That second form is what the published Callirrhoe sample uses.
Here is how the two levers combine on a product announcement, voiced with Puck. The prompt reads "Speak like a friendly presenter on a morning show, upbeat but not rushed." The script reads Big news today. [whispering] We're launching the new dashboard tomorrow. [shouting] And it's free for every team! The prompt sets the energy ceiling, the whisper creates a beat of intimacy, and the shout lands harder because it follows a quiet line. Contrast is what makes tags audible.
💡 One emotion per sentence. Stacking three tags in a line produces a muddle. Place one tag, listen, then decide if the line needs another.
What Gemini TTS Really Costs
Audio is billed in tokens, and Google's pricing page spells out the conversion: 25 audio tokens per second. Everything else follows from that one figure.
One minute of speech is 60 seconds times 25 tokens, so 1,500 output tokens. At $20.00 per million, that is 1,500 times 20 divided by 1,000,000, or $0.03 per minute. The text you send is a rounding error: 150 spoken words is roughly 200 input tokens, about $0.0002.
Audio length
3.1 Flash standard
3.1 Flash batch
2.5 Flash standard
1 minute
$0.03
$0.015
$0.015
10 minutes
$0.30
$0.15
$0.15
1 hour
$1.80
$0.90
$0.90
10 hours
$18.00
$9.00
$9.00
Free Tier and Batch Discounts
Google lists a free tier for both Flash models, while Pro is paid only. Rate limits on free tiers change, so confirm them on the pricing page before you build a launch plan around them. Batch mode, meaning asynchronous jobs that finish later, cuts both rates in half. That makes it the obvious pick for audiobooks, course libraries or any back catalog you can render overnight.
A worked example makes it concrete. Say you voice a course with 20 lessons of 8 minutes each. That is 160 minutes of audio, which costs 160 times $0.03, or $4.80, at the standard 3.1 Flash rate and $2.40 in batch. Stretch it to 100 lessons and you land at $24.00 standard or $12.00 in batch. Re-recording one lesson after a script change costs about 24 cents, which is the real advantage: editing audio stops being a scheduling problem.
💡 Budget rule of thumb. Multiply your finished audio minutes by $0.03 for the standard 3.1 Flash rate. Half of that for batch. These rates come from preview models, so recheck them when you scale.
Calling the Gemini TTS API
With the Python SDK, a TTS request is a normal generate_content call with two changes: an audio response modality and a speech config that names the voice.
A Minimal Python Example
from google import genai
from google.genai import types
import wave
client = genai.Client() # reads your credentials from the environment
response = client.models.generate_content(
model="gemini-3.1-flash-tts-preview",
contents="Say warmly: Welcome back. Today we compare three voices.",
config=types.GenerateContentConfig(
response_modalities=["AUDIO"],
speech_config=types.SpeechConfig(
voice_config=types.VoiceConfig(
prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name="Kore")
)
),
),
)
pcm = response.candidates[0].content.parts[0].inline_data.data
with wave.open("welcome.wav", "wb") as f:
f.setnchannels(1)
f.setsampwidth(2)
f.setframerate(24000)
f.writeframes(pcm)
A few details save debugging time:
Credentials stay out of the file. The client picks them up from an environment variable, so nothing secret sits in your repository.
The style instruction is part of the prompt text. "Say warmly:" before the script does the same job as the Picasso IA prompt field.
The output is raw PCM. It arrives as 24 kHz, mono, 16-bit audio. The wave module wraps it into a playable WAV file, and ffmpeg converts that to MP3 if you need smaller files.
Field names evolve on preview models. Check Google's current reference if a call fails after an SDK update.
Two Speakers in One Call
The API accepts up to two speakers in a single request. Label the lines in your script, for example "Host:" and "Guest:", then map each label to a voice in the speech config. It is a fast route to podcast style dialogue or a two voice FAQ. Pair a firm voice with a lively one so listeners can tell turns apart without visuals. For three or more voices, render each speaker's lines separately and join them in an audio editor.
For long scripts, split by paragraph, send the chunks one at a time and cache every result. A failed request then costs you one paragraph, not the whole chapter. Preview models usually carry tighter rate limits than stable ones, so wrap calls in a retry with exponential backoff. Log the voice, language code and a hash of the script next to each file, and you can regenerate any clip later with the same settings.
Try Gemini TTS in Your Browser
If you only need a handful of clips, skip the code. The Picasso IA model page presents Gemini 3.1 Flash TTS as free to try online with no setup, and every field maps to the API options above.
Paste your script into Text. The limit is 4,000 bytes, enough for a short explainer. It counts bytes, not characters, so Japanese or Arabic scripts fit fewer characters than English.
Choose a Voice. Start with Kore, then audition two others.
Set the Language code that matches your script, such as es-MX for Mexican Spanish.
Write a Prompt describing tone and pace, or leave the default "Say the following." for a neutral read.
Add tags like [whispering] where you want emotion, press generate and download the audio.
For longer pieces, split the script by section and join the files afterward.
Settings Worth Changing
Field
Default
Tip
Text
None
Up to 4,000 bytes. Use tags sparingly
Voice
Kore
30 options. Audition three before choosing
Prompt
Say the following.
Up to 4,000 bytes. Set tone, pace and accent
Language code
en-US
Match the language the script is written in
💡 A flat take usually means a short prompt. Lengthen the prompt, not the script. Say who is speaking, to whom and why, and the pacing and emphasis tend to follow.
Pair It With Music and Transcripts
Speech is rarely the whole project. For a background bed, try Lyria 3 or Music 2.6 and mix it quietly under the narration. For captions, run the audio through Gemini 3 Pro or GPT 4o Transcribe.
Transcription also works as a quality check. Transcribe your generated narration and compare it with the script. Mispronounced product names and skipped words show up as text differences long before a listener complains.
Make Your First Voiceover
Gemini TTS earns its place in a handful of jobs: product demo narration, podcast intros and ad reads, audio versions of articles and newsletters, training videos with a consistent tone, and multilingual tracks built from one script. If your project needs a different sound, three other models in the same category are worth a test: MiniMax Speech 2.8 HD for studio quality voiceovers, ElevenLabs V3 for natural narration and Inworld Realtime TTS 2 when response speed matters most. One more limit worth knowing: both the API and the Picasso IA page work with the 30 prebuilt voices. If you need a voice that sounds like you or your brand, test MiniMax Voice Cloning or Qwen3 TTS instead.
The fastest way to settle the voice question is to hear it. Take one paragraph from your own project, open Gemini 3.1 Flash TTS on Picasso IA, and render it with three voices. Add one [whispering] tag, change the language code to see how your script sounds in a second market, and compare the clips side by side. Within ten minutes you will know which voice is yours, and you can try the same workflow with music and transcription models on Picasso IA whenever the project grows.