Generate speechGenerate musicTranscribe audio

Gemini TTS: Voices, Pricing, API and Languages Broken Down

Gemini TTS turns a script into speech with 30 prebuilt voices, 70+ languages and expressive tags like whispering and laughing. This article breaks down every voice by tone, the real token math behind the price, a working Python API call and a browser route that needs no setup.

Gemini TTS: Voices, Pricing, API and Languages Broken Down
Cristian Da Conceicao
Founder of Picasso IA

A one minute voiceover from Gemini TTS costs about three cents at the standard rate. That single number explains why developers, podcasters and course creators keep asking about it. Google's speech models turn a script into audio with 30 named voices, dozens of languages and a plain-English way to direct the performance. What the demo clips don't show is everything that decides whether it fits your project: which voice suits which job, how the token bill adds up, what the API call looks like and where the limits sit. Below you'll find each of those with real numbers, a working Python example and a no-code route you can try in your browser in under a minute.

What Gemini TTS Actually Does

Gemini TTS is the speech output side of Google's Gemini models. You send text plus an optional instruction, and the model returns audio instead of more text. Classic text to speech engines apply the same fixed delivery to every sentence. A Gemini TTS model reads the instruction and the script together, so "say this like a tired night shift nurse" changes pacing, breath and tone, not just pitch.

A podcast host in a gray hoodie speaking into a boom arm microphone in a warm home studio

From Text to Spoken Audio

Every request has the same three ingredients: a script, a voice and a style instruction. The Gemini 3.1 Flash TTS page on Picasso IA exposes them as four fields: text, voice, prompt and language_code. The defaults are the voice Kore, the language en-US and the prompt "Say the following." Press generate and a finished audio file arrives in seconds. The two sample runs published on the model page took about 6.5 and 4.0 seconds. A loop that short changes how you work: rewrite a line, regenerate it and compare takes the way you would with a human narrator, without the booking fee or the studio hours.

Model Versions in Play

Google ships more than one TTS model, and the choice changes your bill:

  • Gemini 3.1 Flash TTS (preview): priced at $1.00 per million text input tokens and $20.00 per million audio output tokens.
  • Gemini 2.5 Flash Preview TTS: exactly half the price of 3.1 Flash on both sides.
  • Gemini 2.5 Pro Preview TTS: the same rates as 3.1 Flash, with no free tier.

Which one should you pick? Choose 3.1 Flash when you want the newest delivery control. Choose 2.5 Flash when you voice thousands of short strings, such as app notifications, where half the price adds up fast. Choose 2.5 Pro only if you prefer its output, since it costs the same as 3.1 Flash and has no free tier.

Google's speech documentation also lists newer Flash and Flash Lite TTS models. Their rates were not in the pricing table when this article was written, so check the table before you budget around them.

The 30 Voices, Sorted by Tone

Gemini TTS ships 30 prebuilt voices with names borrowed from astronomy and myth. Google attaches a one word label to each, such as "Bright" or "Firm", and those labels are the fastest shortcut for choosing. Treat them as a starting point, because a style prompt can push any voice warmer, faster or flatter.

Overhead close-up of a sound engineer's hands on a mixing console with brushed aluminum faders

Bright and Upbeat

Seven voices lean energetic: Zephyr (Bright), Puck (Upbeat), Autonoe (Bright), Laomedeia (Upbeat), Sadachbia (Lively), Fenrir (Excitable) and Leda (Youthful). Reach for them in ads, social clips, onboarding tours and kids' content, anywhere the listener should feel momentum in the first second.

Warm and Gentle

Six voices sit on the softer side: Sulafat (Warm), Achird (Friendly), Vindemiatrix (Gentle), Achernar (Soft), Enceladus (Breathy) and Aoede (Breezy). They suit meditation scripts, bedtime stories, wellness apps and customer support messages.

Firm and Clear

Ten voices project authority: Kore (Firm), Orus (Firm), Alnilam (Firm), Iapetus (Clear), Erinome (Clear), Charon (Informative), Rasalgethi (Informative), Sadaltager (Knowledgeable), Schedar (Even) and Pulcherrima (Forward). Kore is the default for a reason: it works for tutorials, explainers and product demos without sounding stiff.

The remaining seven fill the gaps. Algieba, Despina, Callirrhoe, Umbriel and Zubenelgenubi are smooth, easy-going or casual, which suits conversation. Algenib (Gravelly) and Gacrux (Mature) bring texture for characters and storytelling.

JobVoices to audition first
Product demoKore, Charon, Sadaltager
Podcast introPuck, Fenrir, Algenib
MeditationVindemiatrix, Achernar, Enceladus
Social adZephyr, Sadachbia, Leda
Audiobook narrationGacrux, Schedar, Sulafat

💡 Audition before you commit. Paste one 30 word paragraph, render it with three voices and listen on the speakers your audience will use. The labels are Google's description. Your ears make the final call.

70+ Languages and Accents

The language field accepts more than 70 languages, and many come in regional flavors. That matters more than it sounds: Mexican and Castilian Spanish differ in vocabulary and rhythm, and listeners notice within a sentence.

Five colleagues around an oak table in a bright translation office comparing scripts in different languages

LanguageCodes available
Englishen-US, en-GB, en-AU, en-IN
Spanishes-ES, es-MX, es-419
Frenchfr-FR, fr-CA
Portuguesept-BR, pt-PT
Chinesecmn-CN, cmn-tw
Arabicar-001, ar-EG

The workflow is simple. Write the script in the target language, set the matching code and keep the voice. In most cases one voice can read every language in the list, so you hold a single brand sound across markets. Other popular codes include de-DE, it-IT, ja-JP, ko-KR, hi-IN, tr-TR, pl-PL, nl-NL, sv-SE and vi-VN.

Rarer Languages in the List

Beyond the big markets, the list reaches Cebuano (ceb-PH), Haitian Creole (ht-HT), Javanese (jv-JV), Malagasy (mg-MG), Maithili (mai-IN), Konkani (kok-IN), Sindhi (sd-IN), Basque (eu-ES), Galician (gl-ES) and even Latin (la-VA). Quality will not be identical across all of them, so ask a native speaker to review anything customer facing. Before you localize a whole course, render a 20 second sample in the target language with two voices and send it to a native speaker with three questions: do the names sound right, does the rhythm feel natural, and does the accent match the region you target? If acronyms trip the model, spell them out phonetically in the script. If your source is a finished video rather than a script, ElevenLabs Dubbing translates it into 90+ languages on the same platform.

Tags and Style Prompts

Two levers shape the delivery. Tags act on single phrases, and the style prompt sets the mood for the whole take.

Close-up of an actor whispering into a pop filter in a vocal booth with a marked-up script

Inline Tags for Single Lines

Tags go in square brackets inside the script. The model page lists [sigh], [laughing], [whispering], [shouting] and [extremely fast], and the published samples show that descriptive tags work too. One sample opens with [like dracula] on the Algenib voice. Another runs [laughs] I did NOT expect that. [sigh] Can you believe it! on Callirrhoe.

TagWhat to expect
[whispering]Hushed, close to the microphone
[laughing]An audible laugh before or inside the line
[shouting]Raised volume and urgency
[sigh]A tired or relieved exhale
[extremely fast]A rapid, compressed pace

Style Prompts for Whole Takes

The prompt field takes plain language up to 4,000 bytes. Short instructions like "Say this in a calm, professional tone" or "Speak with excitement and energy" already move the result. Longer ones build a scene: "You are having a casual conversation with a friend. Say the following in a friendly and amused way." That second form is what the published Callirrhoe sample uses.

Here is how the two levers combine on a product announcement, voiced with Puck. The prompt reads "Speak like a friendly presenter on a morning show, upbeat but not rushed." The script reads Big news today. [whispering] We're launching the new dashboard tomorrow. [shouting] And it's free for every team! The prompt sets the energy ceiling, the whisper creates a beat of intimacy, and the shout lands harder because it follows a quiet line. Contrast is what makes tags audible.

💡 One emotion per sentence. Stacking three tags in a line produces a muddle. Place one tag, listen, then decide if the line needs another.

What Gemini TTS Really Costs

Audio is billed in tokens, and Google's pricing page spells out the conversion: 25 audio tokens per second. Everything else follows from that one figure.

A freelance creator working out a budget in a notebook beside a laptop and a solar calculator

ModelText input per 1M tokensAudio output per 1M tokensFree tier
Gemini 3.1 Flash TTS Preview$1.00$20.00Yes
Gemini 2.5 Flash Preview TTS$0.50$10.00Yes
Gemini 2.5 Pro Preview TTS$1.00$20.00No

Token Math in Plain Numbers

One minute of speech is 60 seconds times 25 tokens, so 1,500 output tokens. At $20.00 per million, that is 1,500 times 20 divided by 1,000,000, or $0.03 per minute. The text you send is a rounding error: 150 spoken words is roughly 200 input tokens, about $0.0002.

Audio length3.1 Flash standard3.1 Flash batch2.5 Flash standard
1 minute$0.03$0.015$0.015
10 minutes$0.30$0.15$0.15
1 hour$1.80$0.90$0.90
10 hours$18.00$9.00$9.00

Free Tier and Batch Discounts

Google lists a free tier for both Flash models, while Pro is paid only. Rate limits on free tiers change, so confirm them on the pricing page before you build a launch plan around them. Batch mode, meaning asynchronous jobs that finish later, cuts both rates in half. That makes it the obvious pick for audiobooks, course libraries or any back catalog you can render overnight.

A worked example makes it concrete. Say you voice a course with 20 lessons of 8 minutes each. That is 160 minutes of audio, which costs 160 times $0.03, or $4.80, at the standard 3.1 Flash rate and $2.40 in batch. Stretch it to 100 lessons and you land at $24.00 standard or $12.00 in batch. Re-recording one lesson after a script change costs about 24 cents, which is the real advantage: editing audio stops being a scheduling problem.

💡 Budget rule of thumb. Multiply your finished audio minutes by $0.03 for the standard 3.1 Flash rate. Half of that for batch. These rates come from preview models, so recheck them when you scale.

Calling the Gemini TTS API

With the Python SDK, a TTS request is a normal generate_content call with two changes: an audio response modality and a speech config that names the voice.

Over-the-shoulder view of a developer typing beside a monitor with a dark code editor

A Minimal Python Example

from google import genai
from google.genai import types
import wave

client = genai.Client()  # reads your credentials from the environment

response = client.models.generate_content(
    model="gemini-3.1-flash-tts-preview",
    contents="Say warmly: Welcome back. Today we compare three voices.",
    config=types.GenerateContentConfig(
        response_modalities=["AUDIO"],
        speech_config=types.SpeechConfig(
            voice_config=types.VoiceConfig(
                prebuilt_voice_config=types.PrebuiltVoiceConfig(voice_name="Kore")
            )
        ),
    ),
)

pcm = response.candidates[0].content.parts[0].inline_data.data

with wave.open("welcome.wav", "wb") as f:
    f.setnchannels(1)
    f.setsampwidth(2)
    f.setframerate(24000)
    f.writeframes(pcm)

A few details save debugging time:

  • Credentials stay out of the file. The client picks them up from an environment variable, so nothing secret sits in your repository.
  • The style instruction is part of the prompt text. "Say warmly:" before the script does the same job as the Picasso IA prompt field.
  • The output is raw PCM. It arrives as 24 kHz, mono, 16-bit audio. The wave module wraps it into a playable WAV file, and ffmpeg converts that to MP3 if you need smaller files.
  • Field names evolve on preview models. Check Google's current reference if a call fails after an SDK update.

Two Speakers in One Call

The API accepts up to two speakers in a single request. Label the lines in your script, for example "Host:" and "Guest:", then map each label to a voice in the speech config. It is a fast route to podcast style dialogue or a two voice FAQ. Pair a firm voice with a lively one so listeners can tell turns apart without visuals. For three or more voices, render each speaker's lines separately and join them in an audio editor.

For long scripts, split by paragraph, send the chunks one at a time and cache every result. A failed request then costs you one paragraph, not the whole chapter. Preview models usually carry tighter rate limits than stable ones, so wrap calls in a retry with exponential backoff. Log the voice, language code and a hash of the script next to each file, and you can regenerate any clip later with the same settings.

Try Gemini TTS in Your Browser

If you only need a handful of clips, skip the code. The Picasso IA model page presents Gemini 3.1 Flash TTS as free to try online with no setup, and every field maps to the API options above.

A content creator typing a script on a laptop in a bright workspace with headphones around the neck

Step by Step in the Browser

  1. Open the Gemini 3.1 Flash TTS page on Picasso IA.
  2. Paste your script into Text. The limit is 4,000 bytes, enough for a short explainer. It counts bytes, not characters, so Japanese or Arabic scripts fit fewer characters than English.
  3. Choose a Voice. Start with Kore, then audition two others.
  4. Set the Language code that matches your script, such as es-MX for Mexican Spanish.
  5. Write a Prompt describing tone and pace, or leave the default "Say the following." for a neutral read.
  6. Add tags like [whispering] where you want emotion, press generate and download the audio.
  7. For longer pieces, split the script by section and join the files afterward.

Settings Worth Changing

FieldDefaultTip
TextNoneUp to 4,000 bytes. Use tags sparingly
VoiceKore30 options. Audition three before choosing
PromptSay the following.Up to 4,000 bytes. Set tone, pace and accent
Language codeen-USMatch the language the script is written in

💡 A flat take usually means a short prompt. Lengthen the prompt, not the script. Say who is speaking, to whom and why, and the pacing and emphasis tend to follow.

Pair It With Music and Transcripts

Speech is rarely the whole project. For a background bed, try Lyria 3 or Music 2.6 and mix it quietly under the narration. For captions, run the audio through Gemini 3 Pro or GPT 4o Transcribe.

A journalist with round glasses transcribing an interview at a cafe table with a laptop and recorder

Transcription also works as a quality check. Transcribe your generated narration and compare it with the script. Mispronounced product names and skipped words show up as text differences long before a listener complains.

Make Your First Voiceover

Gemini TTS earns its place in a handful of jobs: product demo narration, podcast intros and ad reads, audio versions of articles and newsletters, training videos with a consistent tone, and multilingual tracks built from one script. If your project needs a different sound, three other models in the same category are worth a test: MiniMax Speech 2.8 HD for studio quality voiceovers, ElevenLabs V3 for natural narration and Inworld Realtime TTS 2 when response speed matters most. One more limit worth knowing: both the API and the Picasso IA page work with the 30 prebuilt voices. If you need a voice that sounds like you or your brand, test MiniMax Voice Cloning or Qwen3 TTS instead.

A woman in a mustard cardigan listening to an audiobook in a rainy window reading nook

The fastest way to settle the voice question is to hear it. Take one paragraph from your own project, open Gemini 3.1 Flash TTS on Picasso IA, and render it with three voices. Add one [whispering] tag, change the language code to see how your script sounds in a second market, and compare the clips side by side. Within ten minutes you will know which voice is yours, and you can try the same workflow with music and transcription models on Picasso IA whenever the project grows.

Share this article