Generate speechGenerate musicTranscribe audio

OpenAI TTS Voices: Samples, Options and Which to Pick

Thirteen OpenAI TTS voices, three models and no audio preview. This article compares every voice, shows which model supports it, gives a test script to audition them yourself, breaks down real costs, and recommends starting picks for narration, support bots and short video.

OpenAI TTS Voices: Samples, Options and Which to Pick
Cristian Da Conceicao
Founder of Picasso IA

Thirteen voice names, three models, and no audio preview on the page where you have to pick one. That is the first wall most developers hit with OpenAI's speech endpoint. Alloy, ash, ballad, coral, echo, fable, nova, onyx, sage, shimmer, verse, marin and cedar all look equally reasonable in a dropdown, yet they do not behave the same on every model, and they will not sound the same on your script. This article lays out what each of the OpenAI TTS voices is, which model supports it, how to audition all thirteen in under ten minutes, what the bill looks like, and which voice to start with for narration, support bots and short video. It also points to the speech tools on PicassoIA that sit next to OpenAI's offer, including one with a tutorial you can run today.

What OpenAI TTS Actually Offers

OpenAI's text to speech API turns written text into spoken audio through a single endpoint. You send a model name, a voice name and your text, and you get an audio file back. The part that trips people up is that the model you choose decides how many voices you can use and how much control you get over delivery.

Three Models, Three Price Tags

ModelVoicesPriceBest for
tts-19$15 per 1M charactersFast, low-cost drafts and live playback
tts-1-hd9$30 per 1M charactersHigher fidelity audio rendered ahead of time
gpt-4o-mini-tts13$0.60 per 1M text tokens in, $12 per 1M audio tokens outSteerable tone, pace and accent through instructions

The two older models are priced by character count. gpt-4o-mini-tts is priced in tokens, which means your cost follows the length of the audio rather than the length of the script.

A close-up of a silver studio microphone grille on a wooden desk

Output Formats and Streaming

Audio comes back as MP3 by default, with Opus, AAC, FLAC, WAV and PCM as alternatives. Each has a job:

  • MP3 is the safe default for websites and podcasts.
  • Opus suits streaming over the internet because it stays small.
  • AAC plays well on phones and in video editors.
  • FLAC keeps everything, which matters if you plan to edit the file further.
  • WAV and PCM are uncompressed, so they are the pick when latency matters more than file size.

The endpoint can stream audio as it is generated, so playback can begin before the whole clip is finished. That makes a real difference for voice assistants, where a pause of two seconds feels like a dropped call.

All 13 Voices at a Glance

An overhead view of a desk with thirteen small mugs, headphones and a notebook

Nine voices work on every model. Four more work only on gpt-4o-mini-tts. Here is the full list, with a note on how listeners commonly describe each one.

💡 Read this first: OpenAI does not publish personality labels for its voices, so the "commonly described as" column reflects what developers and listeners tend to say, not an official spec. Treat it as a shortlist and let your own ears make the final call.

The Original Nine

VoiceCommonly described asWorks on
alloyNeutral, balanced, all purposeAll three models
ashSofter, slightly gravelly, conversationalAll three models
coralWarm, friendly, approachableAll three models
echoClear, steady, matter of factAll three models
fableStoryteller feel, expressive, often heard as BritishAll three models
novaBright, energetic, upbeatAll three models
onyxDeep, authoritative, calmAll three models
sageMeasured, thoughtful, even pacedAll three models
shimmerLight, clear, gentleAll three models

These nine have been around longest, so they are the most heavily used and the easiest to find opinions about. If your product already ships with alloy or nova, there is rarely a reason to switch unless the new voices solve a specific problem.

The Four Newer Voices

VoiceCommonly described asWorks on
balladSmooth, melodic, gentlegpt-4o-mini-tts only
verseExpressive, dynamic, livelygpt-4o-mini-tts only
marinNatural, polished, flagship qualitygpt-4o-mini-tts only
cedarNatural, grounded, warmgpt-4o-mini-tts only

OpenAI's own documentation recommends marin or cedar for the best quality, which makes them the natural first test for any new project. Because they only run on gpt-4o-mini-tts, choosing them also means you get the instructions field, which is where the real flexibility lives.

Audition Voices and Steer Delivery

A woman wearing headphones compares audio files on a laptop with a printed script beside her

Audio cannot live inside a text page, so the fastest route to real samples is to generate them yourself. Thirteen clips of the same script take a few minutes and cost a few cents.

A Script That Exposes Weak Voices

A good test script is not a nice paragraph. It is a stress test. Use something like this:

Order 4,817 ships on March 3rd to Dr. Okonkwo at 22 Birchwood Lane. Wait, did you say Thursday? Honestly, I did not expect the delivery to arrive early, but here we are. Three things to remember: bring the receipt, check the seal, and call us if anything feels off.

It packs in a long number, a date, a hard surname, a question, a mild emotional beat and a list. Voices that sound great on a single sentence often stumble on one of those. Then run this loop:

from openai import OpenAI
from pathlib import Path

client = OpenAI()
script = "Order 4,817 ships on March 3rd to Dr. Okonkwo at 22 Birchwood Lane. ..."
voices = ["alloy", "ash", "ballad", "coral", "echo", "fable", "nova",
          "onyx", "sage", "shimmer", "verse", "marin", "cedar"]

for voice in voices:
    out = Path(f"sample_{voice}.mp3")
    with client.audio.speech.with_streaming_response.create(
        model="gpt-4o-mini-tts",
        voice=voice,
        input=script,
        instructions="Speak clearly in a warm, even tone.",
    ) as response:
        response.stream_to_file(out)

Listen with good headphones and score each clip on four things:

  1. Numbers and names: does "4,817" come out as one natural phrase?
  2. The question: does the pitch lift on "did you say Thursday?"
  3. Pacing: do the commas feel like breaths or like stops?
  4. Fatigue: would you tolerate this voice for ten minutes straight?

Instructions That Change the Read

On gpt-4o-mini-tts, an instructions field lets you describe how the voice should sound in plain language. OpenAI notes that the model can be prompted for accent, emotional range, intonation, impressions, speed, tone and whispering. The older tts-1 and tts-1-hd models do not support it.

InstructionWhat to expect
"Speak in a cheerful and positive tone."A lifted, friendly read, good for onboarding
"Calm, slow, like a bedtime story."Softer delivery with longer pauses
"Professional and concise, like a news anchor."Tighter pacing, flatter emotion
"Whisper the final sentence."A hushed ending for dramatic effect
"Sound apologetic but confident."Useful for support scripts about delays

💡 Tip: Keep instructions short and concrete. One clear sentence usually beats a paragraph of adjectives, and changing a single instruction at a time makes it obvious what moved the result.

Which Voice Fits Which Job

The table below is a set of starting points to audition, not rules. Voices are subjective, and your script matters more than any label.

JobFirst pickBackup
Audiobooks and long readscedarsage
Support and phone botsmarincoral
Product demosalloyecho
Ads and short clipsnovaverse
Meditation and calm contentballadshimmer

Narration and Long Reads

An older man listens to an audiobook in an armchair beside a lamp

Long listening punishes voices with a strong personality. A bright voice that charms you for fifteen seconds can wear you down by minute ten. For narration, pick something even and unhurried, then use instructions to add warmth. Split the manuscript at paragraph breaks, because each request has an input length cap and shorter chunks give the voice a fresh start. Check the current API reference for the exact limit.

Assistants and Support Bots

A support agent with a headset speaks toward a dual monitor setup in a bright office

For a bot, the winning traits are clarity and a believable tone under mild stress. Test your voice on the worst messages you send: refunds, delays, outages. A voice that sounds cheerful while delivering bad news feels tone deaf, so write instructions like "calm and sincere" into every request. Pair that with streaming so replies start fast.

Ads and Short Clips

A young man records a vertical video in a kitchen with a phone on a tripod

Short clips reward energy. Lean on nova or verse, add an instruction such as "upbeat, quick, conversational," and keep sentences short so the delivery never has to rush. Render two or three versions and pick by ear, since the difference between a good and a great take is often one changed adjective.

Costs and Limits in Real Numbers

An overhead view of a ledger, calculator and invoices on a wooden desk

Per Character vs Per Token

A 10,000 word manuscript runs roughly 60,000 characters and about 67 minutes of speech at a natural pace. For gpt-4o-mini-tts, OpenAI's launch estimate was around 1.5 cents per minute of audio, which is how the last row below is calculated.

ModelRough cost for 10,000 words
tts-1About $0.90
tts-1-hdAbout $1.80
gpt-4o-mini-ttsAbout $1.00

Those figures are small enough that voice quality should decide the model, not price. The bigger cost is your time re-rendering takes that missed the tone, which is another argument for the instructions field.

Language Support

OpenAI lists support for more than 50 languages, from Afrikaans and Arabic to Czech, Danish and Dutch, with the note that the voices are currently optimized for English. In practice that means a non English script works, but the accent can drift toward English phonetics. Test every language you ship, and have a native speaker listen before launch.

Disclosure and Custom Voices

OpenAI's usage policies require a clear disclosure to end users that the voice they are hearing is AI generated and not a human. Build that line into your product, not into the fine print. Custom voices exist, but they sit behind a separate process that requires speaker consent recordings and audio samples, so plan for lead time if you want a brand voice.

Alternatives on PicassoIA

PicassoIA does not list OpenAI's speech models. What it does offer is a large set of other engines in the Generate speech category, plus tools for the two jobs that surround voice work: checking the result and adding sound underneath it.

Other Speech Engines

ModelStandout trait
Gemini 3.1 Flash TTS30 voices, 70+ language codes, style prompts and expressive tags
MiniMax Speech 2.8 HDStudio quality voiceovers
ElevenLabs V3Natural, expressive AI voiceovers
Inworld TTS 1.5 MaxFast voiceovers in 15 languages
ChatterboxVoice cloning with emotion control
Qwen3 TTSClone a voice or design your own

Check Output With Transcription

A quick way to catch mispronounced names is to run the generated audio back through speech to text and compare it to your script. GPT-4o Transcribe and GPT-4o Mini Transcribe both handle this, and Gemini 3 Pro is a solid third opinion. If the transcript drops or changes a word, a listener probably heard the same thing.

Add Music Under the Voice

A music producer adjusts a fader on a large mixing console in a dim home studio

A voiceover with no bed underneath can feel bare. Generate a track in the Generate music category and mix it low behind the narration. ElevenLabs Music, MiniMax Music 2.6, Lyria 3 Pro and Stable Audio 2.5 all turn a text prompt into a track. Ask for "soft instrumental, no vocals, steady tempo" so nothing competes with the words.

Use Gemini 3.1 Flash TTS on PicassoIA

A laptop, a small condenser microphone and a notebook with a handwritten voice list on a bright desk

Since OpenAI's voices are not hosted there, the closest equivalent is Gemini 3.1 Flash TTS, which gives you a similar workflow with a bigger voice list. Here is how to run it:

  1. Open the model page and find the input form.
  2. Paste your script into the Text field. The limit is 4,000 bytes, so split longer scripts.
  3. Pick a voice from the 30 options. The default is Kore, and names such as Puck, Charon, Fenrir, Aoede and Zephyr are good starting points.
  4. Write a style prompt in the Prompt field. The default is "Say the following." Replace it with something like "Speak slowly with confidence" or "Use a calm, friendly tone."
  5. Set the language code. The default is en-US, and more than 70 codes are available.
  6. Generate, listen and adjust. Change one thing per run so you know what caused the difference.
SettingTip
Text tagsAdd [whispering], [laughing], [shouting], [sigh] or [extremely fast] right before the phrase they should affect
PromptDescribe pace, tone and accent in one sentence
VoiceTest three voices on the same paragraph before committing
Language codeMatch it to the script language, not your own

💡 Tip: Run the finished clip through GPT-4o Transcribe to confirm every name and number came out as written.

Try It Yourself on Picasso IA

Choosing among OpenAI TTS voices is a listening job, and the fastest way to get better at it is to generate clips and compare them. Take the test script above, run it through the thirteen voices, then run the same script through Gemini 3.1 Flash TTS and MiniMax Speech 2.8 HD on Picasso IA. Add a soft music bed, check the transcript, and keep the voice that still sounds right on the tenth listen. Open Picasso IA, paste your own script, and start with one voice and one instruction. Your first usable voiceover is closer than the dropdown makes it look.

Share this article