Generate speechGenerate musicTranscribe audio

OpenAI TTS API Pricing: Voices, tts-1-hd and vs ElevenLabs

OpenAI charges $15 per million characters for tts-1, $30 for tts-1-hd and about $0.015 a minute for gpt-4o-mini-tts. This breakdown lists every voice, works out real costs for an article and an audiobook, and sets the numbers against ElevenLabs.

OpenAI TTS API Pricing: Voices, tts-1-hd and vs ElevenLabs
Cristian Da Conceicao
Founder of Picasso IA

OpenAI's text-to-speech API is one of the cheapest ways to put a voice on an app, a course or a podcast, and the price list fits in one sentence: $15 per million characters for tts-1, $30 per million for tts-1-hd, and about $0.015 per minute of audio for the newer gpt-4o-mini-tts. ElevenLabs charges $0.10 per 1,000 characters on its top models, more than three times the tts-1-hd rate. Whether that gap is worth paying depends on voices, languages and how human the delivery has to sound. This page lays out every OpenAI rate, lists the voices model by model, works through real cost examples and sets the numbers beside ElevenLabs so you can choose before writing a line of code.

💡 Quick math: one minute of narration is roughly 900 characters. That costs about $0.014 on tts-1, $0.027 on tts-1-hd and $0.09 on ElevenLabs v3 or Multilingual v2.

What OpenAI Charges Per Character

OpenAI bills its two classic speech models by input characters, not by seconds of audio. You pay for the text you send, so a slow, dramatic read costs the same as a fast one. The tts-1-hd model page lists the speech endpoint (v1/audio/speech) as its only supported route, and the same endpoint serves tts-1 and gpt-4o-mini-tts.

tts-1 and tts-1-hd Rates

The public price list is short. There is no seat fee and no minimum spend: you add credit to your account and each request draws it down.

ModelPricePer 1,000 charactersBest for
tts-1$15 per 1M characters$0.015Low latency apps and assistants
tts-1-hd$30 per 1M characters$0.030Finished, replayed audio
gpt-4o-mini-tts$0.60 per 1M text tokens plus $12 per 1M audio tokensToken based, about $0.015 per minuteSteerable, expressive delivery

Overhead view of a walnut desk with a calculator, a blank invoice, coins and a condenser microphone, used to budget text-to-speech costs

The HD model costs exactly double. For a typical script that is still pocket change, which is why the real decision is rarely about money. The tts-1-hd model page lists request limits between 2,500 and 10,000 per minute depending on your usage tier, far above what most apps ever send.

gpt-4o-mini-tts Bills by Token

The newest model breaks the per-character pattern. It bills $0.60 per million text tokens for the input and $12 per million audio tokens for the output, which works out to about $0.015 per minute of generated audio. That lands almost level with tts-1, but it adds an instructions field where you describe the delivery in plain language, such as "speak slowly, like a patient teacher." The tts-1 models have no such control.

Software developer at a standing desk with two monitors showing blurred code editors, wiring a text-to-speech API into an app

A minimal call with the official Python SDK looks like this:

from openai import OpenAI

client = OpenAI()

with client.audio.speech.with_streaming_response.create(
    model="tts-1-hd",
    voice="coral",
    input="Your order shipped this morning and should arrive on Friday.",
    response_format="mp3",
) as response:
    response.stream_to_file("order-update.mp3")

Swap the model name to gpt-4o-mini-tts and add an instructions string to steer the delivery. Each request accepts up to 4,096 characters of input, so long scripts must be split into chunks and joined afterwards.

The Voices You Actually Get

Voice count is where the three OpenAI models differ most, and it matters more than most pricing pages admit. A cheap model with one voice you dislike is not cheap.

Low-angle view of a podcast host in a corduroy jacket speaking into a broadcast microphone in a warmly lit studio

Nine Voices on tts-1 and tts-1-hd

Both classic models share the same nine voices: alloy, ash, coral, echo, fable, onyx, nova, sage and shimmer. That is enough range for a narrator, a friendly assistant and a deeper announcer, but not for a cast of distinct characters.

Four Extra Voices on the Newer Model

gpt-4o-mini-tts offers 13 voices. It keeps the original nine and adds ballad, verse, marin and cedar. OpenAI recommends marin or cedar when you want the best quality.

ModelVoicesNames
tts-19alloy, ash, coral, echo, fable, onyx, nova, sage, shimmer
tts-1-hd9The same nine
gpt-4o-mini-tts13The nine plus ballad, verse, marin, cedar

A few practical details apply to all three:

  • Output formats: MP3 by default, plus Opus, AAC, FLAC, WAV and PCM. WAV and PCM return fastest.
  • Streaming: the API supports chunked streaming, so playback can begin before the full file is generated.
  • Languages: roughly 99, following Whisper's language support. OpenAI notes the voices are optimized for English, so test your target language before you commit.
  • Disclosure: the usage policies require a clear notice to end users that the voice is AI-generated and not human.

tts-1 or tts-1-hd?

Both models take the same nine voices and the same input, so switching is a one-word change in your code. That makes the choice easy to test: render one paragraph on each and listen on the device your audience actually uses.

Extreme close-up of studio headphones with leather ear pads resting on a wooden table in morning light

What the Extra $15 Buys

OpenAI positions tts-1 as the lower latency option and tts-1-hd as the higher quality one. Expect smoother, steadier delivery on long passages from the HD model, and expect the gap to be easiest to hear on headphones. On a phone speaker inside a noisy app, many listeners will not notice it.

How to Choose Between Them

  • Pick tts-1 for chat replies, voice assistants, notifications and anything where the first sound should arrive fast.
  • Pick tts-1-hd for audiobooks, course lessons, podcasts and ads, audio that people replay and judge closely.
  • Pick gpt-4o-mini-tts when the voice needs a mood: a calm support agent, an excited product announcer, a whisper.

💡 Rule of thumb: for finished content the math is simple. Rendering an 80,000-word audiobook on HD instead of standard adds about $7.20 to the bill, so default to HD whenever listeners will hear the audio more than once.

Real Cost Examples

Every figure below assumes a narration pace of 150 words per minute and about 6 characters per word including spaces, which works out to 900 characters per minute. Your own scripts will vary, so treat these as budget estimates, not invoices.

Jobtts-1tts-1-hdgpt-4o-mini-ttsElevenLabs FlashElevenLabs v3 or Multilingual v2
1,000-word article (6,000 characters)$0.09$0.18about $0.10$0.30$0.60
80,000-word audiobook (480,000 characters)$7.20$14.40about $8.00$24.00$48.00
1 million characters a month$15$30about $16.70$50$100

Audiobook narrator reading a thick manuscript inside a foam-lined vocal booth, seen through the control room glass

A 1,000-Word Article

Turning a blog post into a listenable version costs under a dime on tts-1 and under a quarter on HD. Even at the ElevenLabs top rate it is 60 cents, so for occasional narration the price difference is almost irrelevant. Pick on sound, not cost.

An 80,000-Word Audiobook

A full-length book is where the gap becomes visible: $14.40 on tts-1-hd versus $48 on the ElevenLabs top models, for roughly nine hours of audio. A human narrator costs far more per finished hour than any column in that table, so every option here is a fraction of a studio booking.

One Million Characters a Month

At app scale the tiers separate. A product that voices a million characters monthly pays $15 to $100 depending on the provider and model. Watch retries: if you regenerate every third clip, add a third to every figure. Cache audio for repeated phrases such as greetings and error messages, because the same text never needs to be paid for twice.

OpenAI TTS vs ElevenLabs

The cost table shows who is cheaper. It does not show who sounds better for your use case, or what happens when your script is longer than a single request allows.

Two colleagues comparing text-to-speech options at a shared desk, one pointing at a laptop while the other takes notes

FactorOpenAI tts-1-hdElevenLabs Flash v2.5ElevenLabs Multilingual v2 and v3
API price per 1,000 characters$0.030$0.05$0.10
LanguagesAbout 993229 (v2), 70+ (v3)
Max characters per request4,09640,00010,000 (v2), 5,000 (v3)
Built-in voices9Large library plus voice cloningLarge library plus voice cloning
Speed focusLower latency on tts-1About 75 ms stated by the vendorHigher quality, slower

The ElevenLabs numbers come from mid-2026 pricing breakdowns. The company updates its plans often, so confirm the current figures on its pricing page before you set a budget.

Price Per 1,000 Characters

On the API, tts-1 sits at $0.015, tts-1-hd at $0.030, ElevenLabs Flash and Turbo at $0.05, and Multilingual v2 and v3 at $0.10. That makes ElevenLabs Flash about 1.7 times the cost of tts-1-hd, and its top models about 3.3 times. Subscriptions bundle credits instead: the free plan includes 10,000 credits a month, with one credit per character on the standard models and half a credit on Flash and Turbo.

Voice Quality and Expression

Price is half the story. ElevenLabs built its reputation on expressive, natural delivery, and Eleven v3 adds emotional range and multi-speaker dialogue across 70+ languages. OpenAI's answer is steerability: gpt-4o-mini-tts accepts plain-language instructions about accent, tone, speed and emotion, something neither tts-1 model can do.

If you need a cast of characters that sound different in every scene, the ElevenLabs voice library and cloning give you more room. If you need one reliable, cheap voice, OpenAI wins on cost. The only honest test is to run the same paragraph through both and listen.

Limits, Languages and Licensing

The request limits shape your code. OpenAI caps input at 4,096 characters and Eleven v3 at 5,000, so a long script needs chunking and a way to keep the voice consistent between pieces. Flash v2.5 accepts 40,000 characters, which means fewer joins for long reads.

Licensing differs too. The ElevenLabs free plan has no commercial license, and commercial rights begin at the first paid plan. OpenAI requires you to tell end users the voice is AI-generated. Read both sets of terms before shipping audio to customers.

How to Use ElevenLabs v3 on PicassoIA

Before paying per character, you can hear the ElevenLabs side of this comparison in a browser. PicassoIA's text-to-speech collection lists 24 models, including ElevenLabs v3, Flash v2.5, Turbo v2.5 and Multilingual v2, plus options from MiniMax, Google, Inworld and Resemble AI. Paste a script, pick a voice and the audio comes back in seconds.

Young man in a knit beanie listening to an AI voice sample through a single earbud beside a window

Steps in the Browser

  1. Open the ElevenLabs v3 page in the text-to-speech collection.
  2. Paste your script into the Prompt field.
  3. Choose a voice from the dropdown. The default is Rachel, and the list has more than 25 personas.
  4. Set the language code. The default is en; use es or fr for Spanish or French.
  5. Leave stability at 0.5, similarity boost at 0.75, style at 0 and speed at 1 for the first run.
  6. Run the model, listen, adjust one setting at a time and download the file you like.

Settings That Change the Result

  • Stability: higher values keep the voice steady across a long script, lower values allow more variation.
  • Style: a 0 to 1 slider that moves delivery from neutral narration toward a more theatrical read.
  • Speed: anywhere from 0.25x to 4x, handy for slowing down tutorial narration.
  • Previous and next text: paste the sentences before and after a chunk so intonation stays natural at the join. This is the fix for the per-request character limits mentioned above.

For a quick A/B test, run the same paragraph through Flash v2.5 for speed, MiniMax Speech 2.8 HD for studio-style quality, Gemini 3.1 Flash TTS for its 30 voices and 70+ languages, and Inworld TTS 1.5 Max for fast multilingual voiceovers. Then compare the results against a tts-1-hd sample from your own account.

Music and Transcription Too

Voice is rarely the only audio job. A product video needs a music bed, and a podcast needs a transcript for captions and search.

Musician at a wooden desk with a pad controller and studio monitor in warm sunset light, composing a track

Generate music. For a background track, try MiniMax Music 2.6, Lyria 3 Pro, ElevenLabs Music or Stable Audio 2.5. Describe the genre, mood and tempo in a sentence and let the model draft the track.

Journalist in a cafe typing on a laptop beside a small voice recorder and a cappuccino, turning an interview into text

Transcribe audio. To turn finished narration or an interview into text, use GPT-4o Transcribe, GPT-4o Mini Transcribe or Gemini 3 Pro. A transcript doubles as captions, show notes and a proofreading pass: if the transcript misreads a word, a listener probably will too.

A simple pipeline ties it together: write the script, render the voice, add a music bed, then transcribe the final mix to check the wording.

Try It Yourself Today

Pricing tables only go so far. The fastest way to settle OpenAI versus ElevenLabs is to hear your own script in several voices and watch what each option costs at your volume. Open PicassoIA, paste a paragraph into a text-to-speech model, and compare it against a tts-1-hd clip. Then add a music track, transcribe the result, and generate a few photorealistic images for the cover of your episode or video. Everything you need to test a full audio and visual workflow sits in one place, so experiment freely and keep whichever combination sounds best.

Share this article