Generate speechGenerate musicTranscribe audio

ElevenLabs v4: What's New for Voices and the API

ElevenLabs launched Eleven v4 and v4 Turbo on September 28, 2026. This article breaks down the 90+ languages, stackable audio tags, voice cloning from 10 seconds, new model IDs, latency, and pricing, then shows how to draft scripts with ElevenLabs v3 on PicassoIA today.

ElevenLabs v4: What's New for Voices and the API
Cristian Da Conceicao
Founder of Picasso IA

On September 28, 2026, ElevenLabs shipped two speech models at once: Eleven v4 and Eleven v4 Turbo. If you build voice features, narrate videos, or run a podcast, the short version is this: more languages, audio tags you can stack, a longer request limit, and a fast model aimed at live voice agents. The longer version is below, with the numbers, the API changes, and a practical way to rehearse your scripts today.

💡 Heads up: ElevenLabs is running a two-week launch discount on the API that ends on October 12, 2026. If you plan to test v4 at volume, that date matters more than any feature on this page.

What Eleven v4 Actually Is

ElevenLabs announced both models on the same day, and both are available in its agent platform, its creative studio, and the public API. The pitch is simple: voices that sound like a person acting, not a person reading.

Eleven v4 is the expressive one. It is the model to reach for when performance matters: audiobooks, trailers, character dialogue, ads. Eleven v4 Turbo gives up a little range in exchange for speed. It is built for voice agents, the kind of product where a half-second pause makes a caller wonder whether the line dropped.

Both follow Eleven v3, which already supported inline audio tags. v4 keeps that idea and goes after the weak spots: voices that drift on long reads, a language list that stopped at about 70, and no support for Professional Voice Clones.

A desk with a laptop, printed script, headphones and coffee seen from above

The Numbers at a Glance

SpecEleven v3Eleven v4Eleven v4 Turbo
API model IDeleven_v3eleven_v4eleven_v4_turbo
Languages70+90+90+
Characters per request5,00010,000Not published
SpeedStandardStandardAbout 150 ms to first speech
Built forExpressive narrationExpressive narration, long readsLive voice agents

Those figures come from the ElevenLabs launch post, its model documentation, and press reports from the first week. Where a number is missing, the company has not published it, and I have marked it as not published rather than guess.

Voices That Hold Together

Audio Tags You Can Stack

The headline feature is direction. You write tags inline, in square brackets, and the model acts on them. ElevenLabs gives examples such as [laughs], [said angrily in French accent], [light rain] and [phone buzzing]. Notice that the list mixes emotion, accent, and sound effects in one syntax.

What is new is stacking. According to TechCrunch, you can now put several tags in a row and the model follows the sequence instead of dropping all but one. A line might look like this:

[whispers] The door was already open. [footsteps] [light rain] Somebody had been here.

That is an illustration of the style, not a copy from the docs, so test your own wording. Short, concrete tags tend to be safer than poetic ones.

Consistency Across Long Scripts

If you have ever generated a chapter in pieces, you know the problem. Line 40 sounds like a different narrator than line 3. ElevenLabs says v4 reads the tone, pacing, and context of the whole script and keeps the voice steady, even when you regenerate a single line later.

There is a number attached. The Decoder reports a pronunciation score of 91.7 percent for v4 against 85.6 percent for v3. That is a vendor figure, so treat it as a hint about direction, not a promise about your script.

💡 Tip: Long productions need a plan for retakes. Generate by paragraph, keep each paragraph under the character limit, and re-render only the lines you dislike. v4's consistency is meant to make that workflow safe.

An older male narrator reading a printed book into a microphone in a small vocal booth

90+ Languages and Accents

The language list grows from about 70 to more than 90. TechCrunch notes visible quality gains in Japanese, Brazilian Portuguese, Mandarin, and Cantonese, and press reports mention newly supported languages such as Mongolian and Odia.

One detail is easy to miss. Cloned voices speaking another language keep a native accent for that language and do not slide back toward the original speaker's accent as the text goes on. For anyone localizing a brand voice, that is the difference between a believable Spanish read and an American one with Spanish words.

On outside rankings, ElevenLabs says v4 sits at number one on an independent speech leaderboard, with an Elo score of 1319 reported in launch articles. In blind tests, the company cites roughly 75 percent of listeners preferring v4 over competing models, while The Decoder puts the range at 65 to 81 percent depending on the competitor.

💡 Reality check: Leaderboards use short test sentences. Run your own script, in your own language, before you commit a project to any model.

Voice Cloning Gets Better

Instant Clones from 10 Seconds

Instant Voice Clones need only 10 seconds of audio. ElevenLabs says the new architecture makes cloning faster and holds the voice's identity better across a longer passage. A bigger voice library helps too: launch reports mention more than 17,500 voices to choose from.

The quality of that 10-second sample decides most of the result. Record in a quiet room with no music or fan noise, keep the same distance from the microphone throughout, and speak at your natural pace instead of performing. A clean, boring sample beats an energetic, noisy one every time.

Low-angle close-up of a condenser microphone with a pop filter in a wood-panelled studio

Professional Clones Return

v3 did not support Professional Voice Clones. v4 does, and that is aimed at the highest-fidelity jobs: a narrator whose cloned voice will read hundreds of hours of audio.

Every clone requires verified consent from the voice's owner. That is not a footnote. If you clone a voice for a client, get the permission in writing before you upload a single second of audio.

What Changes for API Developers

Model IDs and Limits

Switching models is a one-field change. Per the ElevenLabs docs, the request body keeps its shape and only model_id moves:

{
  "text": "[laughs] That is not what the invoice said.",
  "model_id": "eleven_v4"
}

Use eleven_v4_turbo for the low-latency version. The per-request limit is 10,000 characters, which ElevenLabs equates to about ten minutes of audio, double the 5,000 characters allowed for eleven_v3. Output comes as MP3, WAV/PCM, or µ-law, and bidirectional streaming over WebSocket is supported, which matters if your text arrives token by token from a language model.

The older models are still around. Per the docs, eleven_flash_v2_5 handles 40,000 characters per request in 32 languages at roughly 75 ms, and eleven_multilingual_v2 supports 29 languages. You can still reach both on PicassoIA through Flash v2.5 and v2 Multilingual.

A software developer typing at a dual-monitor desk with blurred code on the screens

Latency and Streaming

Turbo is the headline here. ElevenLabs reports a median time to first speech of about 150 ms and a median inference latency of about 100 ms. The Decoder compares that with 262 ms for Cartesia Sonic 3.6.

For voice agents, the more interesting claim is about timing. The model can begin producing audio as soon as the language model behind it starts producing its answer, so the caller hears the reply while the rest of it is still being written.

Two caveats. ElevenLabs did not publish a latency figure for standard v4, and no independent latency tests have appeared yet. Measure from your own region, with your own network overhead, because every vendor figure excludes it.

Pricing and the October 12 Deadline

The list price per million characters is $80 for v4 and $40 for v4 Turbo. Until October 12, 2026, the launch offer drops those to $22 and $11. Here is what 100,000 characters, roughly 100 minutes of speech, costs in each case:

ModelList priceLaunch price100,000 characters at list100,000 characters at launch
Eleven v4$80 per million$22 per million$8.00$2.20
Eleven v4 Turbo$40 per million$11 per million$4.00$1.10

If you are evaluating, the discount lets you run a real bake-off cheaply. If you are budgeting a product, plan around the list price, because that is what you will pay once the offer ends.

Which Model for Which Job

A smiling customer support agent wearing a headset at a bright open office

The right pick depends on what the audio is for, not on which model is newest.

JobBest fitWhy
Audiobook or long narrationEleven v4Voice consistency and a 10,000 character limit
Voice agent or phone botEleven v4 TurboAbout 150 ms to first speech
Bulk, cheap, multilingual batchesFlash v2.540,000 characters per request and a lower price
Expressive drafts you can run todayEleven v3Available on PicassoIA now
Steady, established multilingual readsv2 MultilingualA mature model with 29 languages

A practical split for small teams: draft on a faster, cheaper model while the script is still changing, then render the final audio on the expressive one. Voice work gets expensive when you regenerate an entire chapter after every edit, so lock the words first and spend the premium characters last. If your product mixes the two needs, such as a narrated course with a live Q&A assistant, you can use v4 for the lessons and Turbo for the assistant, since both share the same request format.

3 Common Mistakes

  1. Judging by one sentence. A demo line proves nothing about a 40-paragraph chapter. Test your longest, most awkward script.
  2. Using the expressive model for a live agent. If a caller is waiting, Turbo is the right tool even if v4 sounds a little richer.
  3. Budgeting at the launch price. The discount ends on October 12, 2026. Price your project at $80 and $40 per million characters.

How to Use ElevenLabs v3 on PicassoIA

At the time of writing, v4 is not in the PicassoIA catalog. ElevenLabs v3 is, and it is the closest relative: same family, same habit of reading inline direction. That makes it a good place to draft scripts and test voices while v4 reaches more platforms.

Step by Step

  1. Open the ElevenLabs v3 page in the text-to-speech collection.
  2. Paste your script into the Prompt field.
  3. Pick a voice from the list of 26 options. The default is Rachel; others include Aria, Roger, Sarah, and James.
  4. Set the language code, for example en, es, or fr.
  5. Leave the sliders at their defaults for the first run, then press generate and listen.
  6. Adjust one setting at a time and generate again. Changing three things at once tells you nothing.

Two hands adjusting faders on an analog mixing console

Settings That Matter

SettingRangeDefaultWhat it does
Speed0.25 to 4.01Pacing of the whole read
Style0.0 to 1.00Pushes delivery from neutral toward theatrical
Stability0.0 to 1.00.5Higher keeps the voice steadier across a long script
Similarity boost0.0 to 1.00.75How closely the output follows the chosen voice
Previous text / Next textTextEmptyGives the model context so intonation lines up at sentence edges

Two habits will save you time. First, paste the paragraph before and after your current one into the context fields, so each generated chunk joins its neighbors cleanly. Second, try bracketed tags such as [whispers] in the prompt and listen to how the model handles them. That is the same direction style v4 is built around, so the practice carries over.

For other voices to compare, the same collection includes MiniMax Speech 2.8 HD, Gemini 3.1 Flash TTS, and Inworld Realtime TTS 2.

Beyond Speech: Music and Transcripts

A voice model rarely works alone. Most projects that need narration also need a music bed, a transcript, or a translated version. All three are on PicassoIA.

Compose with ElevenLabs Music

ElevenLabs Music turns a text prompt into a song. For narration work, ask for something instrumental, steady, and low in the mix. Write the mood, tempo, and instruments in plain words, such as "slow acoustic guitar, warm room, no vocals, 70 BPM." If you want to compare, Lyria 3 Pro and Music 2.6 take the same kind of prompt.

A musician beside an acoustic guitar adjusting a synth in a small home studio

Transcribe with GPT-4o or Gemini

Once the narration exists, you need captions and a text record. GPT-4o Transcribe turns audio into text, and GPT-4o Mini Transcribe is the lighter option for quick drafts. Gemini 3 Pro is a third path, handy when you want to compare how each model handles names and numbers.

A useful trick: transcribe your generated narration and compare it with the original script. Any mismatch points to a word the voice stumbled over, and you can fix it with a spelling tweak or a pronunciation hint.

Dub Videos into 90+ Languages

ElevenLabs Dubbing takes a finished video and produces versions in other languages. It pairs naturally with v4's wider language list, since localization is where voice consistency pays off most.

A young woman speaking into a microphone while watching a film scene in a dubbing studio

Your Turn to Experiment

The best way to judge a voice model is to give it your script and listen. You do not need a studio, a budget line, or a week of setup to start.

Here is a plan for an afternoon:

  • Write a 150-word script with one emotional turn in the middle.
  • Render it on ElevenLabs v3 with three different voices and one change of the style slider.
  • Add a music bed from ElevenLabs Music.
  • Transcribe the result with GPT-4o Transcribe and check the words against your script.

When v4 reaches your platform of choice, you will already know which tags, voices, and settings suit your material, and you can budget from the list prices with confidence.

Open Picasso IA and try your first script today. Pick a voice, add one stacked tag, and hear the difference direction makes.

Two friends laughing while recording a podcast across a round wooden table

Share this article