Generate speechGenerate musicTranscribe audio

Voice Cloning API: Free, ElevenLabs, Cartesia and MiniMax Compared

Three voice cloning APIs side by side. See what ElevenLabs, Cartesia and MiniMax give you for free, how much audio each needs to clone a voice, what a million characters costs, and where a browser tool beats writing integration code.

Voice Cloning API: Free, ElevenLabs, Cartesia and MiniMax Compared
Cristian Da Conceicao
Founder of Picasso IA

Every voice cloning API pitch starts with the word free, and almost none of them mean it the way you hope. A free ElevenLabs account cannot clone a voice at all. Cartesia's free plan is limited to non-commercial use. In the roundups checked for this comparison, MiniMax lists no free plan and charges $1.50 per cloned voice instead. If you are picking an API for a product, a podcast pipeline or an audiobook workflow, the headline price says very little.

This comparison puts the three providers side by side on the points that decide a project: how much audio each one needs, what you pay before the first finished sentence, what happens to the cloned voice afterwards, and where a browser tool beats writing integration code. Figures come from 2026 pricing roundups as of October 2026. Providers change plans often, so confirm the numbers on the live pricing page before you commit a budget.

What a Voice Cloning API Does

A voice cloning API turns a short recording into a reusable voice ID. You upload the sample once, the provider returns an identifier, and every later text to speech request carries that identifier so new text comes out in the same voice. No retraining per script, no re-uploading the recording.

Instant vs Professional Clones

Providers sell two grades of clone, and the grade drives both quality and price:

  • Instant clone: needs 10 seconds to a couple of minutes of audio and is ready within seconds. It captures timbre and general delivery, but it can flatten emotion on long reads.
  • Professional clone: needs 30 minutes to several hours of audio and trains for hours. It tracks accent, pacing and breath habits much more closely. Most providers gate it behind higher plans and a consent check.

The Request Flow

The integration has the same shape everywhere:

  1. Send the sample file to the cloning endpoint.
  2. Store the voice ID the response returns.
  3. Send text plus the voice ID to the speech endpoint.
  4. Receive an audio file or stream, billed per character or per credit.

Two details change a budget more than the headline rate. The first is where the clone lives: MiniMax removes an unused clone after a week, so the project has to plan for that. The second is what a test costs: MiniMax bills previews separately, and on a credit based plan every test render spends credits too. A project that regenerates the same paragraph twenty times while tuning the delivery pays for all twenty.

Overhead view of a developer desk with a laptop, USB microphone, headphones and a notebook of pencil sketches

Three APIs Side by Side

Here is the short version before the detail. All amounts are USD list prices.

ElevenLabsCartesiaMiniMax
Free plan10,000 credits a month, no commercial license, no instant cloning20,000 credits a month, non-commercial onlyNo free plan listed, $1.50 per cloned voice
Cheapest instant cloningStarter plan, about $5 to $6 a monthPro plan, about $5 a monthPay as you go
Audio for an instant clone1 to 2 minutesAbout 10 seconds10 seconds to 5 minutes
Professional cloneCreator plan ($22) and up, 30+ minutes of audio, voice verificationStartup plan ($49 a month) and upNot listed
Billing unitCharacters, through monthly creditsCreditsCharacters
List price per 1M characters$50 (Flash) to $100 (Multilingual v2, v3)Varies by plan$60 (Turbo) to $100 (HD)

💡 Compare cost per finished minute, not per plan. One million characters is roughly 18 to 20 hours of narration, so even a $100 per million rate comes to under 9 cents per minute of audio.

Cartesia's credits and monthly allowances make a direct price comparison impossible without a test run, which is why it gets its own section below.

A team of colleagues listening to a voice sample on a portable speaker around an oak table in a bright office

ElevenLabs: Polished and Pricey

ElevenLabs is the name most people start with, and the voices justify that. The trade is price, plus a cloning ladder that begins above the free plan.

What the Free Plan Gives You

The free plan includes 10,000 credits a month and has no commercial license, so audio from it cannot go into monetized content or client work. Instant voice cloning is not part of it either: the only voice creation tool on the free plan is Voice Design. Treat it as a way to audition stock voices and test the API shape, nothing more.

Instant, Professional and Price

The Starter plan (about $5 to $6 a month, 30,000 credits) adds a commercial license and instant voice cloning from 1 to 2 minutes of audio. For something closer to the real speaker you step up to Professional Voice Cloning on Creator ($22 a month, 121,000 credits) and above. It asks for a minimum of 30 minutes of audio, recommends 2 to 3 hours, trains for several hours and requires voice verification, which blocks cloning someone who has not agreed.

On the API side, pricing roundups list $0.10 per 1,000 characters on Multilingual v2 and V3, and $0.05 on Flash v2.5. If you want to hear the model quality before paying for any plan, all three run on PicassoIA.

Voice actor with silver hair recording in a padded vocal booth in front of a vintage ribbon microphone

Cartesia: Built for Real Time

Cartesia is built around speed, and its plans show it: even the free tier bundles $1 of voice agent usage next to the credits, which tells you who the product is for.

The Free Tier and the $5 Step

The free plan gives 20,000 model credits a month for personal, non-commercial use. Cloning sits one step up. The Pro plan, about $5 a month or $48 a year, includes instant cloning from a 10 second clip, the shortest requirement of the three providers. Pro Voice Cloning arrives on the Startup plan ($49 a month, or $468 billed yearly) with 1,250,000 credits a month, and the Scale tier lifts that to 8,000,000.

Credits Instead of Characters

Credits make quotes hard to compare. ElevenLabs and MiniMax bill by character, while Cartesia bills in credits whose cost per spoken minute depends on the plan. Run your real script through a trial, divide the credits spent by the characters you sent, and only then set the number beside the other two.

Cartesia is not on PicassoIA. If low latency is what you are testing, Inworld Realtime TTS 1.5 Mini (listed at 120ms) and Inworld Realtime TTS 1.5 Max (sub-200ms) are the closest speed focused options there.

Close-up of a sound engineer's hands sliding the faders on an analog mixing console

MiniMax: Lowest Price Per Character

MiniMax is the outlier: no subscription ladder for cloning, a per-voice fee, and the lowest price per character of the three at the Turbo tier.

Ten Seconds and the 168 Hour Rule

MiniMax Speech 2.8 clones from a 10 second reference. The file can be MP3, M4A or WAV, from 10 seconds up to 5 minutes and under 20 MB. The model IDs are speech-2.8-hd and speech-2.8-turbo.

The catch is that a cloned voice is temporary. If you do not run at least one real synthesis with the voice ID within 168 hours (7 days), it is deleted, and the preview audio generated at clone time does not count. Save the voice ID the moment it comes back and send a short real request straight away.

What One Voice Costs

Cloning costs $1.50 per voice, charged on first synthesis use, and previews are billed separately. Speech runs $100 per million characters on HD and $60 on Turbo. Both tiers are on PicassoIA: Speech 2.8 HD and Speech 2.8 Turbo.

Here is what 100,000 characters of narration (about 17,000 words) costs with one cloned voice, at list prices:

Provider and modelSpeech costClone costTotal
ElevenLabs Flash$5Plan fee$5 plus plan
ElevenLabs Multilingual v2 or v3$10Plan fee$10 plus plan
MiniMax Speech 2.8 Turbo$6$1.50$7.50
MiniMax Speech 2.8 HD$10$1.50$11.50
Cartesia SonicCreditsPlan feeRun a trial

ElevenLabs and Cartesia need a paid plan before you can clone at all, so their real monthly bill is higher than the speech line alone.

Extreme close-up of a condenser microphone and nylon pop filter lit by a soft beam of morning light

Free Options That Actually Work

Free tiers are test beds. ElevenLabs' free plan cannot clone and Cartesia's is non-commercial, so use them to audition voices and check how the API behaves, then budget for a paid plan if the output ships.

Cloning in the Browser

If your goal is a finished voice file rather than an integration, PicassoIA runs cloning with no code:

  • Voice Cloning (MiniMax): a reusable voice profile from a 10 second to 5 minute sample, listed as free to try online.
  • Chatterbox from Resemble AI: voice cloning with emotion control.
  • Qwen3 TTS: clone a voice or design a new one from scratch.

💡 As of this writing, PicassoIA's developer API lists image and video models, so voice cloning there runs from the browser. Check the API page for the current model list before you plan an integration.

Audiobook narrator in profile reading from a hardback book inside a glass fronted recording booth

How to Use Voice Cloning on PicassoIA

  1. Open the Voice Cloning page.
  2. Upload your sample as MP3, M4A or WAV, from 10 seconds to 5 minutes and under 20 MB.
  3. Switch on noise reduction if the room was not silent, and volume normalization if the levels wander.
  4. Pick the speech tier you plan to use: Speech 2.6 HD (the default), Speech 2.6 Turbo, Speech 02 HD or Speech 02 Turbo.
  5. Leave accuracy at its default of 0.7 and only adjust it if results drift from the way you pronounce words.
  6. Run it and keep the voice ID for your synthesis runs.

💡 Use an HD tier for finished narration and a Turbo tier for quick drafts. Drafting on Turbo and rendering the final pass on HD keeps the iteration cheap.

Record a Clean Sample

Every provider's clone is only as good as the recording. The sample matters more than the plan you pay for.

Young woman recording a voice sample on a smartphone inside a closet padded with wool coats and blankets

  • Pick a small, soft room. A closet full of coats beats a bare kitchen. Hard walls add echo that the clone will copy.
  • Stay a hand span from the microphone. Closer causes pops, farther picks up the room.
  • Record 30 to 60 seconds even when the minimum is 10. Read varied sentences with a question, a number and a pause.
  • Keep it clean. One speaker, no music, no fan noise.

Once the clone exists, check it with a round trip. Run the generated audio through GPT-4o Transcribe and compare the text to your script. Any word that comes back different is a spot where the voice slurs. For a podcast or video intro, you can also lay the cloned narration over a track from Music 2.6 or Stable Audio 2.5.

Which One Fits Your Project

Studio audio interface with metal gain knobs next to a pair of studio monitor speakers and folded reading glasses

Your projectBest fitWhy
Live voice agentCartesiaSpeed focused plans, 10 second clone
Long audiobookMiniMax Turbo or HDLowest cost per character
Polished multilingual videoElevenLabsModel range and professional clones
One off voiceoverBrowser toolNo code, no plan fee

Voice Agents and Live Calls

A phone agent lives or dies on the delay before the first word. Cartesia's plans bundle voice agent usage, and its 10 second clone is the fastest to set up. Test the credit cost per call before you scale it.

Audiobooks and Long Narration

Narration is a volume game, so cost per character dominates. MiniMax Turbo at $60 per million characters, or HD at $100, keeps a long book affordable. Split the text into chapters, reuse one voice ID, and remember the 168 hour rule: keep generating inside the window.

Apps With Many Voices

If every user records their own voice, per voice pricing adds up fast. At MiniMax's $1.50, 10,000 users each cloning one voice is $15,000 in clone fees before any speech is generated. Also store each original sample, because a voice that expires after seven idle days has to be recreated.

Questions to Settle First

Before you pick a provider, answer these five:

  • Do you need commercial rights? The free plans on ElevenLabs and Cartesia do not include them.
  • How much audio do you have of the speaker? With only 10 seconds, Cartesia and MiniMax work, while ElevenLabs instant cloning wants 1 to 2 minutes.
  • Does the voice have to stay permanent? MiniMax needs a real synthesis inside seven days.
  • Do you bill by characters or credits? Pick the unit that matches how your own product charges users.
  • Which languages do you ship? Check each provider's language list against yours before you clone, since a voice may not speak every language equally well.

Consent comes first, whichever API you pick. Clone only voices you own or have written permission to use. ElevenLabs requires verification for professional clones, and a free plan's audio carries no commercial rights.

Elderly man with white hair speaking into a small microphone at a sunlit kitchen table beside an old family photograph

A good use of a clone is a family one: a grandparent who agrees to record twenty minutes of stories, preserved in a voice the grandchildren can still hear.

Try It on PicassoIA

The fastest way to compare these providers is to hear your own voice come back. Record 30 seconds in a quiet closet, upload it to Voice Cloning on PicassoIA, and read a paragraph of your own script in the cloned voice. Then switch the tier from Turbo to HD and listen to the difference, or try Chatterbox for emotion control. Ten minutes of testing will tell you more about what you need than any pricing table, and it costs nothing to try.

Share this article