Generate speechGenerate musicTranscribe audio

Fish Audio API: TTS and Voice Cloning Pricing, Line by Line

Fish Audio bills text to speech at $15 per million UTF-8 bytes, transcription at $0.36 per audio hour and voice design at one cent per request. This breakdown turns those rates into real project budgets, compares web plans with the API, and shows where voice cloning fits.

Fish Audio API: TTS and Voice Cloning Pricing, Line by Line
Cristian Da Conceicao
Founder of Picasso IA

Fish Audio sells speech by the byte, and that one detail explains almost every surprise on the invoice. The Fish Audio API charges $15 per million UTF-8 bytes of text for text to speech, $0.36 per audio hour for transcription, and $0.01 per request for voice design. Voice cloning has no separate line on that rate table, which makes budgeting simpler than it first looks.

This breakdown turns those rates into real project budgets: a blog post, an audiobook, a support bot, a month of podcast intros. You will also see how the web plans compare with the API, where rate limits bite, and how the price stacks up against other speech engines. The figures come from Fish Audio's published pricing and developer docs and they can change, so recheck them before you commit money.

💡 Quick answer: one hour of English speech costs about $1.25 through the API. A 1,000-word article, narrated start to finish, costs about 8 cents.

What Fish Audio Charges Per Byte

Condenser microphone with a pop filter in a wood-paneled recording room

The rate card

Every price on the developer side fits in one table.

ServiceModel or featurePrice
Text to speechs2.1-pro$15.00 per million UTF-8 bytes
Text to speechs2-pro$15.00 per million UTF-8 bytes
Text to speechs1$15.00 per million UTF-8 bytes
Text to speechs2.1-pro-free$0.00 per million UTF-8 bytes
Transcriptionboth available models$0.36 per audio hour
Voice designeach successful request$0.01

The docs recommend s2.1-pro for new projects, and the three paid models share one price, so picking between them is a quality decision, not a money decision. The free s2.1-pro variant is listed at zero, which suits prototyping while you test a script.

Fish Audio's own yardstick: one million bytes is roughly 180,000 English words, or about 12 hours of speech. Divide $15 by 12 and you get the number worth memorizing, $1.25 per hour of finished audio.

Why bytes, not characters

A byte is the unit your text takes up once it is encoded as UTF-8. For plain English, one character equals one byte, so $15 per million bytes is also $15 per million characters. Other scripts break that equation:

  • Accented Latin letters and Cyrillic take about 2 bytes each.
  • Chinese, Japanese and Korean characters take about 3 bytes each.
  • Emoji take 4 bytes.
  • Spaces and punctuation are bytes too, so a script padded with stage directions costs more than the same script trimmed.

Overhead view of a desk with a notebook of handwritten cost calculations and printed manuscript pages

💡 Measure before you send: in Python, len(text.encode("utf-8")) returns the exact byte count of a script. Multiply by 0.000015 and you have the price in dollars.

Real Costs for Real Projects

Numbers on a rate card feel abstract until you attach them to work. These examples use Fish Audio's own ratio of 180,000 words per million bytes.

JobAudio lengthBytes sentCost
One 1,000-word articleabout 4 minutesabout 5,600about $0.08
30 daily five-minute briefings2.5 hoursabout 208,000about $3.13
An 80,000-word audiobookabout 5.3 hoursabout 444,000about $6.67
10,000 support replies of 300 charactersabout 36 hours3,000,000$45.00
One million Japanese charactersvariesabout 3,000,000about $45.00

Narration and audiobooks

An entire audiobook for less than the price of a sandwich is the headline here. Even a generous production run with five full re-renders of an 80,000-word title lands near $33. The bill follows the text you send, not the seconds that come back, so slowing the voice with the speed control (the docs allow 0.5 to 2.0) does not raise the price.

Woman in a gray sweater reading a printed manuscript into a studio microphone

Bots and high volume

Short messages add up quietly. A support bot answering 10,000 questions with 300-character replies burns three million bytes, which is $45. Cache any reply that repeats. A greeting spoken 50,000 times should be generated once and stored.

Non-Latin scripts

Budget in bytes, not characters, whenever the script is not English. A million Japanese or Chinese characters is about $45, three times the English figure, and Cyrillic lands near $30. The audio is not longer, but the meter runs on encoded size.

Four ways to trim the bill

Small habits matter once volume grows:

  1. Prototype on the free variant. Draft your pacing and punctuation on s2.1-pro-free, check its limits in the docs, then render finals on a paid model.
  2. Strip the extras. Remove stage directions, repeated whitespace and leftover markup before the text goes out.
  3. Cache repeating lines. Greetings, menu prompts and disclaimers should be generated once and stored.
  4. Render in paragraphs. A typo then costs you one paragraph, not a whole chapter.

Voice Cloning and What It Costs

The rate card has no cloning fee. Based on the pricing page, generating speech with a cloned voice is billed as ordinary text to speech, per byte. What changes is the workflow.

Instant cloning versus saved voices

The docs describe two routes:

  • Instant cloning: pass reference audio directly with the speech request. Nothing is stored, which fits one-off jobs.
  • Persistent voices: create a voice once, receive a reusable voice ID, and call it from any later request. The default fast training mode makes the voice usable almost immediately.

Saved voices are private by default. You can switch one to unlisted, which gives a shareable link, or public, which places it in the community Voice Library.

Sound engineer's hands resting on the faders of an analog mixing console

A request in the TTS endpoint looks like this:

POST /v1/tts
Authorization: Bearer <your token>
model: s2.1-pro

{
  "text": "Your script goes here.",
  "reference_id": "<voice id>",
  "format": "mp3",
  "latency": "balanced"
}

The latency field takes balanced, which Fish Audio describes as roughly 300 ms to first audio, or normal, which favors stability. Output formats are mp3, wav, opus and pcm. Fish Audio's materials also mention inline emotion tags for tone control, so check the API reference for the exact syntax before you rely on them.

Samples, consent and rights

A good clone starts with a good recording. The docs ask for clean, mono, single-speaker audio, at least 10 seconds per clip, and note that a minute or two of clear speech improves fidelity. Accepted files are wav, mp3, m4a and opus, and background noise removal is on by default. Marketing pages quote a 15-second sample across 30 or more languages, so treat 10 to 15 seconds as the floor, not the target.

Hands holding a smartphone to record a voice sample in a quiet living room

Rights matter more than price. Paid subscribers may use verified voices that they own for commercial work, while free plan output is limited to personal, non-commercial projects. Check the license terms for your plan before publishing anything commercial, and if you clone a client's or a contractor's voice, get written permission first.

💡 Rule of thumb: clone your own voice, a voice you hired for the purpose, or a voice with a signed release. Nothing else.

Web Plans Versus the API

Fish Audio also sells monthly plans for people who work in the web studio instead of writing code.

PlanMonthly priceCreditsStated generation timeExtras
Free$08,000about 7 minutes500 characters per generation, 3 public voice slots
Plus$11250,000up to 200 minutes15,000 characters per generation, 10 private voices, 1 professional voice
Pro$752,000,000up to 1,620 minutes30,000 characters per generation, 5 professional voices, 3 team seats
Max$74925,000,000up to 6,250 minutes15 professional voices, 10 team seats
EnterpriseCustom, billed annuallyCustomCustomZero data retention, on-premise deployment, SOC2

Annual billing lowers the monthly rate: Plus is $132 a year, Pro $900, and Max $8,988.

Three colleagues comparing printed price sheets around a wooden table

Where the API wins

Convert each plan's stated minutes into API spend at $1.25 per hour and the gap is large.

PlanStated minutesSame audio through the APIPlan price
Plus200about $4.17$11
Pro1,620about $33.75$75
Max6,250about $130.21$749

Those minute counts are the plan page's own figures, so the comparison is approximate. Still, the pattern is clear: a plan buys the web studio, voice slots, and team seats, while raw audio is cheaper by the byte. Developers, batch jobs and automated pipelines belong on the API. Editors and marketers who never open a terminal may happily pay for the interface.

Here is a worked month for a small podcast team. It generates 40 hours of speech ($50.00), transcribes 40 hours of raw guest tape ($14.40) and tries 20 new voice designs ($0.20). The total is $64.60. That is less than the $75 Pro plan, which states only 27 hours of generation, and the team keeps full control over automation and storage.

Transcription, Voice Design and Limits

Transcription at $0.36 an hour

Fish Audio's speech recognition costs $0.36 per audio hour, billed on processed duration rounded to the nearest second. Ten hours of interviews cost $3.60, and a hundred hours cost $36. Pair it with TTS and you can build a loop that transcribes a recording, edits the text, and speaks the result back in a cloned voice.

Studio headphones resting on a stack of highlighted transcript pages

Voice design at one cent

Voice design creates new voices from a written description instead of a recording, and one request can return several candidates. The charge is $0.01 per successful POST request, and it applies once no matter how many candidate voices come back. A hundred design attempts cost one dollar, so experiment freely.

Concurrency tiers

Speed limits depend on how much you have prepaid:

TierPrepaid spendConcurrent requests
Starterunder $1005
Elevated$100 or more15
High Volume$1,000 or more50
EnterpriseCustomCustom

A tier switches on as soon as you reach its prepaid threshold, and the balance does not need to be spent first. If you plan to narrate around 80 hours a month ($100 of speech), a $100 top-up buys both the audio and three times the parallelism.

Narrow aisle between black server racks in a data center

How the Price Compares

Prices below are as reported in Fish Audio's developer comparison and a third-party pricing review. Treat them as a starting point and confirm on each vendor's page.

EngineReported price per million characters
Fish Audio$15 (per million bytes)
Google Cloud TTS, standard voices$4
Amazon Polly, standard / neural$4 / $16
Azure TTS, neural voices$15
ElevenLabs Flash v2.5about $60
ElevenLabs v2 Multilingualabout $120 to $165

What the table hides

Price is not quality. Third-party pricing reviews report that Fish Audio's S2 Pro scores well in blind listening tests, while ElevenLabs holds the more mature feature set. A cheap engine that needs three re-renders per paragraph is no longer cheap. Non-Latin scripts also tilt the math, because bytes triple the cost for Japanese and Chinese.

The fix is a bake-off: run the same 200-word script through several engines and judge by ear. On Picasso IA you can do that in one place with ElevenLabs v3, MiniMax Speech 2.8 HD, Qwen3 TTS, Chatterbox and Gemini 3.1 Flash TTS.

Run Voice Cloning on PicassoIA

If you want to hear a clone before writing integration code, MiniMax Voice Cloning turns a short recording into a reusable voice profile. PicassoIA lists it as free to try online, with no coding needed.

Pick the model and the sample

  1. Open the MiniMax Voice Cloning page and upload your voice file: MP3, M4A or WAV, from 10 seconds to 5 minutes, under 20 MB.
  2. Switch on noise reduction if the recording has hiss or room tone, and volume normalization if the level wobbles.
  3. Leave accuracy at the default of 0.7 at first. It sets the text validation threshold between 0 and 1.
  4. Choose the speech tier to train on. The default is Speech 2.6 HD, and Speech 2.6 Turbo is the faster option.

Generate speech with the clone

Run the model and you get back a voice profile you can reuse across as many text to speech runs as the project needs. Run your script through the speech tier you trained on, then compare the result against what you hear from the Fish Audio API.

Podcaster wearing headphones looking at a laptop in a cozy home studio

💡 Better sample, better clone: record 30 seconds at your natural pace in a carpeted, curtained room, with no music underneath.

The same workspace handles the rest of an audio pipeline. Add a bed track with Stable Audio 2.5, MiniMax Music 2.6 or ElevenLabs Music, then proofread the narration by running it back through GPT-4o Transcribe or Gemini 3 Pro.

Build Your Voice Pipeline on Picasso IA

Pricing tables answer what a voice costs. Only your own ears answer whether it is worth it. Open Picasso IA, upload a short sample into MiniMax Voice Cloning, read the same paragraph through two or three speech models, and note which one you would trust with your brand.

Then add music and a transcript check, and you have a draft of a full audio workflow before spending a cent on an API contract. A simple test plan works well:

  • Clone a 30-second sample of your own voice.
  • Narrate the same 200 words with three different speech models.
  • Transcribe each result and count the words the engine got wrong.
  • Price the winner using the byte math above.

Start with one short script today, compare the results side by side, and let the clips you actually enjoy decide which engine gets your budget.

Share this article