Fish Audio API: TTS and Voice Cloning Pricing, Line by Line
Fish Audio bills text to speech at $15 per million UTF-8 bytes, transcription at $0.36 per audio hour and voice design at one cent per request. This breakdown turns those rates into real project budgets, compares web plans with the API, and shows where voice cloning fits.
Fish Audio sells speech by the byte, and that one detail explains almost every surprise on the invoice. The Fish Audio API charges $15 per million UTF-8 bytes of text for text to speech, $0.36 per audio hour for transcription, and $0.01 per request for voice design. Voice cloning has no separate line on that rate table, which makes budgeting simpler than it first looks.
This breakdown turns those rates into real project budgets: a blog post, an audiobook, a support bot, a month of podcast intros. You will also see how the web plans compare with the API, where rate limits bite, and how the price stacks up against other speech engines. The figures come from Fish Audio's published pricing and developer docs and they can change, so recheck them before you commit money.
💡 Quick answer: one hour of English speech costs about $1.25 through the API. A 1,000-word article, narrated start to finish, costs about 8 cents.
What Fish Audio Charges Per Byte
The rate card
Every price on the developer side fits in one table.
Service
Model or feature
Price
Text to speech
s2.1-pro
$15.00 per million UTF-8 bytes
Text to speech
s2-pro
$15.00 per million UTF-8 bytes
Text to speech
s1
$15.00 per million UTF-8 bytes
Text to speech
s2.1-pro-free
$0.00 per million UTF-8 bytes
Transcription
both available models
$0.36 per audio hour
Voice design
each successful request
$0.01
The docs recommend s2.1-pro for new projects, and the three paid models share one price, so picking between them is a quality decision, not a money decision. The free s2.1-pro variant is listed at zero, which suits prototyping while you test a script.
Fish Audio's own yardstick: one million bytes is roughly 180,000 English words, or about 12 hours of speech. Divide $15 by 12 and you get the number worth memorizing, $1.25 per hour of finished audio.
Why bytes, not characters
A byte is the unit your text takes up once it is encoded as UTF-8. For plain English, one character equals one byte, so $15 per million bytes is also $15 per million characters. Other scripts break that equation:
Accented Latin letters and Cyrillic take about 2 bytes each.
Chinese, Japanese and Korean characters take about 3 bytes each.
Emoji take 4 bytes.
Spaces and punctuation are bytes too, so a script padded with stage directions costs more than the same script trimmed.
💡 Measure before you send: in Python, len(text.encode("utf-8")) returns the exact byte count of a script. Multiply by 0.000015 and you have the price in dollars.
Real Costs for Real Projects
Numbers on a rate card feel abstract until you attach them to work. These examples use Fish Audio's own ratio of 180,000 words per million bytes.
Job
Audio length
Bytes sent
Cost
One 1,000-word article
about 4 minutes
about 5,600
about $0.08
30 daily five-minute briefings
2.5 hours
about 208,000
about $3.13
An 80,000-word audiobook
about 5.3 hours
about 444,000
about $6.67
10,000 support replies of 300 characters
about 36 hours
3,000,000
$45.00
One million Japanese characters
varies
about 3,000,000
about $45.00
Narration and audiobooks
An entire audiobook for less than the price of a sandwich is the headline here. Even a generous production run with five full re-renders of an 80,000-word title lands near $33. The bill follows the text you send, not the seconds that come back, so slowing the voice with the speed control (the docs allow 0.5 to 2.0) does not raise the price.
Bots and high volume
Short messages add up quietly. A support bot answering 10,000 questions with 300-character replies burns three million bytes, which is $45. Cache any reply that repeats. A greeting spoken 50,000 times should be generated once and stored.
Non-Latin scripts
Budget in bytes, not characters, whenever the script is not English. A million Japanese or Chinese characters is about $45, three times the English figure, and Cyrillic lands near $30. The audio is not longer, but the meter runs on encoded size.
Four ways to trim the bill
Small habits matter once volume grows:
Prototype on the free variant. Draft your pacing and punctuation on s2.1-pro-free, check its limits in the docs, then render finals on a paid model.
Strip the extras. Remove stage directions, repeated whitespace and leftover markup before the text goes out.
Cache repeating lines. Greetings, menu prompts and disclaimers should be generated once and stored.
Render in paragraphs. A typo then costs you one paragraph, not a whole chapter.
Voice Cloning and What It Costs
The rate card has no cloning fee. Based on the pricing page, generating speech with a cloned voice is billed as ordinary text to speech, per byte. What changes is the workflow.
Instant cloning versus saved voices
The docs describe two routes:
Instant cloning: pass reference audio directly with the speech request. Nothing is stored, which fits one-off jobs.
Persistent voices: create a voice once, receive a reusable voice ID, and call it from any later request. The default fast training mode makes the voice usable almost immediately.
Saved voices are private by default. You can switch one to unlisted, which gives a shareable link, or public, which places it in the community Voice Library.
The latency field takes balanced, which Fish Audio describes as roughly 300 ms to first audio, or normal, which favors stability. Output formats are mp3, wav, opus and pcm. Fish Audio's materials also mention inline emotion tags for tone control, so check the API reference for the exact syntax before you rely on them.
Samples, consent and rights
A good clone starts with a good recording. The docs ask for clean, mono, single-speaker audio, at least 10 seconds per clip, and note that a minute or two of clear speech improves fidelity. Accepted files are wav, mp3, m4a and opus, and background noise removal is on by default. Marketing pages quote a 15-second sample across 30 or more languages, so treat 10 to 15 seconds as the floor, not the target.
Rights matter more than price. Paid subscribers may use verified voices that they own for commercial work, while free plan output is limited to personal, non-commercial projects. Check the license terms for your plan before publishing anything commercial, and if you clone a client's or a contractor's voice, get written permission first.
💡 Rule of thumb: clone your own voice, a voice you hired for the purpose, or a voice with a signed release. Nothing else.
Web Plans Versus the API
Fish Audio also sells monthly plans for people who work in the web studio instead of writing code.
Plan
Monthly price
Credits
Stated generation time
Extras
Free
$0
8,000
about 7 minutes
500 characters per generation, 3 public voice slots
Plus
$11
250,000
up to 200 minutes
15,000 characters per generation, 10 private voices, 1 professional voice
Pro
$75
2,000,000
up to 1,620 minutes
30,000 characters per generation, 5 professional voices, 3 team seats
Max
$749
25,000,000
up to 6,250 minutes
15 professional voices, 10 team seats
Enterprise
Custom, billed annually
Custom
Custom
Zero data retention, on-premise deployment, SOC2
Annual billing lowers the monthly rate: Plus is $132 a year, Pro $900, and Max $8,988.
Where the API wins
Convert each plan's stated minutes into API spend at $1.25 per hour and the gap is large.
Plan
Stated minutes
Same audio through the API
Plan price
Plus
200
about $4.17
$11
Pro
1,620
about $33.75
$75
Max
6,250
about $130.21
$749
Those minute counts are the plan page's own figures, so the comparison is approximate. Still, the pattern is clear: a plan buys the web studio, voice slots, and team seats, while raw audio is cheaper by the byte. Developers, batch jobs and automated pipelines belong on the API. Editors and marketers who never open a terminal may happily pay for the interface.
Here is a worked month for a small podcast team. It generates 40 hours of speech ($50.00), transcribes 40 hours of raw guest tape ($14.40) and tries 20 new voice designs ($0.20). The total is $64.60. That is less than the $75 Pro plan, which states only 27 hours of generation, and the team keeps full control over automation and storage.
Transcription, Voice Design and Limits
Transcription at $0.36 an hour
Fish Audio's speech recognition costs $0.36 per audio hour, billed on processed duration rounded to the nearest second. Ten hours of interviews cost $3.60, and a hundred hours cost $36. Pair it with TTS and you can build a loop that transcribes a recording, edits the text, and speaks the result back in a cloned voice.
Voice design at one cent
Voice design creates new voices from a written description instead of a recording, and one request can return several candidates. The charge is $0.01 per successful POST request, and it applies once no matter how many candidate voices come back. A hundred design attempts cost one dollar, so experiment freely.
Concurrency tiers
Speed limits depend on how much you have prepaid:
Tier
Prepaid spend
Concurrent requests
Starter
under $100
5
Elevated
$100 or more
15
High Volume
$1,000 or more
50
Enterprise
Custom
Custom
A tier switches on as soon as you reach its prepaid threshold, and the balance does not need to be spent first. If you plan to narrate around 80 hours a month ($100 of speech), a $100 top-up buys both the audio and three times the parallelism.
How the Price Compares
Prices below are as reported in Fish Audio's developer comparison and a third-party pricing review. Treat them as a starting point and confirm on each vendor's page.
Price is not quality. Third-party pricing reviews report that Fish Audio's S2 Pro scores well in blind listening tests, while ElevenLabs holds the more mature feature set. A cheap engine that needs three re-renders per paragraph is no longer cheap. Non-Latin scripts also tilt the math, because bytes triple the cost for Japanese and Chinese.
If you want to hear a clone before writing integration code, MiniMax Voice Cloning turns a short recording into a reusable voice profile. PicassoIA lists it as free to try online, with no coding needed.
Pick the model and the sample
Open the MiniMax Voice Cloning page and upload your voice file: MP3, M4A or WAV, from 10 seconds to 5 minutes, under 20 MB.
Switch on noise reduction if the recording has hiss or room tone, and volume normalization if the level wobbles.
Leave accuracy at the default of 0.7 at first. It sets the text validation threshold between 0 and 1.
Run the model and you get back a voice profile you can reuse across as many text to speech runs as the project needs. Run your script through the speech tier you trained on, then compare the result against what you hear from the Fish Audio API.
💡 Better sample, better clone: record 30 seconds at your natural pace in a carpeted, curtained room, with no music underneath.
Pricing tables answer what a voice costs. Only your own ears answer whether it is worth it. Open Picasso IA, upload a short sample into MiniMax Voice Cloning, read the same paragraph through two or three speech models, and note which one you would trust with your brand.
Then add music and a transcript check, and you have a draft of a full audio workflow before spending a cent on an API contract. A simple test plan works well:
Clone a 30-second sample of your own voice.
Narrate the same 200 words with three different speech models.
Transcribe each result and count the words the engine got wrong.
Price the winner using the byte math above.
Start with one short script today, compare the results side by side, and let the clips you actually enjoy decide which engine gets your budget.