Generate speechLarge Language ModelsGenerate music
Text to Speech API Free: Pricing, Python and Unlimited Options
Every free text to speech API has different rules: monthly allowances that reset, twelve month trials that expire, and open source models with no meter. See real prices, working Python code, and a browser workflow for testing voices first.
Type "free text to speech API" into a search bar and you get three very different answers. A monthly allowance from a cloud giant. A Python package that quietly borrows someone else's servers. An open source model you run on your own machine. All three are free, but only one is truly unlimited, and it is not the one most people expect.
This article sorts the options by what they really cost, shows working Python for each, and ends with a way to audition voices in your browser before you write any integration code. The figures come from provider pricing pages and published summaries as of October 2026. Prices move, so confirm the numbers before you build a budget on them.
What Free Really Means
Providers use the word "free" for three different arrangements, and mixing them up is how a prototype turns into a surprise invoice.
Recurring Monthly Allowances
Google Cloud and Microsoft Azure reset a free allowance every month with no expiry date. Google gives 4 million characters of Standard voices, and Azure's F0 tier gives 0.5 million characters of neural voices. Stay under the line and the bill is zero indefinitely. Go over on Google and you pay the normal per character rate. Azure's F0 tier is capped, so you hit a ceiling instead of a bill.
Twelve Month Trials
Amazon Polly's allowance applies to new accounts for their first 12 months. AWS lists 1 million neural characters a month, 500 thousand long form characters and 100 thousand generative characters, and published summaries put standard voices at 5 million. After month twelve the same traffic costs $4 per million characters on standard voices and $16 on neural. A prototype that is free in January can carry a real line item by the next January.
ElevenLabs sits in between. Its free plan gives roughly 10,000 characters a month, which is only ten to fifteen minutes of audio, and it asks for attribution and rules out commercial use. That is plenty for testing a voice and far too small for a product. The same voice family is also available on PicassoIA as ElevenLabs v3.
Open Source and Unlimited
The only truly unlimited option is software you run yourself. Piper runs on a plain CPU with more than 100 voices across 30+ languages. Kokoro is an 82 million parameter model that produces natural speech and reportedly runs close to 100 times faster than real time on a GPU. Nobody meters you because nobody is hosting you. The price is your own hardware, your setup time, and the job of checking each voice's license before you sell anything made with it.
Free Tiers Compared Side by Side
The Pricing Table
Provider and tier
Free allowance
Does it expire?
Paid rate after
Google Cloud Standard
4M characters a month
No
$4 per 1M
Google Cloud Neural2
About 1M characters a month
No
$16 per 1M
Azure Neural (F0)
0.5M characters a month
No
Pay as you go
Amazon Polly Standard
5M characters a month
After 12 months
$4 per 1M
Amazon Polly Neural
1M characters a month
After 12 months
$16 per 1M
ElevenLabs Free
About 10K characters a month
No
Paid plans
Piper or Kokoro
No limit
No
Your hardware
💡 Allowances are counted per account, per month and per voice tier. Standard and premium voices draw from separate pools, so test the exact voice you plan to ship.
What a Million Characters Buys
AWS's own pricing examples work out to about 23 hours and 8 minutes of speech per million characters. Use that as a rule of thumb. Google's 4 million Standard characters is roughly 90 hours of audio a month, and Azure's 0.5 million is about 11 hours.
A 1,500 word blog post runs near 9,000 characters. At that size, Azure's free tier narrates about 55 posts a month and Google's narrates more than 400. A read aloud button on a mid size blog fits inside a free tier, as long as you store each audio file instead of regenerating it on every page view.
Count carefully. Most providers meter every character you send, including spaces, punctuation and markup tags, so strip HTML before synthesis and send plain text. Measure a week of real traffic before you trust any estimate.
Python Options That Cost Nothing
Three packages handle most quick scripts. They install with pip, need no billing account and work in a few lines, but they are not equal.
gTTS in Four Lines
from gtts import gTTS
tts = gTTS("Your order has shipped.", lang="en")
tts.save("order.mp3")
gTTS is an unofficial wrapper around the speech endpoint behind Google Translate. There is no credential and no published quota, which is exactly the problem: Google can throttle or change the endpoint without notice. It suits a classroom script. It is risky for a customer facing feature.
pyttsx3 Runs Offline
import pyttsx3
engine = pyttsx3.init()
engine.setProperty("rate", 170)
engine.save_to_file("Your order has shipped.", "order.wav")
engine.runAndWait()
pyttsx3 drives the voices already installed in your operating system: SAPI5 on Windows, NSSpeechSynthesizer on macOS and eSpeak on Linux. There is no network call and no limit. There is also a voice that sounds like a 2005 GPS unit, so judge it by your own ears before committing.
edge-tts for Better Voices
import asyncio
import edge_tts
async def main():
speech = edge_tts.Communicate("Your order has shipped.", "en-US-AriaNeural")
await speech.save("order.mp3")
asyncio.run(main())
edge-tts talks to the same read aloud service used by the Edge browser. The voices are neural quality and the package is free, but it is an unofficial client with no service agreement behind it. Treat it like gTTS: good for personal projects, a gamble when paying users depend on it.
Local Neural Models
Piper and Kokoro are the two to try first. Piper takes a downloaded voice file and writes audio straight from the command line:
echo "Your order has shipped." | piper --model en_US-lessac-medium.onnx --output_file order.wav
It is light enough for a Raspberry Pi 4, which makes it the usual pick for offline kiosks, home automation and anything that must keep working without internet. Kokoro needs more memory but sounds noticeably more natural, and it flies on a GPU.
Calling a Cloud API From Python
When you need an SLA, SSML markup or dozens of languages, a cloud provider is the honest route. The pattern is the same everywhere: authenticate, send text, receive audio bytes.
A Google Cloud Example
from google.cloud import texttospeech
client = texttospeech.TextToSpeechClient()
response = client.synthesize_speech(
input=texttospeech.SynthesisInput(text="Your order has shipped."),
voice=texttospeech.VoiceSelectionParams(
language_code="en-US", name="en-US-Standard-C"
),
audio_config=texttospeech.AudioConfig(
audio_encoding=texttospeech.AudioEncoding.MP3
),
)
with open("order.mp3", "wb") as f:
f.write(response.audio_content)
Authentication comes from the GOOGLE_APPLICATION_CREDENTIALS environment variable, which points at a service account file, so no secret ever sits inside the script. Swap en-US-Standard-C for a Neural2 voice and the call stays identical. Only the pool you draw from, and the price, changes.
Cache Before You Pay
import hashlib
import pathlib
def cached_speech(text, synthesize):
name = hashlib.sha256(text.encode()).hexdigest() + ".mp3"
path = pathlib.Path("audio_cache") / name
if not path.exists():
path.parent.mkdir(exist_ok=True)
path.write_bytes(synthesize(text))
return path
Hashing the text means a repeated sentence costs nothing the second time. For a read aloud feature this habit alone usually keeps traffic inside the free tier. Two more habits help: split long text at sentence boundaries, since most providers cap a single request at a few thousand characters, and wrap each call in a retry with exponential backoff for rate limit responses.
Unlimited Options and Their Catch
Search for "unlimited" and the same three promises come up. Each one has a cost.
Self hosting costs time. You install models, manage dependencies, watch GPU memory and update everything when a release breaks something. For a few thousand requests a month a free cloud tier is cheaper than your hours. At millions of requests, running Piper or Kokoro on your own server starts to win.
Unofficial endpoints break. gTTS and edge-tts depend on services whose owners never promised them to developers. They can slow down, change format or disappear. If a feature matters to your customers, give it a fallback voice from a supported provider.
"Unlimited" plans meter something else. Some subscription services advertise unlimited speech and then cap concurrency, commercial rights or voice quality. Read the fair use paragraph before you commit a product to the plan.
Quality varies more than the price tag suggests. A Standard voice and a neural voice from the same provider can sound worlds apart, so a generous allowance on the cheaper tier may not be the voice you actually wanted.
Volume is why this matters. A read aloud feature that narrates every article for visitors, including people who rely on audio to read the web, can burn through a monthly allowance in one busy week.
Picking the Right Free Route
Your situation
Best free route
Watch out for
Weekend script or demo
gTTS or edge-tts
Unofficial, can break
Offline app or device
pyttsx3 or Piper
Flat voices from pyttsx3
Small production feature
Google Standard or Azure F0
Allowance resets monthly
High volume, fixed cost
Self hosted Kokoro or Piper
GPU time and voice licenses
Audiobook or video voiceover
A browser tool with no code
No API access
Most real projects mix two rows: a local or free cloud voice for development, and a supported provider for the version customers touch.
Before you ship, run a three sentence test on whichever voice you pick: one with a price such as $1,200.50, one with an abbreviation such as Dr. or SQL, and one with a name that is hard to pronounce. Those three lines expose most normalization problems, and they cost less than a hundred characters of your allowance.
Use Speech 2.8 HD on PicassoIA
Sometimes you do not need an integration at all, only a good voiceover file. Speech 2.8 HD on PicassoIA runs in the browser, so you can audition voices and settings before deciding what to build. At the time of writing, PicassoIA's developer API lists image and video models rather than speech, so this is a browser workflow, not a drop in replacement for the providers above.
Paste Your Script
Open the model page and paste up to 10,000 characters into the text field. Insert pauses with markers such as <#0.5#>, which adds half a second of silence. Break a long script into scenes and generate each one separately so a single bad line never forces a full redo.
Set Voice and Emotion
Choose a voice (the default is Wise_Woman) and an emotion: auto, happy, sad, angry, fearful, disgusted, surprised, calm, fluent or neutral. The language hint supports more than 40 languages, so one script can become a Spanish or Japanese version with a single dropdown change. Then tune the delivery:
Speed: 0.5 to 2.0, with 1.0 as the default
Pitch: minus 12 to plus 12 semitones
Volume: 0 to 10, with 1.0 as the default
💡 For explainer videos, try speed 1.1 and keep pitch within two semitones of zero. Bigger shifts start to sound artificial.
Export in the Right Format
MP3 is the default and reaches 256 kbps, which suits podcasts and video. WAV and FLAC are lossless for editing, and PCM gives raw audio for pipelines. Switch on subtitle metadata when you want sentence timestamps for captions.
Other models worth auditioning on the same script:
Script, voice and soundtrack are three separate jobs, and each has a model. Draft the script with Claude Sonnet 5 or Gemini 3.5 Flash, generate the narration with Speech 2.8 HD, then lay a music bed underneath at low volume, made with Lyria 3, Music 2.6 or Stable Audio 2.5.
A simple plan gets you from zero to a working voice in one sitting:
Write a 150 word script and generate it with three different voices.
Pick the best one and write down its speed, pitch and emotion settings.
Choose the free route from the table that fits your monthly volume.
Store every file you generate, so each sentence is paid for only once.
Open PicassoIA, paste one paragraph from your own project and try three voices with three different emotions. Ten minutes of auditioning tells you more than any pricing page, and it settles the voice before you write a line of integration code. While you are there, generate a thumbnail image or a short video to pair with the audio, then build from the voice you liked best.