Generate speechLarge Language ModelsGenerate music

Text to Speech API Free: Pricing, Python and Unlimited Options

Every free text to speech API has different rules: monthly allowances that reset, twelve month trials that expire, and open source models with no meter. See real prices, working Python code, and a browser workflow for testing voices first.

Text to Speech API Free: Pricing, Python and Unlimited Options
Cristian Da Conceicao
Founder of Picasso IA

Type "free text to speech API" into a search bar and you get three very different answers. A monthly allowance from a cloud giant. A Python package that quietly borrows someone else's servers. An open source model you run on your own machine. All three are free, but only one is truly unlimited, and it is not the one most people expect.

This article sorts the options by what they really cost, shows working Python for each, and ends with a way to audition voices in your browser before you write any integration code. The figures come from provider pricing pages and published summaries as of October 2026. Prices move, so confirm the numbers before you build a budget on them.

What Free Really Means

Providers use the word "free" for three different arrangements, and mixing them up is how a prototype turns into a surprise invoice.

Recurring Monthly Allowances

Google Cloud and Microsoft Azure reset a free allowance every month with no expiry date. Google gives 4 million characters of Standard voices, and Azure's F0 tier gives 0.5 million characters of neural voices. Stay under the line and the bill is zero indefinitely. Go over on Google and you pay the normal per character rate. Azure's F0 tier is capped, so you hit a ceiling instead of a bill.

Twelve Month Trials

Amazon Polly's allowance applies to new accounts for their first 12 months. AWS lists 1 million neural characters a month, 500 thousand long form characters and 100 thousand generative characters, and published summaries put standard voices at 5 million. After month twelve the same traffic costs $4 per million characters on standard voices and $16 on neural. A prototype that is free in January can carry a real line item by the next January.

ElevenLabs sits in between. Its free plan gives roughly 10,000 characters a month, which is only ten to fifteen minutes of audio, and it asks for attribution and rules out commercial use. That is plenty for testing a voice and far too small for a product. The same voice family is also available on PicassoIA as ElevenLabs v3.

Open Source and Unlimited

The only truly unlimited option is software you run yourself. Piper runs on a plain CPU with more than 100 voices across 30+ languages. Kokoro is an 82 million parameter model that produces natural speech and reportedly runs close to 100 times faster than real time on a GPU. Nobody meters you because nobody is hosting you. The price is your own hardware, your setup time, and the job of checking each voice's license before you sell anything made with it.

Startup team sorting free tier options on a whiteboard

Free Tiers Compared Side by Side

The Pricing Table

Provider and tierFree allowanceDoes it expire?Paid rate after
Google Cloud Standard4M characters a monthNo$4 per 1M
Google Cloud Neural2About 1M characters a monthNo$16 per 1M
Azure Neural (F0)0.5M characters a monthNoPay as you go
Amazon Polly Standard5M characters a monthAfter 12 months$4 per 1M
Amazon Polly Neural1M characters a monthAfter 12 months$16 per 1M
ElevenLabs FreeAbout 10K characters a monthNoPaid plans
Piper or KokoroNo limitNoYour hardware

💡 Allowances are counted per account, per month and per voice tier. Standard and premium voices draw from separate pools, so test the exact voice you plan to ship.

What a Million Characters Buys

AWS's own pricing examples work out to about 23 hours and 8 minutes of speech per million characters. Use that as a rule of thumb. Google's 4 million Standard characters is roughly 90 hours of audio a month, and Azure's 0.5 million is about 11 hours.

A 1,500 word blog post runs near 9,000 characters. At that size, Azure's free tier narrates about 55 posts a month and Google's narrates more than 400. A read aloud button on a mid size blog fits inside a free tier, as long as you store each audio file instead of regenerating it on every page view.

Count carefully. Most providers meter every character you send, including spaces, punctuation and markup tags, so strip HTML before synthesis and send plain text. Measure a week of real traffic before you trust any estimate.

Printed invoice, calculator and espresso on a wooden desk

Python Options That Cost Nothing

Three packages handle most quick scripts. They install with pip, need no billing account and work in a few lines, but they are not equal.

gTTS in Four Lines

from gtts import gTTS

tts = gTTS("Your order has shipped.", lang="en")
tts.save("order.mp3")

gTTS is an unofficial wrapper around the speech endpoint behind Google Translate. There is no credential and no published quota, which is exactly the problem: Google can throttle or change the endpoint without notice. It suits a classroom script. It is risky for a customer facing feature.

Developer typing Python code in a coworking space

pyttsx3 Runs Offline

import pyttsx3

engine = pyttsx3.init()
engine.setProperty("rate", 170)
engine.save_to_file("Your order has shipped.", "order.wav")
engine.runAndWait()

pyttsx3 drives the voices already installed in your operating system: SAPI5 on Windows, NSSpeechSynthesizer on macOS and eSpeak on Linux. There is no network call and no limit. There is also a voice that sounds like a 2005 GPS unit, so judge it by your own ears before committing.

edge-tts for Better Voices

import asyncio
import edge_tts

async def main():
    speech = edge_tts.Communicate("Your order has shipped.", "en-US-AriaNeural")
    await speech.save("order.mp3")

asyncio.run(main())

edge-tts talks to the same read aloud service used by the Edge browser. The voices are neural quality and the package is free, but it is an unofficial client with no service agreement behind it. Treat it like gTTS: good for personal projects, a gamble when paying users depend on it.

Local Neural Models

Piper and Kokoro are the two to try first. Piper takes a downloaded voice file and writes audio straight from the command line:

echo "Your order has shipped." | piper --model en_US-lessac-medium.onnx --output_file order.wav

It is light enough for a Raspberry Pi 4, which makes it the usual pick for offline kiosks, home automation and anything that must keep working without internet. Kokoro needs more memory but sounds noticeably more natural, and it flies on a GPU.

Raspberry Pi connected to a small speaker on a workbench

Calling a Cloud API From Python

When you need an SLA, SSML markup or dozens of languages, a cloud provider is the honest route. The pattern is the same everywhere: authenticate, send text, receive audio bytes.

A Google Cloud Example

from google.cloud import texttospeech

client = texttospeech.TextToSpeechClient()

response = client.synthesize_speech(
    input=texttospeech.SynthesisInput(text="Your order has shipped."),
    voice=texttospeech.VoiceSelectionParams(
        language_code="en-US", name="en-US-Standard-C"
    ),
    audio_config=texttospeech.AudioConfig(
        audio_encoding=texttospeech.AudioEncoding.MP3
    ),
)

with open("order.mp3", "wb") as f:
    f.write(response.audio_content)

Authentication comes from the GOOGLE_APPLICATION_CREDENTIALS environment variable, which points at a service account file, so no secret ever sits inside the script. Swap en-US-Standard-C for a Neural2 voice and the call stays identical. Only the pool you draw from, and the price, changes.

Technician walking through a data center aisle

Cache Before You Pay

import hashlib
import pathlib

def cached_speech(text, synthesize):
    name = hashlib.sha256(text.encode()).hexdigest() + ".mp3"
    path = pathlib.Path("audio_cache") / name
    if not path.exists():
        path.parent.mkdir(exist_ok=True)
        path.write_bytes(synthesize(text))
    return path

Hashing the text means a repeated sentence costs nothing the second time. For a read aloud feature this habit alone usually keeps traffic inside the free tier. Two more habits help: split long text at sentence boundaries, since most providers cap a single request at a few thousand characters, and wrap each call in a retry with exponential backoff for rate limit responses.

Unlimited Options and Their Catch

Search for "unlimited" and the same three promises come up. Each one has a cost.

Self hosting costs time. You install models, manage dependencies, watch GPU memory and update everything when a release breaks something. For a few thousand requests a month a free cloud tier is cheaper than your hours. At millions of requests, running Piper or Kokoro on your own server starts to win.

Unofficial endpoints break. gTTS and edge-tts depend on services whose owners never promised them to developers. They can slow down, change format or disappear. If a feature matters to your customers, give it a fallback voice from a supported provider.

"Unlimited" plans meter something else. Some subscription services advertise unlimited speech and then cap concurrency, commercial rights or voice quality. Read the fair use paragraph before you commit a product to the plan.

Quality varies more than the price tag suggests. A Standard voice and a neural voice from the same provider can sound worlds apart, so a generous allowance on the cheaper tier may not be the voice you actually wanted.

Volume is why this matters. A read aloud feature that narrates every article for visitors, including people who rely on audio to read the web, can burn through a monthly allowance in one busy week.

Man listening to audio on a smartphone at a cafe terrace

Picking the Right Free Route

Your situationBest free routeWatch out for
Weekend script or demogTTS or edge-ttsUnofficial, can break
Offline app or devicepyttsx3 or PiperFlat voices from pyttsx3
Small production featureGoogle Standard or Azure F0Allowance resets monthly
High volume, fixed costSelf hosted Kokoro or PiperGPU time and voice licenses
Audiobook or video voiceoverA browser tool with no codeNo API access

Most real projects mix two rows: a local or free cloud voice for development, and a supported provider for the version customers touch.

Before you ship, run a three sentence test on whichever voice you pick: one with a price such as $1,200.50, one with an abbreviation such as Dr. or SQL, and one with a name that is hard to pronounce. Those three lines expose most normalization problems, and they cost less than a hundred characters of your allowance.

Woman recording a voiceover with a condenser microphone

Use Speech 2.8 HD on PicassoIA

Sometimes you do not need an integration at all, only a good voiceover file. Speech 2.8 HD on PicassoIA runs in the browser, so you can audition voices and settings before deciding what to build. At the time of writing, PicassoIA's developer API lists image and video models rather than speech, so this is a browser workflow, not a drop in replacement for the providers above.

Paste Your Script

Open the model page and paste up to 10,000 characters into the text field. Insert pauses with markers such as <#0.5#>, which adds half a second of silence. Break a long script into scenes and generate each one separately so a single bad line never forces a full redo.

Set Voice and Emotion

Choose a voice (the default is Wise_Woman) and an emotion: auto, happy, sad, angry, fearful, disgusted, surprised, calm, fluent or neutral. The language hint supports more than 40 languages, so one script can become a Spanish or Japanese version with a single dropdown change. Then tune the delivery:

  • Speed: 0.5 to 2.0, with 1.0 as the default
  • Pitch: minus 12 to plus 12 semitones
  • Volume: 0 to 10, with 1.0 as the default

💡 For explainer videos, try speed 1.1 and keep pitch within two semitones of zero. Bigger shifts start to sound artificial.

Export in the Right Format

MP3 is the default and reaches 256 kbps, which suits podcasts and video. WAV and FLAC are lossless for editing, and PCM gives raw audio for pipelines. Switch on subtitle metadata when you want sentence timestamps for captions.

Woman adjusting audio settings on a monitor in a bright studio

Other models worth auditioning on the same script:

Make Your First Voiceover

Script, voice and soundtrack are three separate jobs, and each has a model. Draft the script with Claude Sonnet 5 or Gemini 3.5 Flash, generate the narration with Speech 2.8 HD, then lay a music bed underneath at low volume, made with Lyria 3, Music 2.6 or Stable Audio 2.5.

Musician's hands on an analog mixing console

A simple plan gets you from zero to a working voice in one sitting:

  1. Write a 150 word script and generate it with three different voices.
  2. Pick the best one and write down its speed, pitch and emotion settings.
  3. Choose the free route from the table that fits your monthly volume.
  4. Store every file you generate, so each sentence is paid for only once.

Open PicassoIA, paste one paragraph from your own project and try three voices with three different emotions. Ten minutes of auditioning tells you more than any pricing page, and it settles the voice before you write a line of integration code. While you are there, generate a thumbnail image or a short video to pair with the audio, then build from the voice you liked best.

Share this article