Large Language ModelsGenerate speechGenerate images
Text to Speech MCP Server: Local TTS for Claude With Kokoro
Build a text to speech MCP server that gives Claude a private voice running on your own computer. See how Kokoro's 82 million parameter model works, get a 60 line Python server, connect it to Claude Desktop and Claude Code, and weigh it against hosted voices.
Claude can draft a 2,000 word article in under a minute, and then you spend ten minutes squinting at it. A text to speech MCP server closes that gap. Claude calls a tool, a small voice model on your own machine turns the text into audio, and the result plays through your speakers before your coffee gets cold.
Kokoro is the model that makes this practical. It has just 82 million parameters, ships under the Apache 2.0 license, and produces 24 kHz speech that holds up next to much larger systems. There is no API bill, no upload of your drafts to a third party, and no rate limit. This article shows how the pieces connect, gives you a working Python server of about 60 lines, wires it into Claude Desktop and Claude Code, and then looks honestly at when a hosted voice is the better call.
Why Run Text to Speech Locally
Cloud voices sound great, but they come with strings attached. Every request leaves your machine, every character is metered, and every outage becomes your outage. A local server flips each of those points.
Privacy Without Extra Effort
When Claude reads a draft contract, a private journal entry, or an unreleased product spec aloud, the text passes through your MCP server and nothing else. Kokoro runs inference on your CPU or GPU, writes a WAV file to disk, and the trail ends there. One honest caveat: Claude itself still receives whatever you type, so "local" describes the voice step, not the whole conversation.
Cost and Offline Use
Hosted voice APIs usually charge per character or per minute. That is fine for a podcast intro and painful for an agent that narrates every answer all day. Once Kokoro's weights are downloaded, one more sentence costs a few seconds of electricity. Long documents, repeated retakes, and experiments with different voices all cost the same: nothing.
The first launch downloads the model from Hugging Face. After that the server works on a plane, in a basement, or behind a corporate firewall. Latency depends on your hardware, not on a network round trip, so behavior stays identical at 3 a.m. and at 3 p.m.
What Kokoro Actually Is
Size, License, and Training
Kokoro-82M is an open-weight text to speech model built on the StyleTTS 2 architecture with an ISTFTNet vocoder. Version 1.0 arrived on January 27, 2025, after an earlier release on December 25, 2024. The authors trained on a few hundred hours of permissively licensed and synthetic audio, which is why the Apache 2.0 license sits comfortably with commercial projects.
Spec
Value
Parameters
82 million
License
Apache 2.0
Output sample rate
24 kHz
Architecture
StyleTTS 2 with ISTFTNet vocoder
Voices in v1.0
54
Languages in v1.0
8
System dependency
espeak-ng
💡 Small does not mean weak. Kokoro's size is what lets it run on a laptop. For plain narration of English prose it is hard to separate from many paid voices in a casual listen. Expressiveness on command, such as whispering or laughing, is where larger hosted models pull ahead.
Languages and Voice Codes
Voice names follow a pattern: the first letter is the language or accent, the second is the gender. af_heart is an American English female voice, and bm_george is a British English male voice. The pipeline also needs a matching lang_code so the text is turned into phonemes correctly.
Code
Language
Example voices
a
American English
af_heart, am_michael
b
British English
bf_emma, bm_george
e
Spanish
ef_dora
f
French
ff_siwis
h
Hindi
hf_alpha
i
Italian
if_sara
j
Japanese
jf_alpha
p
Brazilian Portuguese
pf_dora
z
Mandarin Chinese
zf_xiaobei
American and British English count as one language in the "8 languages" figure. Japanese and Mandarin pull in extra packages from the misaki family, so read the project README before you switch them on.
How the MCP Server Works
The Three Moving Parts
The setup has only three pieces, and each one has a narrow job:
The client. Claude Desktop or Claude Code speaks the Model Context Protocol and decides when a tool is worth calling.
The server. A small Python process that the client starts over stdio. It exposes a few tools and nothing else.
Kokoro. Loaded once into memory inside that process, so later calls skip the slow startup.
The flow is short. Claude decides to call speak, sends the text plus a voice name, the server synthesizes the audio, plays it, and returns the file path as plain text so Claude can tell you where the recording landed.
Ready-Made Servers Worth Trying
You do not have to write everything yourself. Several community projects already wrap Kokoro for MCP:
kristofferv98/MCP_tts_server offers more than one TTS engine, including Kokoro, with streaming playback.
kokoro-tts-mcp by scottschram runs Kokoro-82M with MLX acceleration on Apple Silicon.
koroko-speech-mcp by hammeiam is a compact speech server built around Kokoro.
Pre-built servers save an afternoon. Writing your own, as below, costs about an hour and gives you full control over text cleanup, voice defaults, and where files land.
Build the Server Step by Step
Install the Dependencies
Use a virtual environment so the PyTorch stack stays away from your system Python. Python 3.10, 3.11, or 3.12 is the safe choice.
Kokoro also needs the espeak-ng phonemizer on your system:
macOS: brew install espeak-ng
Debian or Ubuntu: sudo apt-get install espeak-ng
Windows: install the espeak-ng release package, then open a fresh terminal
Write the Server File
Save this as kokoro_server.py. It exposes two tools, loads the English pipeline at startup, cleans markdown out of the text, and plays audio without blocking Claude.
import os
import re
import time
from pathlib import Path
import numpy as np
import sounddevice as sd
import soundfile as sf
from kokoro import KPipeline
from mcp.server.fastmcp import FastMCP
SAMPLE_RATE = 24000
OUT_DIR = Path(os.environ.get("KOKORO_OUT", Path.home() / "kokoro_audio"))
OUT_DIR.mkdir(parents=True, exist_ok=True)
mcp = FastMCP("kokoro-tts")
pipelines = {}
def get_pipeline(lang_code: str) -> KPipeline:
if lang_code not in pipelines:
pipelines[lang_code] = KPipeline(lang_code=lang_code)
return pipelines[lang_code]
def clean_text(text: str) -> str:
text = re.sub(r"`{3}.*?`{3}", " code block omitted. ", text, flags=re.S)
text = re.sub(r"https?://\S+", "link", text)
text = re.sub(r"[#*_`>]+", "", text)
return re.sub(r"[ \t]+", " ", text).strip()
@mcp.tool()
def speak(text: str, voice: str = "af_heart", speed: float = 1.0,
lang_code: str = "a", play: bool = True) -> str:
"""Read text aloud with Kokoro. Returns the path of the saved WAV file."""
pipeline = get_pipeline(lang_code)
parts = []
for _, _, audio in pipeline(clean_text(text), voice=voice, speed=speed):
if audio is not None:
parts.append(audio.detach().cpu().numpy())
if not parts:
return "No audio was produced. Check the text and the voice name."
wave = np.concatenate(parts)
path = OUT_DIR / f"speech_{time.strftime('%Y%m%d_%H%M%S')}.wav"
sf.write(path, wave, SAMPLE_RATE)
if play:
sd.play(wave, SAMPLE_RATE)
return f"Saved {len(wave) / SAMPLE_RATE:.1f}s of audio to {path}"
@mcp.tool()
def list_voices() -> str:
"""List a few Kokoro voices and the lang_code each one needs."""
return (
"a: af_heart, af_bella, am_michael | b: bf_emma, bm_george | "
"e: ef_dora | f: ff_siwis | j: jf_alpha"
)
if __name__ == "__main__":
get_pipeline("a")
mcp.run()
Three design choices matter here. The pipeline is cached per language, so switching between English and Spanish does not reload anything twice. Playback uses sd.play, which returns immediately, so Claude never waits for the audio to finish. And nothing in the file calls print, because on a stdio server stdout belongs to the protocol. A stray print corrupts the message stream.
Register It With Claude
For Claude Desktop, open claude_desktop_config.json. On macOS it lives in ~/Library/Application Support/Claude/, and on Windows in %APPDATA%\Claude\. Point the command at the interpreter inside your virtual environment, not at a bare python.
claude mcp add kokoro-tts -- /absolute/path/to/.venv/bin/python /absolute/path/to/kokoro_server.py
Restart the client, then test with a plain sentence: "Use the speak tool to say hello in the bm_george voice." If you hear a British gentleman greet you, the whole chain works.
Make Claude Speak Naturally
Write for the Ear
Text that reads well on screen often sounds clumsy aloud. Give Claude a standing instruction so it writes for listening from the first draft:
When I ask you to read something aloud, call the speak tool. Write for the ear: short sentences, no bullet symbols, no URLs, and spell out numbers or acronyms when they are hard to say.
The clean_text function in the server is your safety net, but it cannot rescue a sentence that has no natural pauses. A few habits improve the output at once:
Split long passages into paragraphs. The pipeline breaks text on line breaks by default, which keeps each chunk short and the pacing steady.
Respell tricky names phonetically. If a brand name comes out wrong, write it the way it sounds.
Keep speed between 0.9 and 1.1. Beyond that range, speech starts to feel hurried or sluggish.
Where Hands-Free Voice Pays Off
A voice tool earns its place in moments when your eyes or hands are busy:
Cooking. Ask Claude to adapt a recipe for six people and read the steps aloud while flour dusts your fingers.
Proofreading. Hearing a draft exposes clunky rhythm and repeated words that your eyes slide past.
Commutes and walks. Request a spoken summary of yesterday's notes or a long thread before you sit down.
Accessibility. Spoken output can help people with low vision or reading difficulties work with long text more comfortably.
Speed, Hardware, and Common Fixes
What to Expect on Your Machine
Kokoro runs on an ordinary laptop CPU, and a GPU or Apple Silicon speeds it up further. Rather than trust anyone's benchmark, run your own test: send a 200 word paragraph and compare the render time with the playback time. If rendering finishes first, you have real-time headroom and long documents will feel instant.
Two habits keep the experience smooth. Load the pipeline at startup, as the server above does, because the first synthesis is always the slowest. And keep the model resident: the MCP client keeps your server process alive between calls, so the weights stay in memory.
💡 Tip: Run the server once from a terminal before you register it. Any missing dependency shows up as a readable Python error instead of a vague "server failed" banner in the client.
Five Problems and Their Fixes
Symptom
Likely cause
Fix
Client reports the server failed to start
Config points at the wrong Python
Use the absolute path to the virtual environment interpreter
Tool call hangs, then errors
A print call wrote to stdout
Remove prints and log to stderr instead
Phoneme or espeak error
espeak-ng is not on the PATH
Install it, then open a new terminal
File saved but no sound
Audio output device or PortAudio issue
Install PortAudio on Linux, or pick a device in sounddevice
Names pronounced oddly
The phonemizer guessed wrong
Respell the name as it sounds
Local Kokoro Versus Hosted Voices
Local is not always the answer. Kokoro is excellent for private, repeatable, free narration, but some jobs need more than 54 voices can offer.
When Hosted Voices Win
Reach for a hosted model when you need voice cloning, wide language support, or fine emotional control. The text to speech collection on Picasso IA has plenty of options, all one click away:
A sensible split: use Kokoro for daily reading, drafts, and anything private, and switch to a hosted voice for the final take of a video, podcast, or product demo.
Pairing With a Language Model
The server only speaks. What it speaks depends on the model writing the words. For long drafts and careful rewriting, Claude Sonnet 5 and Claude Opus 4.7 are strong choices, while Claude 4.5 Haiku keeps quick spoken answers snappy. Match the model to the job, then let Kokoro handle the delivery.
Your Turn: Create With Picasso IA
You now have a private voice for Claude that costs nothing per sentence. The next step is to give your projects a face. The same afternoon you finish this server, you can build thumbnails, article headers, and scene art on Picasso IA to sit next to your audio.
Try Seedream 4.5 for photorealistic scenes, or Flux 2 Pro when you want sharp detail and clean composition. When a project needs a studio voice, test Speech 2.8 HD alongside your local Kokoro setup and compare the results by ear.
Open Picasso IA, write one prompt, and see what you get. Then try a second one with a different camera angle. Experiment freely, because the best prompts come from trying, listening, and adjusting.