AI Video Dubbing API: Translate Videos Automatically
An AI video dubbing API turns one video into many languages, but the pieces differ: some return audio, some return lip-synced video. See the five stages behind a request, the dubbing models on Picasso IA, a job runner pattern, and the checks to run before you publish.
Dubbing a video used to mean a translator, a voice actor, a booked studio, and a sound engineer nudging syllables into the gaps between lip movements. An AI video dubbing API squeezes that chain into a request: you send a video and a target language, and a finished clip comes back in someone else's mother tongue. The catch is that "API" means different things depending on where you look. Some services return a dubbed audio track, some return a lip-synced video, and some only hand you building blocks (transcription, translation, speech) and leave the assembly to you.
This article sorts out which is which. You will see what happens inside a dubbing request, which dubbing models you can run on Picasso IA today, how to wrap a job runner around any asynchronous dubbing service, and what to check before a translated video goes live. One honest note up front: on Picasso IA the dubbing models run in the browser, while the developer API currently lists four image and video models. Both facts matter if you plan an automated pipeline, so both are spelled out below.
What a Dubbing API Really Does
Most dubbing APIs look simple from the outside: one request in, one translated video out. Inside, the same five stages run every time, and knowing them tells you where a bad result came from.
Five Stages Behind One Request
Stage
What happens
Where it usually goes wrong
1. Transcription
Speech becomes timed text
Names, jargon, and crosstalk get misheard
2. Translation
Text moves into the target language
Literal phrasing, wrong formality level
3. Voice synthesis
Translated lines are spoken, ideally in a clone of the original voice
Flat delivery, emotion lost
4. Timing alignment
New audio is stretched or trimmed to fit each scene
Sentences overrun the cut
5. Lip sync and mux
Mouth movement is adjusted and audio is laid back on the video
Drift, mouth shapes that miss the sound
A "one call" service runs all five behind a single endpoint. That is fast and leaves little to configure. A building-block setup exposes each stage separately, which lets you put a human translator in stage two or swap the voice in stage three without redoing everything else.
Dubbing vs Subtitles vs Voiceover
Format
Original voice
Viewer effort
Mouth match
Best fit
Subtitles
Kept
Must read while watching
Not needed
Dense talk, muted autoplay
Voiceover
Muted under a translator
Listens only
None
Interviews, documentaries
AI dubbing
Cloned or replaced
Listens only
Optional
Tutorials, ads, courses, social video
The booth above is the traditional route: one script, one human voice, one session per language. It still wins for flagship campaigns. But for a library of 200 tutorials, a webinar series, or a product demo that changes every quarter, nobody books 200 sessions times ten languages. That is where an automated dubbing pipeline pays for itself.
💡 Tip: Many people scroll with the sound off, so a dub does not replace captions. Dub the audio and keep any burned-in captions in the same language as the new track.
Picking a Dubbing Route
Video Translate for Lip-Synced Output
Video Translate is the closest thing on Picasso IA to a one-call dubbing service. You upload an MP4, pick an output language, and get back a dubbed video whose lip movements are adjusted to the new audio. The language list runs past 150 entries, including regional variants such as Spanish (Mexico), Portuguese (Brazil), and Arabic (Morocco), so one model reaches most markets a small team will ever target.
It offers two modes. Speed is built for quick turnaround and precision, the default, favors output quality. The example clips on the model page took 169 and 183 seconds to generate, so plan on a few minutes per short video rather than a few seconds.
ElevenLabs Dubbing for Voice Cloning
Dubbing from ElevenLabs takes a different angle. It supports more than 90 languages and dialects, detects the source language automatically, and keeps the original speaker's voice, emotion, and timing. A cloning strength setting from 0 to 10 (default 7) decides how closely the new voice should match the original: higher for interviews where identity matters, lower for a more natural delivery in the target language.
It also accepts a public URL, such as a direct file link, a YouTube link, or a TikTok link, so you can dub a video without downloading it first. The listed output type is audio, so check what you get back for your file and lay the new track onto the original video if you need a finished MP4.
Mixing Parts Into Your Own Pipeline
If you need control per language, assemble the chain yourself: a transcription model for stage one, a text-to-speech or voice cloning model for stage three, and a lip sync model for stage five. The sections further down list the options. The price of that flexibility is orchestration work, which the job runner pattern later in this article handles.
Route
Input
Languages
Mouth match
Best for
Video Translate
MP4 upload
150+
Built in
Finished videos for new markets
ElevenLabs Dubbing
File or public URL
90+
Timing preserved
Podcasts, interviews, YouTube links
Custom chain
Anything
Depends on parts
Add a lip sync model
Per-language control, human review
How to Use Video Translate
Here is the fastest way to turn one English video into a Spanish one, using the browser version of Video Translate on Picasso IA.
Run a Clip Step by Step
Open the model page and sign in to your Picasso IA account.
Upload an MP4 in the video field. Start with a 15 to 30 second excerpt that has clear speech and a visible face, not the full forty-minute webinar.
Choose the output language. The default is English. Pick Spanish for a broad audience, or a regional entry such as Spanish (Mexico) when the campaign targets one country.
Select the mode. Leave it on precision for anything you plan to publish. Switch to speed when you are only checking that the script and pacing work.
Run the model and wait a few minutes.
Review the result with a native speaker before you queue the other nine languages.
Parameters Worth Changing
Model
Parameter
Options
Change it when
Video Translate
mode
speed, precision (default)
Drafts use speed, final cuts use precision
Video Translate
output_language
150+ names and regional variants
The audience speaks a specific regional dialect
ElevenLabs Dubbing
target_language
BCP-47 tags such as es-MX, pt-BR, en-GB
Always required
ElevenLabs Dubbing
source_language
auto (default) or a tag
Detection picks the wrong language on accented speech
ElevenLabs Dubbing
cloning_strength
0 to 10, default 7
Raise it to match the speaker, lower it for natural delivery
ElevenLabs Dubbing
source_url or file
Public link or upload
Dub a hosted video without downloading it
💡 Tip: Regional variants are not cosmetic. A Mexican Spanish dub and a Spain Spanish dub differ in vocabulary and rhythm, and a marketing clip that sounds foreign to its own audience loses the trust a dub is supposed to build.
Teachers and course creators get the most from this. A recorded lecture dubbed into Hindi, Tamil, or Portuguese (Brazil) reaches students who would otherwise rely on subtitles at reading speed.
Building a Dubbing Pipeline
What the Picasso IA API Offers
Picasso IA has a developer API at https://api.picassoia.com/v1, authenticated with a Bearer token. It follows the asynchronous pattern most AI APIs use: create a prediction for a model, poll its status, then fetch the result. The API page lists four models: two for images and two for video, namely Picasso IA Video and Seedance 2.5 Lite. The video pair produce clips with synchronized audio.
Two limits shape any automation. An account can have up to 5 predictions queued or running at once, shared across its tokens and MCP connections. And at the time of writing the page states that creating predictions needs an Infinite plan while also stating that API predictions are currently free and use no credits. Check the page for the current terms before you build around it.
The practical consequence for dubbing: no dubbing, translation, lip sync, or text-to-speech model is exposed through the API today. Those models run in the browser. A sensible split is to generate your source footage through the API, for example a 5 second Picasso IA Video clip or a 5 or 10 second Seedance 2.5 Lite clip, then run the dubbing step in the browser or through a provider whose API does expose dubbing.
A Job Runner Pattern
Any dubbing API with asynchronous jobs needs the same wrapper: submit, poll with backoff, respect the concurrency cap, and never lose a job id. The sketch below uses placeholder routes. Swap in the real endpoints and field names from whichever provider you choose.
import os
import time
import requests
from multiprocessing.pool import ThreadPool
BASE_URL = os.environ["PROVIDER_BASE_URL"] # your dubbing provider
TOKEN = os.environ["PROVIDER_TOKEN"]
HEADERS = {"Authorization": f"Bearer {TOKEN}"}
SOURCE_URL = os.environ["SOURCE_VIDEO_URL"] # public link to your MP4
LANGUAGES = ["es-MX", "pt-BR", "de", "ja", "hi"]
MAX_PARALLEL = 5 # stay under the provider cap
def submit(language):
r = requests.post(
f"{BASE_URL}/dubbing/jobs", # placeholder route
headers=HEADERS,
json={"source_url": SOURCE_URL, "target_language": language},
timeout=60,
)
r.raise_for_status()
job_id = r.json()["id"]
print(f"{language}: submitted {job_id}") # log it before waiting
return job_id
def wait(job_id, timeout=3600):
delay, deadline = 5, time.time() + timeout
while time.time() < deadline:
job = requests.get(
f"{BASE_URL}/dubbing/jobs/{job_id}", headers=HEADERS, timeout=60
).json()
if job["status"] == "succeeded":
return job["output_url"]
if job["status"] in ("failed", "canceled"):
raise RuntimeError(job.get("error", "dubbing failed"))
time.sleep(delay)
delay = min(delay * 1.5, 60) # gentle backoff
raise TimeoutError(job_id)
def dub(language):
try:
return language, wait(submit(language))
except Exception as exc: # one bad language must not stop the rest
return language, f"ERROR: {exc}"
with ThreadPool(MAX_PARALLEL) as pool:
for language, result in pool.imap(dub, LANGUAGES):
print(language, result)
Three habits in that code are worth copying. Log the job id the moment you get it, so a crash never orphans paid work. Poll with a growing delay instead of hammering the status route. And catch errors per language, so a failed Japanese dub does not stop the German one.
A quick sizing example: if each job takes about three minutes, like the example clips above, and your provider runs 5 at once, 20 languages need four rounds, so roughly 12 minutes. Treat that as an estimate, because real queues vary.
Lip Sync After Translation
Why Timing Slips After Translation
Translated sentences rarely match the length of the original. Spanish and German often run longer than English, which forces the dub to speed up, trim, or rewrite. Even when the audio fits, the mouth still forms the original shapes: a closed-lip "p" or "b" in English may land on an open vowel in the dub. Viewers notice this most on close-ups, so a talking-head clip needs more care than a screen recording with a voice over it.
Lip Sync Models to Chain
When your dubbing step returns audio only, or when you built the chain yourself, a dedicated lip sync model re-times the mouth to the new track. Picasso IA lists several in the lipsync category:
The usual order is: dub the audio, lay it onto the original video, then pass video and new audio to the lip sync model. Skip the step when the speaker is off camera, seen from behind, or small in frame, because nobody can check the mouth.
Voices and Transcripts
Transcribe the Source First
A custom chain starts with an accurate transcript. Picasso IA lists GPT-4o Transcribe, GPT-4o Mini Transcribe, and Gemini 3 Pro in the speech-to-text category. Whatever you pick, read the transcript before translating. Fix product names, acronyms, and people's names once, and every language benefits. A short glossary that stays the same across all targets prevents your brand from being translated in nine different ways.
Choose a Voice per Language
For stage three, the text-to-speech category has plenty of options:
Clone only voices you have permission to use, and keep a written record of that permission next to the project. A brand voice should stay consistent across languages, so pick one voice style per language and reuse it for the whole series.
Quality Checks Before Publishing
Three Common Mistakes
1. Skipping the native review. Automatic dubbing is good, and still wrong in ways only a native ear catches: a formal register where the brand is casual, a number read in the wrong order, a joke that translates into nonsense. Ten minutes with a bilingual reviewer is cheaper than a reshoot.
2. Forgetting what is written on screen. Dubbing changes the audio. It does not touch slides, interface text, lower thirds, or captions burned into the footage. Plan a second pass that replaces those, or the Spanish voice will narrate an English screen.
3. Ignoring the soundtrack. Music, applause, and room tone sit under the original speech. Listen for them in the dubbed file. If the music drops out or stutters, mix the original background back in under the new voice.
A short pre-publish checklist keeps batches honest:
Names and numbers read correctly in every language
Length within a few seconds of the original, with no sentence cut off
Voice consistent across all videos in the series
On-screen text replaced or intentionally left in the source language
Audio levels matched to the original, with music still present
One native listener has signed off per language
Try Dubbing on Picasso IA
You do not need a studio, a translator on retainer, or a week of editing to see how this works. Open Video Translate, upload a 20 second clip, and run it in two languages. Then run the same clip through ElevenLabs Dubbing with the cloning strength at 3 and again at 9, and listen to the difference side by side.
If you want fresh footage to dub, generate it first: Picasso IA creates images and short videos with synchronized audio, and the whole catalog is listed on the all models page. Experiment, compare, and keep the version your audience would actually press play on.