Lipsync videosGenerate speechGenerate videos

AI Video Dubbing API: Translate Videos Automatically

An AI video dubbing API turns one video into many languages, but the pieces differ: some return audio, some return lip-synced video. See the five stages behind a request, the dubbing models on Picasso IA, a job runner pattern, and the checks to run before you publish.

AI Video Dubbing API: Translate Videos Automatically
Cristian Da Conceicao
Founder of Picasso IA

Dubbing a video used to mean a translator, a voice actor, a booked studio, and a sound engineer nudging syllables into the gaps between lip movements. An AI video dubbing API squeezes that chain into a request: you send a video and a target language, and a finished clip comes back in someone else's mother tongue. The catch is that "API" means different things depending on where you look. Some services return a dubbed audio track, some return a lip-synced video, and some only hand you building blocks (transcription, translation, speech) and leave the assembly to you.

This article sorts out which is which. You will see what happens inside a dubbing request, which dubbing models you can run on Picasso IA today, how to wrap a job runner around any asynchronous dubbing service, and what to check before a translated video goes live. One honest note up front: on Picasso IA the dubbing models run in the browser, while the developer API currently lists four image and video models. Both facts matter if you plan an automated pipeline, so both are spelled out below.

A man watching a dubbed cooking video on his phone with headphones around his neck

What a Dubbing API Really Does

Most dubbing APIs look simple from the outside: one request in, one translated video out. Inside, the same five stages run every time, and knowing them tells you where a bad result came from.

Five Stages Behind One Request

StageWhat happensWhere it usually goes wrong
1. TranscriptionSpeech becomes timed textNames, jargon, and crosstalk get misheard
2. TranslationText moves into the target languageLiteral phrasing, wrong formality level
3. Voice synthesisTranslated lines are spoken, ideally in a clone of the original voiceFlat delivery, emotion lost
4. Timing alignmentNew audio is stretched or trimmed to fit each sceneSentences overrun the cut
5. Lip sync and muxMouth movement is adjusted and audio is laid back on the videoDrift, mouth shapes that miss the sound

A "one call" service runs all five behind a single endpoint. That is fast and leaves little to configure. A building-block setup exposes each stage separately, which lets you put a human translator in stage two or swap the voice in stage three without redoing everything else.

Dubbing vs Subtitles vs Voiceover

FormatOriginal voiceViewer effortMouth matchBest fit
SubtitlesKeptMust read while watchingNot neededDense talk, muted autoplay
VoiceoverMuted under a translatorListens onlyNoneInterviews, documentaries
AI dubbingCloned or replacedListens onlyOptionalTutorials, ads, courses, social video

A voice actor speaking into a condenser microphone in a padded recording booth

The booth above is the traditional route: one script, one human voice, one session per language. It still wins for flagship campaigns. But for a library of 200 tutorials, a webinar series, or a product demo that changes every quarter, nobody books 200 sessions times ten languages. That is where an automated dubbing pipeline pays for itself.

💡 Tip: Many people scroll with the sound off, so a dub does not replace captions. Dub the audio and keep any burned-in captions in the same language as the new track.

Picking a Dubbing Route

Video Translate for Lip-Synced Output

Video Translate is the closest thing on Picasso IA to a one-call dubbing service. You upload an MP4, pick an output language, and get back a dubbed video whose lip movements are adjusted to the new audio. The language list runs past 150 entries, including regional variants such as Spanish (Mexico), Portuguese (Brazil), and Arabic (Morocco), so one model reaches most markets a small team will ever target.

It offers two modes. Speed is built for quick turnaround and precision, the default, favors output quality. The example clips on the model page took 169 and 183 seconds to generate, so plan on a few minutes per short video rather than a few seconds.

ElevenLabs Dubbing for Voice Cloning

Dubbing from ElevenLabs takes a different angle. It supports more than 90 languages and dialects, detects the source language automatically, and keeps the original speaker's voice, emotion, and timing. A cloning strength setting from 0 to 10 (default 7) decides how closely the new voice should match the original: higher for interviews where identity matters, lower for a more natural delivery in the target language.

It also accepts a public URL, such as a direct file link, a YouTube link, or a TikTok link, so you can dub a video without downloading it first. The listed output type is audio, so check what you get back for your file and lay the new track onto the original video if you need a finished MP4.

Mixing Parts Into Your Own Pipeline

If you need control per language, assemble the chain yourself: a transcription model for stage one, a text-to-speech or voice cloning model for stage three, and a lip sync model for stage five. The sections further down list the options. The price of that flexibility is orchestration work, which the job runner pattern later in this article handles.

RouteInputLanguagesMouth matchBest for
Video TranslateMP4 upload150+Built inFinished videos for new markets
ElevenLabs DubbingFile or public URL90+Timing preservedPodcasts, interviews, YouTube links
Custom chainAnythingDepends on partsAdd a lip sync modelPer-language control, human review

An overhead view of a desk with a laptop, a printed world map, sticky notes, earbuds, and a mug

How to Use Video Translate

Here is the fastest way to turn one English video into a Spanish one, using the browser version of Video Translate on Picasso IA.

Run a Clip Step by Step

  1. Open the model page and sign in to your Picasso IA account.
  2. Upload an MP4 in the video field. Start with a 15 to 30 second excerpt that has clear speech and a visible face, not the full forty-minute webinar.
  3. Choose the output language. The default is English. Pick Spanish for a broad audience, or a regional entry such as Spanish (Mexico) when the campaign targets one country.
  4. Select the mode. Leave it on precision for anything you plan to publish. Switch to speed when you are only checking that the script and pacing work.
  5. Run the model and wait a few minutes.
  6. Review the result with a native speaker before you queue the other nine languages.

Three colleagues around an oak table reviewing a paused video on a laptop

Parameters Worth Changing

ModelParameterOptionsChange it when
Video Translatemodespeed, precision (default)Drafts use speed, final cuts use precision
Video Translateoutput_language150+ names and regional variantsThe audience speaks a specific regional dialect
ElevenLabs Dubbingtarget_languageBCP-47 tags such as es-MX, pt-BR, en-GBAlways required
ElevenLabs Dubbingsource_languageauto (default) or a tagDetection picks the wrong language on accented speech
ElevenLabs Dubbingcloning_strength0 to 10, default 7Raise it to match the speaker, lower it for natural delivery
ElevenLabs Dubbingsource_url or filePublic link or uploadDub a hosted video without downloading it

💡 Tip: Regional variants are not cosmetic. A Mexican Spanish dub and a Spain Spanish dub differ in vocabulary and rhythm, and a marketing clip that sounds foreign to its own audience loses the trust a dub is supposed to build.

Teachers and course creators get the most from this. A recorded lecture dubbed into Hindi, Tamil, or Portuguese (Brazil) reaches students who would otherwise rely on subtitles at reading speed.

A teacher at a whiteboard in an empty classroom with a camera on a tripod recording the lecture

Building a Dubbing Pipeline

What the Picasso IA API Offers

Picasso IA has a developer API at https://api.picassoia.com/v1, authenticated with a Bearer token. It follows the asynchronous pattern most AI APIs use: create a prediction for a model, poll its status, then fetch the result. The API page lists four models: two for images and two for video, namely Picasso IA Video and Seedance 2.5 Lite. The video pair produce clips with synchronized audio.

Two limits shape any automation. An account can have up to 5 predictions queued or running at once, shared across its tokens and MCP connections. And at the time of writing the page states that creating predictions needs an Infinite plan while also stating that API predictions are currently free and use no credits. Check the page for the current terms before you build around it.

The practical consequence for dubbing: no dubbing, translation, lip sync, or text-to-speech model is exposed through the API today. Those models run in the browser. A sensible split is to generate your source footage through the API, for example a 5 second Picasso IA Video clip or a 5 or 10 second Seedance 2.5 Lite clip, then run the dubbing step in the browser or through a provider whose API does expose dubbing.

A developer typing on a laptop at a cafe window table with a coffee and a notebook

A Job Runner Pattern

Any dubbing API with asynchronous jobs needs the same wrapper: submit, poll with backoff, respect the concurrency cap, and never lose a job id. The sketch below uses placeholder routes. Swap in the real endpoints and field names from whichever provider you choose.

import os
import time
import requests
from multiprocessing.pool import ThreadPool

BASE_URL = os.environ["PROVIDER_BASE_URL"]      # your dubbing provider
TOKEN = os.environ["PROVIDER_TOKEN"]
HEADERS = {"Authorization": f"Bearer {TOKEN}"}
SOURCE_URL = os.environ["SOURCE_VIDEO_URL"]     # public link to your MP4
LANGUAGES = ["es-MX", "pt-BR", "de", "ja", "hi"]
MAX_PARALLEL = 5                                # stay under the provider cap


def submit(language):
    r = requests.post(
        f"{BASE_URL}/dubbing/jobs",             # placeholder route
        headers=HEADERS,
        json={"source_url": SOURCE_URL, "target_language": language},
        timeout=60,
    )
    r.raise_for_status()
    job_id = r.json()["id"]
    print(f"{language}: submitted {job_id}")    # log it before waiting
    return job_id


def wait(job_id, timeout=3600):
    delay, deadline = 5, time.time() + timeout
    while time.time() < deadline:
        job = requests.get(
            f"{BASE_URL}/dubbing/jobs/{job_id}", headers=HEADERS, timeout=60
        ).json()
        if job["status"] == "succeeded":
            return job["output_url"]
        if job["status"] in ("failed", "canceled"):
            raise RuntimeError(job.get("error", "dubbing failed"))
        time.sleep(delay)
        delay = min(delay * 1.5, 60)            # gentle backoff
    raise TimeoutError(job_id)


def dub(language):
    try:
        return language, wait(submit(language))
    except Exception as exc:                    # one bad language must not stop the rest
        return language, f"ERROR: {exc}"


with ThreadPool(MAX_PARALLEL) as pool:
    for language, result in pool.imap(dub, LANGUAGES):
        print(language, result)

Three habits in that code are worth copying. Log the job id the moment you get it, so a crash never orphans paid work. Poll with a growing delay instead of hammering the status route. And catch errors per language, so a failed Japanese dub does not stop the German one.

A quick sizing example: if each job takes about three minutes, like the example clips above, and your provider runs 5 at once, 20 languages need four rounds, so roughly 12 minutes. Treat that as an estimate, because real queues vary.

Lip Sync After Translation

Why Timing Slips After Translation

Translated sentences rarely match the length of the original. Spanish and German often run longer than English, which forces the dub to speed up, trim, or rewrite. Even when the audio fits, the mouth still forms the original shapes: a closed-lip "p" or "b" in English may land on an open vowel in the dub. Viewers notice this most on close-ups, so a talking-head clip needs more care than a screen recording with a voice over it.

A presenter speaking to a camera on a tripod with a softbox light in a small home studio

Lip Sync Models to Chain

When your dubbing step returns audio only, or when you built the chain yourself, a dedicated lip sync model re-times the mouth to the new track. Picasso IA lists several in the lipsync category:

ModelWhat the listing promises
Lipsync 2 ProSync lips to audio
React 1Realistic lipsync on any video
Kling Lip SyncMatch mouth to audio in any video
Lipsync PrecisionDubbing with a precision focus
Lipsync SpeedDubbing in seconds, speed first

The usual order is: dub the audio, lay it onto the original video, then pass video and new audio to the lip sync model. Skip the step when the speaker is off camera, seen from behind, or small in frame, because nobody can check the mouth.

Voices and Transcripts

Transcribe the Source First

A custom chain starts with an accurate transcript. Picasso IA lists GPT-4o Transcribe, GPT-4o Mini Transcribe, and Gemini 3 Pro in the speech-to-text category. Whatever you pick, read the transcript before translating. Fix product names, acronyms, and people's names once, and every language benefits. A short glossary that stays the same across all targets prevents your brand from being translated in nine different ways.

A sound engineer's hands on the faders of an analog mixing console

Choose a Voice per Language

For stage three, the text-to-speech category has plenty of options:

ModelWhy you would pick it
ElevenLabs v2 MultilingualVoiceover in 30+ languages
Gemini 3.1 Flash TTS30 voices across 70+ languages
Speech 2.8 HDStudio-quality voiceovers
MiniMax Voice CloningCustom voices from a sample
Qwen3 TTSClone a voice or design your own

Clone only voices you have permission to use, and keep a written record of that permission next to the project. A brand voice should stay consistent across languages, so pick one voice style per language and reuse it for the whole series.

Quality Checks Before Publishing

Three Common Mistakes

1. Skipping the native review. Automatic dubbing is good, and still wrong in ways only a native ear catches: a formal register where the brand is casual, a number read in the wrong order, a joke that translates into nonsense. Ten minutes with a bilingual reviewer is cheaper than a reshoot.

A bilingual reviewer with headphones listening carefully to a dubbed clip at a laptop

2. Forgetting what is written on screen. Dubbing changes the audio. It does not touch slides, interface text, lower thirds, or captions burned into the footage. Plan a second pass that replaces those, or the Spanish voice will narrate an English screen.

3. Ignoring the soundtrack. Music, applause, and room tone sit under the original speech. Listen for them in the dubbed file. If the music drops out or stutters, mix the original background back in under the new voice.

A short pre-publish checklist keeps batches honest:

  • Names and numbers read correctly in every language
  • Length within a few seconds of the original, with no sentence cut off
  • Voice consistent across all videos in the series
  • On-screen text replaced or intentionally left in the source language
  • Audio levels matched to the original, with music still present
  • One native listener has signed off per language

Try Dubbing on Picasso IA

You do not need a studio, a translator on retainer, or a week of editing to see how this works. Open Video Translate, upload a 20 second clip, and run it in two languages. Then run the same clip through ElevenLabs Dubbing with the cloning strength at 3 and again at 9, and listen to the difference side by side.

If you want fresh footage to dub, generate it first: Picasso IA creates images and short videos with synchronized audio, and the whole catalog is listed on the all models page. Experiment, compare, and keep the version your audience would actually press play on.

Share this article