Generate videosLipsync videosVisual Effects

Kling Avatar 2.0 API: Lip Sync Pricing and Examples

Kling Avatar 2.0 turns one photo and one audio file into a talking video. This article lists per-second API prices on fal, PoYo, and Picsart, shows a working request, runs the cost math on real clip lengths, and compares talking-photo models you can run on PicassoIA.

Kling Avatar 2.0 API: Lip Sync Pricing and Examples
Cristian Da Conceicao
Founder of Picasso IA

Kling Avatar 2.0 takes one portrait photo and one audio file and returns a talking video where the lips, eyes, and head motion follow the speech. The model is sold per second of audio, which makes Kling Avatar 2.0 API lip sync pricing easy to predict once you know which provider you call and which tier you pick. This article puts the real numbers side by side: $0.056 per second on fal, $0.035 on PoYo, 2 credits on Picsart, with Pro rates roughly double each of those. You also get a working request, three production-style examples, input rules that prevent bad renders, and a tutorial for the closest lip sync models on PicassoIA.

What Kling Avatar 2.0 Does

Kling Avatar 2.0 is an audio-driven avatar model from the Kling family. You hand it a still image and a speech recording, and it animates the face so the mouth shapes follow the syllables, the eyes hold contact with the lens, and the head and shoulders move the way a real speaker's would. No camera footage and no 3D rig is involved, and the identity of the person in the photo stays stable from the first frame to the last.

A photographer lighting a studio portrait of a smiling woman in a navy blazer

One Photo, One Audio File

The input contract is short. Every provider I checked asks for the same three things:

  • One image of a person, animal, cartoon, or stylized character. PoYo accepts JPEG, PNG, and WebP files up to 10 MB.
  • One audio file with the speech. PoYo accepts 2 to 60 seconds and up to 5 MB, while Hedra lists a range of 2 to 300 seconds, so check the limit of the provider you plan to call.
  • An optional text prompt that describes emotion, gesture, or camera behavior, for example "calm presenter, small nod at the end of each sentence."

The output is an MP4. Hedra lists 16:9, 9:16, and 1:1 as supported aspect ratios, which fits YouTube, Reels, and square feed posts.

Standard vs Pro Tiers

Both tiers take identical inputs. The difference is output quality and price.

StandardPro
Positioned forBulk generation, rapid prototyping, educational videosMarketing videos, client work, hero content
Price versus the other tierBaselineAbout 2x Standard
fal endpointfal-ai/kling-video/ai-avatar/v2/standardfal-ai/kling-video/ai-avatar/v2/pro
InputsImage, audio, optional promptSame

💡 Resolution claims differ by provider. APIPass lists 720p for Standard and 1080p for Pro, while Hedra lists 720p native for Pro. Render one test clip and check the file before you commit a batch.

Kling Avatar 2.0 API Pricing Compared

Pricing is where providers differ most. Billing is per second of audio, so a longer voiceover costs more no matter how simple the photo is.

A desk with a calculator, notebook, and espresso cup seen from above

Per-Second Rates by Provider

ProviderStandardProBilling notes
fal$0.056 per second$0.115 per secondRates as quoted by PoYo and Hedra
PoYo$0.035 per second (7 credits)$0.070 per second (14 credits)Audio duration rounded up to the next second
Picsart API2 credits per second4 credits per secondCredit price depends on your Picsart plan
APIPassCredits per secondCredits per secondThe figures I found were inconsistent

PoYo's page puts its Standard rate about 38% under fal's and its Pro rate about 39% under, which matches the arithmetic: $0.035 against $0.056, and $0.070 against $0.115.

💡 Check before you budget. Two snapshots of the APIPass pricing page disagreed by roughly a factor of ten, so I left its rate out of the cost math below. Confirm any provider's price in its own dashboard before you queue a large batch.

Cost Examples at Real Volumes

Billing is linear: seconds of audio times the rate. Round each clip up first. A 12.3 second voiceover bills as 13 seconds, so on fal Standard it costs 13 × $0.056, which is $0.73.

Clip lengthfal Standardfal ProPoYo StandardPoYo Pro
10 seconds$0.56$1.15$0.35$0.70
30 seconds$1.68$3.45$1.05$2.10
60 seconds$3.36$6.90$2.10$4.20

Volume is where the gap grows. A batch of 500 clips at 30 seconds each is 15,000 seconds of audio:

Provider and tier500 clips at 30 seconds
fal Standard$840
fal Pro$1,725
PoYo Standard$525
PoYo Pro$1,050

Pro earns its premium when a face stays on screen for a full minute, such as a paid ad. Standard is the sensible default for internal training clips, drafts, and A/B tests of scripts.

💡 Budget for retries. While you tune photos and prompts, some clips will need a second render. Add a buffer of 20 to 30 percent to your first batch estimate, then trim it once your settings are stable.

Request Format and Limits

Every provider wraps the same model in a slightly different envelope. Two examples show the shape.

A developer typing at a standing desk with two monitors

Inputs and File Limits

ParameterRequiredNotes
image_url / imageYesJPEG, PNG, or WebP up to 10 MB. APIPass also asks for at least 300 px and an aspect ratio between 1:2.5 and 2.5:1
audio_url / audioYesMP3, WAV, M4A, or AAC on APIPass, 5 MB maximum at PoYo and APIPass
promptNoDefaults to "." on fal. APIPass accepts up to 2,500 characters
modeAPIPass onlystd or pro

A Working fal Request

On fal, each tier is its own endpoint and the job runs through a queue. This JavaScript call uses the official client and waits for the result:

import { fal } from "@fal-ai/client";

const result = await fal.subscribe("fal-ai/kling-video/ai-avatar/v2/standard", {
  input: {
    image_url: "https://example.com/presenter.jpg",
    audio_url: "https://example.com/voiceover.mp3",
    prompt: "Calm presenter, relaxed shoulders, small nod after each sentence",
  },
  logs: true,
});

console.log(result.data.video.url, result.data.duration);

Swap the endpoint to fal-ai/kling-video/ai-avatar/v2/pro for the Pro tier. The response holds a video file with a download URL and a duration in seconds, a handy number to log against your invoice.

APIPass uses a single create-task route and a mode field instead:

{
  "model": "kling/avatar-v2",
  "callBackUrl": "https://your-domain.com/api/callback",
  "input": {
    "image": "https://example.com/avatar-image.jpg",
    "audio": "https://example.com/speech-audio.mp3",
    "prompt": "Professional spokesperson with confident gestures",
    "mode": "std"
  }
}

You POST that body to https://api.apipass.dev/api/v1/jobs/createTask with a bearer token in the Authorization header, and the reply returns a taskId. PoYo follows the same pattern: an asynchronous submit call, then either a webhook callback or status polling by task ID.

💡 Use webhooks in production. Avatar renders are queued jobs, not instant responses. A webhook keeps your workers free, while a polling loop burns requests while you wait.

Lip Sync Examples That Work

Three setups show where the per-second price pays back.

Spokesperson Videos

A woman in a charcoal blazer presenting a small amber bottle to camera

A skincare brand wants 40 product clips of 20 seconds each, one per ingredient. That is 800 seconds of audio. On fal Standard it costs $44.80, on PoYo Pro it costs $56, and on fal Pro it costs $92. Compare that with booking a presenter and a studio day for every variation. The same portrait is reused for every clip, so the face stays consistent across the whole series.

💡 Prompt to try: "Warm, confident presenter, steady eye contact, small head nods, relaxed shoulders."

Course Narration

A lecturer recording an online lesson in a home study

A 10 minute lesson is 600 seconds of narration. Split it into ten 60 second segments so each one fits the strictest audio limit I found, then reuse one photo for all of them. On PoYo Standard the lesson costs $21.00, and on fal Standard it costs $33.60. Cutting at sentence boundaries keeps the joins invisible, because the speaker is already at rest between sentences.

Multilingual Versions

An audio engineer adjusting a fader while a speaker records behind glass

One photo plus one script in three languages gives you three clips. At 30 seconds each on fal Standard, that is 90 seconds and $5.04. Generate each language's audio with a text-to-speech model such as MiniMax Speech 2.8 HD or ElevenLabs V3, then feed each file to the avatar model. If you already have a finished video, HeyGen Video Translate dubs it directly instead.

Stylized characters and animals also work, since the model accepts cartoons and non-human faces. Test a few renders before you build a whole series around a mascot.

Prepare Inputs for Better Results

Most bad renders trace back to the inputs, not the model. Fix these before you spend credits.

Photo Rules

A front-facing portrait of a man with short gray hair and a navy shirt

  • Face the camera. A straight-on head and shoulders frame gives the model the most facial geometry to animate.
  • Keep the mouth closed or neutral. A wide smile with visible teeth in the source can look odd once the lips start moving.
  • Show ears and hairline. Clear edges help the head keep a stable outline during motion.
  • Light the face evenly. Soft window light from the front beats a hard side shadow.
  • Start at 300 px or more on the short side. If your portrait is small, upscale it first with Crystal Upscaler, which is built for portraits.

Audio and Prompt Rules

A man with a silver beard narrating into a microphone in a foam-lined corner

  • One speaker per file. Overlapping voices give the mouth nothing clear to follow.
  • No music under the voice. Mix the music in after the avatar render.
  • Trim silence at both ends. You pay for every second of audio, including dead air.
  • Stay inside the provider limit. The shortest cap I found is 60 seconds, and the minimum is 2 seconds.
  • Keep the prompt short. One line on emotion, one on gesture, one on camera is enough. A long prompt competes with the audio instead of supporting it.

How to Use Kling Lip Sync

Kling Avatar 2.0 itself is not listed in the PicassoIA lipsync catalog at the time of writing. The Kling option that is there, Kling Lip Sync, solves a neighboring problem: it matches the mouth of the speaker in an existing video clip to a new audio track, or to a script you type. It accepts MP4 and MOV clips between 2 and 10 seconds at 720p to 1080p.

A woman at a bright workspace with a video timeline on her monitor

Step by Step

  1. Open the Kling Lip Sync page on PicassoIA.
  2. Upload your clip or paste its link in video_url. Use an MP4 or MOV under 100 MB, 2 to 10 seconds long.
  3. Add the voice. Either upload an MP3, WAV, M4A, or AAC file under 5 MB in audio_file, or type a script in text.
  4. If you typed a script, choose a voice_id from the English or Chinese list and set voice_speed.
  5. Run the generation, preview the result, and download the clip.

Parameter Tips

  • Use either audio or text for a given run. The text field is meant for when you have no audio file.
  • voice_id defaults to en_AOT. Listen to two or three voices on the same line before you pick one.
  • voice_speed defaults to 1. Nudge it until the speech rhythm matches the pace of the footage.
  • video_url and video_id cannot be combined, so pick one source per run.
  • Cut longer scenes into 10 second pieces before you upload them.

PicassoIA Alternatives for Talking Photos

When you start from a photo instead of a video, three PicassoIA models follow the same photo plus audio pattern as Kling Avatar 2.0.

ModelInputVoiceOutput notes
Omni Human 1.5Photo, audio, optional promptYour audio, under 35 secondsFaces, full-body shots, and illustrated characters. Fast mode and a seed for repeatable results
Fabric 1.0Photo and audioYour audio480p or 720p
P Video AvatarPhoto, plus a script or audio30+ built-in voices in 10 languages720p or 1080p
Kling Lip SyncVideo, plus audio or textYour audio or a built-in voice2 to 10 second clips

A quick way to choose:

💡 Test one clip before a batch. Run the same portrait and the same 10 seconds of audio through two models, then compare the mouth shapes on the hard sounds like "p", "b", and "m".

Try Your Own Lip Sync Video

Pick one clip and test it before you plan a batch: one portrait, ten seconds of clean audio, one tier. Run it through Omni Human 1.5 or Fabric 1.0 on Picasso IA, compare the mouth shapes against the audio, and keep the settings that win as your template. Then generate the portrait itself with the platform's image models, record or synthesize the voice, and animate it. Open Picasso IA, upload your first photo, and see how far a single still image and a short voice recording can go.

Share this article