Kling Avatar 2.0 API: Lip Sync Pricing and Examples
Kling Avatar 2.0 turns one photo and one audio file into a talking video. This article lists per-second API prices on fal, PoYo, and Picsart, shows a working request, runs the cost math on real clip lengths, and compares talking-photo models you can run on PicassoIA.
Kling Avatar 2.0 takes one portrait photo and one audio file and returns a talking video where the lips, eyes, and head motion follow the speech. The model is sold per second of audio, which makes Kling Avatar 2.0 API lip sync pricing easy to predict once you know which provider you call and which tier you pick. This article puts the real numbers side by side: $0.056 per second on fal, $0.035 on PoYo, 2 credits on Picsart, with Pro rates roughly double each of those. You also get a working request, three production-style examples, input rules that prevent bad renders, and a tutorial for the closest lip sync models on PicassoIA.
What Kling Avatar 2.0 Does
Kling Avatar 2.0 is an audio-driven avatar model from the Kling family. You hand it a still image and a speech recording, and it animates the face so the mouth shapes follow the syllables, the eyes hold contact with the lens, and the head and shoulders move the way a real speaker's would. No camera footage and no 3D rig is involved, and the identity of the person in the photo stays stable from the first frame to the last.
One Photo, One Audio File
The input contract is short. Every provider I checked asks for the same three things:
One image of a person, animal, cartoon, or stylized character. PoYo accepts JPEG, PNG, and WebP files up to 10 MB.
One audio file with the speech. PoYo accepts 2 to 60 seconds and up to 5 MB, while Hedra lists a range of 2 to 300 seconds, so check the limit of the provider you plan to call.
An optional text prompt that describes emotion, gesture, or camera behavior, for example "calm presenter, small nod at the end of each sentence."
The output is an MP4. Hedra lists 16:9, 9:16, and 1:1 as supported aspect ratios, which fits YouTube, Reels, and square feed posts.
Standard vs Pro Tiers
Both tiers take identical inputs. The difference is output quality and price.
💡 Resolution claims differ by provider. APIPass lists 720p for Standard and 1080p for Pro, while Hedra lists 720p native for Pro. Render one test clip and check the file before you commit a batch.
Kling Avatar 2.0 API Pricing Compared
Pricing is where providers differ most. Billing is per second of audio, so a longer voiceover costs more no matter how simple the photo is.
Per-Second Rates by Provider
Provider
Standard
Pro
Billing notes
fal
$0.056 per second
$0.115 per second
Rates as quoted by PoYo and Hedra
PoYo
$0.035 per second (7 credits)
$0.070 per second (14 credits)
Audio duration rounded up to the next second
Picsart API
2 credits per second
4 credits per second
Credit price depends on your Picsart plan
APIPass
Credits per second
Credits per second
The figures I found were inconsistent
PoYo's page puts its Standard rate about 38% under fal's and its Pro rate about 39% under, which matches the arithmetic: $0.035 against $0.056, and $0.070 against $0.115.
💡 Check before you budget. Two snapshots of the APIPass pricing page disagreed by roughly a factor of ten, so I left its rate out of the cost math below. Confirm any provider's price in its own dashboard before you queue a large batch.
Cost Examples at Real Volumes
Billing is linear: seconds of audio times the rate. Round each clip up first. A 12.3 second voiceover bills as 13 seconds, so on fal Standard it costs 13 × $0.056, which is $0.73.
Clip length
fal Standard
fal Pro
PoYo Standard
PoYo Pro
10 seconds
$0.56
$1.15
$0.35
$0.70
30 seconds
$1.68
$3.45
$1.05
$2.10
60 seconds
$3.36
$6.90
$2.10
$4.20
Volume is where the gap grows. A batch of 500 clips at 30 seconds each is 15,000 seconds of audio:
Provider and tier
500 clips at 30 seconds
fal Standard
$840
fal Pro
$1,725
PoYo Standard
$525
PoYo Pro
$1,050
Pro earns its premium when a face stays on screen for a full minute, such as a paid ad. Standard is the sensible default for internal training clips, drafts, and A/B tests of scripts.
💡 Budget for retries. While you tune photos and prompts, some clips will need a second render. Add a buffer of 20 to 30 percent to your first batch estimate, then trim it once your settings are stable.
Request Format and Limits
Every provider wraps the same model in a slightly different envelope. Two examples show the shape.
Inputs and File Limits
Parameter
Required
Notes
image_url / image
Yes
JPEG, PNG, or WebP up to 10 MB. APIPass also asks for at least 300 px and an aspect ratio between 1:2.5 and 2.5:1
audio_url / audio
Yes
MP3, WAV, M4A, or AAC on APIPass, 5 MB maximum at PoYo and APIPass
prompt
No
Defaults to "." on fal. APIPass accepts up to 2,500 characters
mode
APIPass only
std or pro
A Working fal Request
On fal, each tier is its own endpoint and the job runs through a queue. This JavaScript call uses the official client and waits for the result:
import { fal } from "@fal-ai/client";
const result = await fal.subscribe("fal-ai/kling-video/ai-avatar/v2/standard", {
input: {
image_url: "https://example.com/presenter.jpg",
audio_url: "https://example.com/voiceover.mp3",
prompt: "Calm presenter, relaxed shoulders, small nod after each sentence",
},
logs: true,
});
console.log(result.data.video.url, result.data.duration);
Swap the endpoint to fal-ai/kling-video/ai-avatar/v2/pro for the Pro tier. The response holds a video file with a download URL and a duration in seconds, a handy number to log against your invoice.
APIPass uses a single create-task route and a mode field instead:
You POST that body to https://api.apipass.dev/api/v1/jobs/createTask with a bearer token in the Authorization header, and the reply returns a taskId. PoYo follows the same pattern: an asynchronous submit call, then either a webhook callback or status polling by task ID.
💡 Use webhooks in production. Avatar renders are queued jobs, not instant responses. A webhook keeps your workers free, while a polling loop burns requests while you wait.
Lip Sync Examples That Work
Three setups show where the per-second price pays back.
Spokesperson Videos
A skincare brand wants 40 product clips of 20 seconds each, one per ingredient. That is 800 seconds of audio. On fal Standard it costs $44.80, on PoYo Pro it costs $56, and on fal Pro it costs $92. Compare that with booking a presenter and a studio day for every variation. The same portrait is reused for every clip, so the face stays consistent across the whole series.
💡 Prompt to try: "Warm, confident presenter, steady eye contact, small head nods, relaxed shoulders."
Course Narration
A 10 minute lesson is 600 seconds of narration. Split it into ten 60 second segments so each one fits the strictest audio limit I found, then reuse one photo for all of them. On PoYo Standard the lesson costs $21.00, and on fal Standard it costs $33.60. Cutting at sentence boundaries keeps the joins invisible, because the speaker is already at rest between sentences.
Multilingual Versions
One photo plus one script in three languages gives you three clips. At 30 seconds each on fal Standard, that is 90 seconds and $5.04. Generate each language's audio with a text-to-speech model such as MiniMax Speech 2.8 HD or ElevenLabs V3, then feed each file to the avatar model. If you already have a finished video, HeyGen Video Translate dubs it directly instead.
Stylized characters and animals also work, since the model accepts cartoons and non-human faces. Test a few renders before you build a whole series around a mascot.
Prepare Inputs for Better Results
Most bad renders trace back to the inputs, not the model. Fix these before you spend credits.
Photo Rules
Face the camera. A straight-on head and shoulders frame gives the model the most facial geometry to animate.
Keep the mouth closed or neutral. A wide smile with visible teeth in the source can look odd once the lips start moving.
Show ears and hairline. Clear edges help the head keep a stable outline during motion.
Light the face evenly. Soft window light from the front beats a hard side shadow.
Start at 300 px or more on the short side. If your portrait is small, upscale it first with Crystal Upscaler, which is built for portraits.
Audio and Prompt Rules
One speaker per file. Overlapping voices give the mouth nothing clear to follow.
No music under the voice. Mix the music in after the avatar render.
Trim silence at both ends. You pay for every second of audio, including dead air.
Stay inside the provider limit. The shortest cap I found is 60 seconds, and the minimum is 2 seconds.
Keep the prompt short. One line on emotion, one on gesture, one on camera is enough. A long prompt competes with the audio instead of supporting it.
How to Use Kling Lip Sync
Kling Avatar 2.0 itself is not listed in the PicassoIA lipsync catalog at the time of writing. The Kling option that is there, Kling Lip Sync, solves a neighboring problem: it matches the mouth of the speaker in an existing video clip to a new audio track, or to a script you type. It accepts MP4 and MOV clips between 2 and 10 seconds at 720p to 1080p.
You have a photo and only a script:P Video Avatar writes the voice for you.
You have footage and need new audio:Kling Lip Sync is the right tool.
You need a voice first:Realtime TTS 2 and ElevenLabs V3 produce audio you can feed to any of the models above.
💡 Test one clip before a batch. Run the same portrait and the same 10 seconds of audio through two models, then compare the mouth shapes on the hard sounds like "p", "b", and "m".
Try Your Own Lip Sync Video
Pick one clip and test it before you plan a batch: one portrait, ten seconds of clean audio, one tier. Run it through Omni Human 1.5 or Fabric 1.0 on Picasso IA, compare the mouth shapes against the audio, and keep the settings that win as your template. Then generate the portrait itself with the platform's image models, record or synthesize the voice, and animate it. Open Picasso IA, upload your first photo, and see how far a single still image and a short voice recording can go.