Generate videosLipsync videosVisual Effects

InfiniteTalk ComfyUI Workflow and API: Free Lip Sync Avatars

InfiniteTalk is an open source model that turns one photo and an audio file into a talking video with synced lips, head, and body motion. This article shows the ComfyUI workflow, the settings that matter, hosted and self-hosted API calls, and a browser based alternative for quick results.

InfiniteTalk ComfyUI Workflow and API: Free Lip Sync Avatars
Cristian Da Conceicao
Founder of Picasso IA

You have a photo, a voice recording, and no camera crew. InfiniteTalk turns those two files into a video where the person in the picture actually speaks, with lips, head, and shoulders moving in time with the audio. The weights are open under the Apache 2.0 license, ComfyUI can run them on your own GPU, and a hosted API can run them when you do not have one. This article walks through both routes: the ComfyUI workflow step by step, the settings that change the result, the API calls, and a no-install alternative on PicassoIA for the days when a local setup is more work than the job deserves.

What InfiniteTalk Actually Does

InfiniteTalk comes from the MeiGen-AI team. Code, weights, and the technical report went public on August 19, 2025. It is built on the 14B, 480P checkpoint of the Wan 2.1 image to video model, and it reads audio through a wav2vec2 encoder. Both 480p and 720p output are supported.

The model has a lineage. The same group's earlier MultiTalk handled audio-driven conversation videos, and the ComfyUI wrapper still lists it next to InfiniteTalk. What InfiniteTalk adds is unlimited-length generation and whole-body motion on top of lip movement. The repository's news section now also points to a successor family called LongCat Video Avatar, so check it when you want the newest weights from the same team.

A young man watches a paused video frame of a woman mid-sentence on a matte monitor

Sparse-Frame Dubbing in Plain Words

Most older lip sync tools repaint a patch around the mouth and leave the rest of the face frozen. The result looks like a talking mask. InfiniteTalk uses what its authors call sparse-frame video dubbing: it keeps a handful of reference frames to hold identity steady, then generates everything in between so that lips, head turns, posture, and facial expressions all follow the audio. Long clips are built from chained context windows, with a default window of 81 frames, which is how a talking video can run far past the usual five seconds.

Image Mode vs Video Mode

ModeYou provideBest forLength
Image to videoOne photo plus an audio fileNarrators, avatars, singers, spokespeopleUp to about one minute before quality drops
Video to videoAn existing clip plus new audioDubbing, replacing a voice, keeping camera movementUnlimited in theory

💡 Start with image mode. It needs the fewest inputs and fails in the most readable ways, so you see what the audio settings do before a source video enters the mix.

Is It Really Free?

The license is free. The compute is not. That difference decides which route you should pick.

RouteWhat it costsWhat you trade
Local ComfyUIYour GPU and electricitySetup time, VRAM limits, slow renders on small cards
Hosted APIBilled per second of output. At the time of writing, WaveSpeed lists roughly $0.03 per second at 480p and $0.06 per second at 720p, up to 10 minutes per jobOngoing spend, less control over the graph
Browser model on PicassoIAFree to try on several lipsync modelsA different model, shorter clips

Free also has a time price. A first local install eats a chunk of an afternoon in node installs and downloads, and the model files alone run to tens of gigabytes. Judge the route by how many clips you plan to make, not by the sticker price.

The Hardware Reality Check

InfiniteTalk sits on a 14 billion parameter video model, so memory is the first wall you hit. The official repository ships a low VRAM mode (--num_persistent_param_in_dit 0) that keeps weights off the GPU between steps, plus an FP8 quantized checkpoint. In ComfyUI, Kijai's FP8 scaled files serve the same purpose. I will not quote a VRAM number, because it moves with resolution, frame count, and quantization. Test with a 480p, 10 second clip before you commit an evening to a 720p render.

Macro close-up of a graphics card seated in an open PC case

Build the ComfyUI Workflow

ComfyUI support comes through Kijai's ComfyUI-WanVideoWrapper. The InfiniteTalk music template on comfy.org lists three node packs: the WanVideoWrapper, audio-separation-nodes-comfyui, and comfyui-kjnodes. Together they load the model, split voice from music, and handle the image and video plumbing.

Install the Custom Nodes

  1. Open a terminal in ComfyUI/custom_nodes and clone the wrapper.
  2. Install its requirements with pip install -r requirements.txt (portable builds use the bundled python_embeded interpreter instead).
  3. Add the audio separation and kjnodes packs through ComfyUI Manager.
  4. Restart ComfyUI so the new nodes register.
cd ComfyUI/custom_nodes
git clone https://github.com/kijai/ComfyUI-WanVideoWrapper
pip install -r ComfyUI-WanVideoWrapper/requirements.txt

Top-down view of a desk with a laptop, an SSD, and a handwritten setup checklist

Download the Model Files

The wrapper's README sorts models into four folders: text_encoders, clip_vision, diffusion_models, and vae. InfiniteTalk adds its own weights on top of the Wan 2.1 base.

FileFolderNotes
Wan 2.1 image to video 14Bdiffusion_modelsBase model, FP8 scaled versions exist
UMT5 text encodertext_encodersSame one the Wan 2.1 workflows use
Wan 2.1 VAEvaeShared with other Wan workflows
CLIP vision modelclip_visionReads the reference photo
InfiniteTalk Single (fp16, about 5.1 GB)diffusion_modelsOne speaker
InfiniteTalk Multi (fp16, about 5.1 GB)diffusion_modelsTwo speakers

Both InfiniteTalk files live in the InfiniteTalk folder of Kijai's WanVideo_comfy repository on Hugging Face. If a loader dropdown stays empty after you copy a file in, refresh ComfyUI and check the folder again.

Wire the Graph

Node names shift between wrapper versions, so open the example workflow from the wrapper's example_workflows folder instead of building from scratch. The logic is always the same:

  • Image branch: load the portrait, resize it to the target resolution, and encode it with the CLIP vision model.
  • Audio branch: load the audio, split the voice from any music, and pass it through the wav2vec2 encoder to get audio embeddings.
  • Model branch: load the Wan 2.1 base, attach the InfiniteTalk weights as an extra model, and add a speed LoRA if you want one.
  • Sampler: feed image embeds, text embeds, and audio embeds into the sampler, then decode with the VAE.
  • Output: combine the frames and the original audio into an MP4.

💡 Resize before you sample. A portrait that does not match the target aspect ratio gets stretched or cropped in ways that look like a bad lip sync but are really a bad input.

Prepare Audio and Portrait

The graph is only half the result. The two input files decide the rest.

Audio That Syncs Cleanly

Voice only is the safest input. Background music under the voice muddies the mouth shapes, which is why the ComfyUI template includes audio separation nodes that pull vocals out of a mix. Trim dead air from the start, even out the loudness, and export as WAV when you can.

No recording? Generate the voice instead. Speech 2.8 HD and ElevenLabs v3 both turn a script into a clean voiceover you can feed straight into the audio branch.

Image mode loses quality after about a minute, so write for scenes, not monologues. Break a long script into 20 to 40 second sections, render each one from the same portrait and the same seed, and join them in any video editor. Cutting on a pause hides the seam, and the repeated portrait keeps the identity stable from scene to scene.

Hands turning a knob on an audio interface beside a condenser microphone

A Source Photo That Works

The portrait is the first frame of your avatar, so every flaw travels through the whole video.

  • Face the camera, with a relaxed or lightly closed mouth.
  • Show the shoulders, because InfiniteTalk animates posture too.
  • Use even light with no hard shadow across the lips.
  • Keep hands, microphones, and hair away from the mouth.

No suitable photo? Create one with Seedream 4.5 and ask for a neutral expression, front-facing, soft window light.

A woman in her fifties in a navy blazer facing the camera with a neutral expression

Settings That Change the Result

Two dials do most of the work: the number of sampling steps and the audio CFG scale. Everything else is secondary until those behave.

Steps and Audio CFG

The official defaults use 40 sampling steps. For audio CFG, the repository recommends a value between 3 and 5 for good lip sync. In practice, a higher value pushes the mouth to follow the audio more strictly, and pushing it too far makes the face stiff.

SettingStarting valueWhat it does
Sampling steps40Detail and stability, slower renders
Audio CFG3 to 5Strength of the audio to mouth link
Resolution480p first, 720p for finalsSharpness against render time and memory
SeedFixed while testingLets you change one setting at a time

Fast Mode With LoRAs

Forty steps on a 14B model is slow. The repository describes two shortcuts: 8 steps with the FusionX LoRA, or 4 steps with the lightx2v LoRA. When you use a speed LoRA, drop the audio CFG to about 2. You lose some fine texture in skin and hair, so use fast mode for drafts and for clips that will be viewed small, then rerun the best take at full steps.

When a render goes wrong, the symptom usually points at one cause:

SymptomLikely causeFix
Lips drift away from the audioAudio CFG too low, or music under the voiceRaise audio CFG toward 4 and separate the vocals
Stiff, mask-like faceAudio CFG too highLower it by one point and rerun with the same seed
Face slowly changes in a long clipImage mode past its limit of about a minuteSplit the script into scenes and join them in an editor
Out of memory errorResolution or frame window too largeDrop to 480p, use the FP8 files, or enable the low VRAM setting
Mouth opens during silenceBreath noise or room hum read as speechClean the audio with a noise gate before loading it

💡 Fix the seed, change one thing. Render the same five seconds at 480p while you adjust the audio CFG. Once the mouth looks right, raise the resolution.

Call InfiniteTalk Through an API

An API turns a ten minute manual process into a function call. You have two options, depending on who owns the GPU.

The Hosted API Route

WaveSpeed hosts InfiniteTalk at POST https://api.wavespeed.ai/api/v3/wavespeed-ai/infinitetalk. The request takes image, audio, resolution (480p or 720p), seed, an optional prompt, and an optional mask_image. You get back a prediction ID, then poll a result URL every couple of seconds until the status reaches a terminal state. Audio up to 10 minutes is accepted per job.

import os, time, requests

TOKEN = os.environ["WAVESPEED_TOKEN"]
HEADERS = {"Authorization": f"Bearer {TOKEN}"}
BASE = "https://api.wavespeed.ai/api/v3"

job = requests.post(
    f"{BASE}/wavespeed-ai/infinitetalk",
    headers=HEADERS,
    json={
        "image": "https://example.com/portrait.jpg",
        "audio": "https://example.com/voice.wav",
        "resolution": "480p",
    },
).json()["data"]

while True:
    result = requests.get(f"{BASE}/predictions/{job['id']}/result", headers=HEADERS).json()["data"]
    if result["status"] not in ("created", "processing"):
        break
    time.sleep(2)

print(result["status"], result.get("outputs"))

Treat the snippet as a starting point and check the model page for the current schema before you ship it. Expect roughly 10 to 30 seconds of waiting per second of video, depending on resolution and queue load.

Your Own ComfyUI Endpoint

ComfyUI already runs an HTTP server, so your workflow is an API the moment you export it. Use Save (API Format), change the image and audio filenames inside the JSON, then post it to /prompt.

import json, time, requests

COMFY = "http://127.0.0.1:8188"
workflow = json.load(open("infinitetalk_api.json"))

prompt_id = requests.post(f"{COMFY}/prompt", json={"prompt": workflow}).json()["prompt_id"]

while True:
    history = requests.get(f"{COMFY}/history/{prompt_id}").json()
    if prompt_id in history:
        break
    time.sleep(3)

print(history[prompt_id]["outputs"])

Files named in the workflow must already sit in ComfyUI's input folder, or be uploaded to the /upload/image endpoint first. Fetch the finished video from /view using the filename listed in the outputs.

Low-angle view of a row of server racks with a technician at the far end

💡 Never expose port 8188 to the open internet. ComfyUI has no login. Put it behind a VPN or a reverse proxy with authentication before anyone outside your network can reach it.

Where Avatars Actually Help

A talking avatar earns its place when recording a person on camera is slow, expensive, or impossible to repeat.

  • Lessons and training: one instructor photo and a script give you consistent video modules, and updating a sentence means rerendering instead of reshooting.
  • Product walkthroughs: a spokesperson avatar can narrate the same page in several languages by swapping the audio.
  • Podcast clips: the Multi model handles two speakers, which suits interview snippets for social media.
  • Music and character work: singers, mascots, and illustrated characters work as long as the face is clearly visible.

A teacher recording a lesson on a smartphone tripod beside a whiteboard

Two friends laughing at a podcast table with boom arm microphones

💡 Get permission. Animate your own face, a licensed character, or someone who agreed to it, and label the result as AI generated when it could be mistaken for a real recording.

Skip the Setup on PicassoIA

PicassoIA does not host InfiniteTalk itself. It does offer several photo plus audio models that deliver the same kind of result in a browser tab, with no nodes, no downloads, and no VRAM math.

Pick a Lipsync Model

ModelInputGood to know
Omni Human 1.5Photo plus audioAudio under 35 seconds, optional text prompt, fast mode
Fabric 1.0Photo plus audioChoice of 480p or 720p output
Wan 2.2 S2VPhoto, audio, and promptPrompt shapes motion and mood, 81 frames per chunk by default
Kling Avatar v2Photo plus audio (MP3, WAV, M4A, or AAC)Standard or Pro mode, handles cartoons and animals too
P Video AvatarPhoto plus a typed script or your own audio30+ built-in voices, up to 1080p
Lipsync 2 ProExisting video plus WAV audioActive speaker detection and five sync modes

The last row matters if you came for video to video dubbing: Lipsync 2 Pro works on footage you already shot. Which route fits depends on the job:

  • Choose local ComfyUI for clips longer than a minute, the Multi model with two speakers, dubbing that keeps the original camera moves, or full control over every sampler setting.
  • Choose the hosted API when your own app has to create avatars on demand and you would rather pay per second than keep a GPU awake.
  • Choose a PicassoIA model when you want a result in the next ten minutes, your audio runs under 35 seconds, and speed matters more than matching InfiniteTalk's exact look.

Steps for Your First Clip

  1. Open the Omni Human 1.5 page on PicassoIA.
  2. Upload a front-facing portrait. Faces, full-body shots, and illustrated characters all work.
  3. Upload an MP3 or WAV file under 35 seconds. English, Spanish, Japanese, Korean, Chinese, and Indonesian audio are supported.
  4. Add an optional prompt to direct camera movement or gestures, such as "she smiles and nods slightly".
  5. Switch on fast mode for quick drafts, or leave it off for the best detail.
  6. Set a seed if you want to reproduce a result, then generate and download the video.

💡 Write the script first, record it with Speech 2.8 HD, and keep each clip under the 35 second limit. Short clips stitched together look more natural than one long take.

Make Your First Avatar

Pick one portrait, record 20 seconds of audio, and run the same clip through two or three models. Comparing results side by side takes minutes and tells you more than any spec sheet. Open Omni Human 1.5 or Fabric 1.0 on PicassoIA, upload your two files, and watch a still photo start talking. If you want a local pipeline later, you will already know what a good input looks like.

A young woman smiling at her smartphone in a bright loft studio

Share this article