Generate videosGenerate imagesVisual Effects

LTX 2.5 vs MiniMax H3: Best Open Video Model?

Two open-weight video models arrived in August 2026: Lightricks LTX 2.5 and MiniMax H3. This side by side compares specs, 4K versus 2K output, clip length, audio, API prices, local hardware needs and license restrictions, so you can pick the right one for your projects.

LTX 2.5 vs MiniMax H3: Best Open Video Model?
Cristian Da Conceicao
Founder of Picasso IA

Two video models with downloadable weights landed within ten days of each other this summer, and creators have been arguing about them ever since. MiniMax announced H3 on July 31, 2026 as an API product, then released weights on August 3. Lightricks followed on August 11 with LTX 2.5. Both generate picture and sound in a single pass. Both promise cinematic output. The license fine print, though, makes the choice far less obvious than the spec sheets suggest.

This comparison sticks to what matters when you have a deadline: resolution, clip length, audio, speed, price, local hardware, and who is actually allowed to use the weights. I did not run a lab benchmark, so every figure below comes from published specifications and launch reports, and I flag the numbers that are vendor claims.

💡 Quick verdict: LTX 2.5 is the safer open choice for most people. MiniMax H3 offers 2K output and richer reference inputs through its API, but the weights you can actually download are more limited than the headline suggests.

Video editor comparing two paused clips on a dual-monitor desk in a bright loft studio

The Short Answer

Who Wins in One Line

If you want to run a video model on your own machine, build it into a pipeline, or fine-tune it, LTX 2.5 wins for most readers. Its weights are on Hugging Face, it worked in ComfyUI from day one, and its community license is free for organizations under $10 million in annual recurring revenue. If you want 2K output with dialogue and heavy reference control, and you are happy paying per second through an API, MiniMax H3 deserves a test.

The honest answer to "best open video model" depends on how much weight you put on the word open.

NeedBetter pickWhy
Self-hosting with few legal questionsLTX 2.5Open weights and a community license free under $10M ARR
Longest single clipLTX 2.5Up to 20 seconds on the Fast tier
Highest listed resolutionLTX 2.5Up to 4K, against 2K for H3
Reference-heavy directionMiniMax H3Up to 9 images, 3 video clips and 3 audio clips as input
Multilingual dialogueMiniMax H3Lip sync reported in 11 languages
Lowest cost per clipLTX 2.5$0.09 per second on the Fast tier

What "Open" Means Here

Open weights does not mean open source, and it does not mean free for everyone. It means the model files can be downloaded and run on your own hardware. The license then decides what you may do with them. Three questions sort it out:

  • Download rights: can you get the files at all?
  • Use rights: can you ship paid client work with them?
  • Region rights: are you even permitted to run them where you live?

Both vendors use custom community licenses with revenue thresholds, so read the text before you build a product on either one.

Technician's hands lifting a graphics card out of an open PC case on a steel workbench

LTX 2.5 at a Glance

Specs Worth Knowing

Lightricks built LTX 2.5 as a 22B-parameter audio-video model that generates picture and sound together. The headline feature is native multishot generation. Ask for a wide, a medium and a close-up, and the model renders them as one sequence that holds the character, lighting and voice steady from cut to cut instead of stitching separate clips together.

Other additions worth noting:

  • Automatic duration: the model predicts clip length from the action you describe.
  • 4K HDR output: high-resolution, high-dynamic-range video for finishing work.
  • RAW workflow: a path built for professional colour pipelines.
  • Editing beta: a mode for revising existing footage instead of regenerating it.

On speed, Lightricks claims a 10 second clip in 6.8 seconds on two NVIDIA GB200 superchips. That is a vendor number on very expensive hardware, so treat it as a ceiling, not a promise for your desk. Through the managed API, the same job reportedly took 23.7 seconds at 1080p, which is still quick.

License and Ecosystem

The weights sit on Hugging Face in several flavours: the full dev checkpoint, a distilled version for speed, and int8 and NVFP4 quantizations for smaller memory budgets. ComfyUI support was ready on release day, with official workflow templates for text-to-video and image-to-video.

The LTX-2.x Community License allows fine-tuning and self-hosting and costs nothing for organizations under $10 million in annual recurring revenue. It bars you from building a competing AI system with the model, and it prohibits military use without a separate agreement.

Cinematographer adjusting a cinema camera on a tripod at a coastal cliff during golden hour

MiniMax H3 at a Glance

What H3 Brings

H3 is MiniMax's omni-modal video model, the successor to the Hailuo 2.3 line. It generates 4 to 15 seconds of video at 24 fps, up to 2K, with native stereo sound (dialogue, effects and room tone) produced in the same pass. Launch reports also list lip-synced dialogue in 11 languages, and the base model is reported at 33.1B parameters.

Where H3 stands out is input flexibility. The API reportedly accepts up to 9 reference images, 3 reference video clips (15 seconds in total) and 3 audio clips, 12 files in all, plus a prompt of up to 7,000 characters. Audio references cannot be sent alone. For character consistency across shots, that is a lot of directorial control.

Launch-week API pricing was reported at $0.13 per second of 2K video, or $7.80 per minute, with the first five reference images free and $0.04 for each one after that. Check MiniMax's current rate card before you budget.

Team of colleagues reviewing paused footage on a large wall display in a glass meeting room

The License Catch

H3 started life as an API-only product. The weights arrived on August 3 under the MiniMax H3 Community License, and three details matter:

  • Only H3-Base was released. The two checkpoints, FL2VA and Ref2VA, are both CFG-distilled.
  • The 2K stage stayed closed. Local generation works at a 768 pixel short edge. The H3-Regenerate-2K module was not open-sourced, and the API runs the full 2K pipeline.
  • Territory limits. Several license write-ups, including hosting-provider breakdowns, say the license excludes the United States, the European Union, the United Kingdom and South Korea, reportedly even for using the model's outputs inside those regions. A revenue threshold of $20 million has also been reported.

💡 Read the license text yourself before relying on any third-party summary, this one included. If the territory exclusions apply to where you live, the open weights are off the table, however good the spec sheet looks.

Head to Head Comparison

Resolution and Clip Length

FeatureLTX 2.5MiniMax H3
ReleaseAugust 11, 2026API July 31, weights August 3, 2026
Model size22B parameters33.1B parameters
Open weightsFull dev, distilled, int8, NVFP4H3-Base only (FL2VA, Ref2VA)
Top resolutionUp to 4K HDR2K via API, 768p short edge locally
Clip lengthUp to 20 seconds (Fast tier)4 to 15 seconds
AudioJoint audio and videoNative stereo, dialogue in 11 languages
MultishotNative multishot sequencesMultiple shots inside one generation
Reference inputsImage to video, first and last frame9 images, 3 videos, 3 audio clips
LicenseLTX-2.x Community, free under $10M ARRMiniMax H3 Community, region limits reported

On paper, LTX 2.5 has more headroom: up to 4K and up to 20 seconds. Long clips come with a trade-off, since the 20 second ceiling applies at 1080p or below. H3 tops out at 2K and 15 seconds through the API and at 768p locally. Resolution is not quality, though. A clean 1080p clip with believable motion beats a 4K clip with melting hands.

Faces, Motion and Audio

Close-ups expose every weakness in a video model: skin texture, blinking, hair strands, lip shapes. H3 markets reference-based consistency and multilingual lip sync, which suggests MiniMax is aiming at dialogue-driven scenes. LTX 2.5 goes after the same target with joint audio-video generation and multishot sequences that keep a voice steady across cuts.

I cannot honestly rank their faces from spec sheets alone, and neither can anyone else without side by side renders. Test with your own material: one face, one line of dialogue, one camera move.

Close-up portrait showing skin texture and fine hair detail by a window

Sound is the other half of the picture:

  • LTX 2.5: audio is generated jointly with the video, and on LTX 2.5 Fast it can be switched on or off per render.
  • H3: stereo output at 32 kHz, according to one spec listing, with dialogue, effects and room tone in the same pass.

Sound engineer in headphones at a mixing desk in a treated recording room

Speed and Cost

Per-second pricing makes the math easy. These are launch-week API prices as reported, and the resolutions differ (H3 at 2K, LTX at up to 1080p), so this is not an apples to apples test.

Clip lengthLTX 2.5 Fast ($0.09/s)LTX 2.5 Pro ($0.12/s)H3 at 2K ($0.13/s)
5 seconds$0.45$0.60$0.65
10 seconds$0.90$1.20$1.30
15 seconds$1.35Not available (10 s cap)$1.95
20 seconds$1.80Not available (10 s cap)Not available (15 s cap)

The Fast tier is the cheapest way to iterate. The Pro tier is positioned around prompt adherence, faces and typography, so it makes sense for the final render once the Fast drafts look right.

Overhead view of a desk with a notebook, stopwatch, calculator and coffee cup

Running Them on Your Own Hardware

LTX 2.5 on a Workstation

Lightricks ships a distilled checkpoint for speed and quantized builds for tighter memory budgets. NVFP4 builds target NVIDIA's newest architecture, so older cards will likely do better with int8. I could not find a firm minimum VRAM figure in the launch reports, so check the model card for your exact GPU.

A sensible first run: load the distilled checkpoint in ComfyUI, render a 10 second clip at 1080p, and only then push toward 4K or longer durations.

Workstation tower with the side panel removed on a garage workbench

H3 Local Limits

The H3 repository is large. One hosting-provider walkthrough describes the full download, with every quantization included, at roughly 600 GB. The base model is 33.1B dense parameters against 22B for LTX 2.5, so expect a heavier load on your GPU. ComfyUI added four H3 nodes: EmptyMiniMaxH3LatentAV, MiniMaxH3ImageToVideo, MiniMaxH3ReferenceToVideo and MiniMaxH3SigmaShift.

What you get locally is the 768p base result. The 2K polish that fills the marketing pages comes from the API.

How to Use LTX 2.5 on PicassoIA

You do not need a GPU to try this model. LTX 2.5 Fast runs on Picasso IA and takes a text prompt or a single photo. At the time of writing, H3 itself is not in the catalogue. The newest MiniMax models there are Hailuo 2.3 and Hailuo 2.3 Fast, which makes a fair side test for the MiniMax look.

Step by Step

  1. Open the LTX 2.5 Fast page in the text-to-video collection.
  2. Pick your input. Write a prompt alone, or upload a first frame photo for image-to-video.
  3. Optional: add a last frame image and the model will interpolate a transition between the two.
  4. Set the aspect ratio, resolution, duration and frame rate (see the table below).
  5. Leave Generate Audio on if you want synced sound.
  6. Generate, preview the clip, then download it or adjust the prompt and run again.

Settings and Prompt Tips

SettingOptionsNotes
Aspect ratio16:9, 9:16Landscape for web, portrait for social
Resolution720p, 1080p, 2K, 4KDefault is 1080p; higher resolutions cap duration
Duration2 to 20 secondsOver 10 seconds only at 720p or 1080p with 24 or 25 fps
Frame rate24, 25, 48, 5048 and 50 fps cap at 10 seconds
Generate audioOn or offOn by default

For first frames, build the still with an image model such as Seedream 5 Pro, then feed it to the video model. You control composition before any motion is added.

Write prompts in the order things happen: subject, action, camera move, light, then sound. For example:

A fisherman hauls a rope on a small wooden boat at dawn. Slow dolly-in from the stern, sea mist drifting across the water, low amber light on wet planks. Creaking wood, gentle waves and distant gulls.

Screenwriter typing prompts on a laptop at a cafe table by a rain-streaked window

Which One Should You Pick?

Pick LTX 2.5 When

  • You want to self-host or fine-tune without region questions.
  • Clips longer than 15 seconds matter to you.
  • You iterate a lot and need the lowest cost per draft.
  • You work in ComfyUI and want day-one workflow templates.
  • You need 4K HDR for a finishing pipeline.

Pick H3 When

  • You direct with many references: faces, footage and voice clips in one request.
  • Your scenes depend on multilingual dialogue.
  • You are fine using the API and 2K output matters more than local control.
  • You have confirmed the license allows your location and your revenue band.

If neither fits, run the same prompt through Kling v3 Video or Veo 3.1 for a third opinion.

Run Your Own Comparison

Specs only go so far. The fastest way to settle this is to render the same shot twice and watch it with the sound on. Open LTX 2.5 Fast on Picasso IA, paste one prompt, and render it at 1080p and again at 4K. Then run the same description through Hailuo 2.3 and compare faces, motion and audio side by side.

Want a custom first frame? Generate it with Picasso IA's image models, upload it as the starting image, and let the video model take over. A few experiments will tell you more than any spec table, including this one.

Share this article