Presets#
A preset bundles a model pack with resolution, frame count, sampler settings and the LoRAs to fuse. Pick it per shoot. Two families ship with the starter set.
LTX (image-to-video, audio and video)#
Fast, starts from your reference image, generates picture and sound together.
- Best for: establishing shots, motion from a still, short clips with ambient sound.
- Clip length: up to about 10 seconds. Longer clips lose the audio envelope and the sound falls apart before the picture does.
- The upsample variants render at a lower rate and stretch the frames: the result plays in slow motion. Use them only when that is the intent.
- The reference image is the first frame, so the aspect-ratio rule in Reference images and voices is strict.
MiniMax H3 (reference-to-video, speech)#
Builds the shot from the prompt with a face reference and, optionally, a voice reference. Lip-sync and dialogue are its strength.
- Best for: dialogue scenes, characters that must look the same across shots, any shot with spoken lines.
- Frame counts follow a
17n + 5rule (for example 124, 141, 311, 345). The preset list only offers valid counts. Longer takes are supported up to 345 frames; very long takes take tens of minutes on a 16 GB card. - Resolution: the 768-square presets are the sweet spot for identity and speed. Larger frames cost more than they give.
- Character LoRAs are fused at strength 0.5 by default. Strength 1.0 destroys motion and lighting; do not raise it.
Choosing#
| You want | Use |
|---|---|
| A still to come alive, ambient sound | LTX |
| Someone to say a line, lips in sync | MiniMax H3 |
| The same person across a whole scene | MiniMax H3 with a character set |
| A longer take | MiniMax H3 (up to 345 frames) |
Custom presets (your own LoRA, a different resolution) are built in Models and presets.