Skip to content

Presets#

A preset bundles a model pack with resolution, frame count, sampler settings and the LoRAs to fuse. Pick it per shoot. Two families ship with the starter set.

LTX (image-to-video, audio and video)#

Fast, starts from your reference image, generates picture and sound together.

  • Best for: establishing shots, motion from a still, short clips with ambient sound.
  • Clip length: up to about 10 seconds. Longer clips lose the audio envelope and the sound falls apart before the picture does.
  • The upsample variants render at a lower rate and stretch the frames: the result plays in slow motion. Use them only when that is the intent.
  • The reference image is the first frame, so the aspect-ratio rule in Reference images and voices is strict.

MiniMax H3 (reference-to-video, speech)#

Builds the shot from the prompt with a face reference and, optionally, a voice reference. Lip-sync and dialogue are its strength.

  • Best for: dialogue scenes, characters that must look the same across shots, any shot with spoken lines.
  • Frame counts follow a 17n + 5 rule (for example 124, 141, 311, 345). The preset list only offers valid counts. Longer takes are supported up to 345 frames; very long takes take tens of minutes on a 16 GB card.
  • Resolution: the 768-square presets are the sweet spot for identity and speed. Larger frames cost more than they give.
  • Character LoRAs are fused at strength 0.5 by default. Strength 1.0 destroys motion and lighting; do not raise it.

Choosing#

You want Use
A still to come alive, ambient sound LTX
Someone to say a line, lips in sync MiniMax H3
The same person across a whole scene MiniMax H3 with a character set
A longer take MiniMax H3 (up to 345 frames)

Custom presets (your own LoRA, a different resolution) are built in Models and presets.