Skip to content

Reference images and voices#

Reference image#

Every shoot has a reference image. What it does depends on the preset family:

  • Image-to-video (LTX): the reference is the first frame. The video starts from it and moves according to the prompt.
  • Reference-to-video (MiniMax H3): the reference is a face and appearance anchor. The model builds the shot from the prompt and keeps the person looking like the reference.

The aspect-ratio rule

The aspect ratio of the reference must match the output aspect ratio of the preset. A mismatched image is cropped silently before generation; the result then ignores parts of the prompt and drifts from the reference. Crop the image in the Edit tab first.

Choosing a face reference. A sharp, front-lit, eye-level frame of the face at the size it should appear on screen. Avoid sunglasses, motion blur and crowded frames. Use at most two reference images per scene: each extra reference multiplies the render time and degrades the result.

Voice references#

Speech presets (MiniMax H3) can anchor a voice. Record or import a short clean clip of the speaker in the Voice references panel of the shoot and attach it to the shoot. Two voice anchors per shoot are fine and cost only a few percent of render time.

  • The clip should be the voice alone: no music, no room echo, a few seconds of natural speech.
  • The line to be spoken still goes into the prompt in quotes; the reference only shapes the timbre.
  • Without a voice reference the model picks a voice per take, so it changes between takes.

Identity across shots#

Reference images give a likeness lottery: roughly three takes out of four keep the face. For a character that appears in many shots, train a character LoRA and build a character set for the scene; the LoRA carries the identity even in shots without a reference.