Prompting#
A video prompt is a shot description, not a caption. The rules below come from hundreds of takes; each one decides whether a take is usable.
Describe the shot, in order#
- Framing and camera: "medium close-up, eye level, static camera".
- Subject and appearance: who is in frame and what they wear. Keep the description consistent with the reference image; do not fight it.
- Action and motion: one clear action per shot. "She turns to the window and smiles" works; three actions in ten seconds do not.
- Setting and light: place, time of day, light source.
- Sound: ambient sound, or the dialogue block below. For a silent shot, say so.
Specific beats vague. A prompt that names the lens, the light and the action holds the model to it; a loose prompt lets the seed decide.
Dialogue#
- Write directions in English and put the spoken line in quotes, in the
language it should be spoken:
The man says: "Bu akşam gelmiyorum." - Type the line with the proper characters of that language. A Turkish line
written in plain ASCII (
sasirdiminstead ofşaşırdım) is read out letter by letter. - One speaker per line. For two speakers, give each their own tagged line.
- Match the take length to the line: a 10-second take cannot carry a 25-word sentence. Leave a beat of silence at the end rather than cutting the last word.
- Numbers and years are spoken differently between takes. If the reading matters, write it out in words.
Shot cuts inside a take#
MiniMax H3 presets accept a second shot block inside one take (a [Shot 2]
line followed by its own framing and action). Use it for a reaction cut
without spending a second take. Keep both shots in the same setting.
Things that do not work#
- Two or more people described in detail without a reference for each: the model swaps features between them. Give each person a reference and keep it to two references per scene.
- "Change nothing else" phrasing: it freezes the camera and the motion.
- Long negative lists. Say what you want instead.