Limits and performance#
Envelope#
| LTX | MiniMax H3 | |
|---|---|---|
| Frames per take | up to about 257 | 17n + 5, up to 345 |
| Audio | up to about 10 s stays clean | full take, speech and ambience |
| Resolution sweet spot | preset native | 768 square |
| References | first frame | up to 2 face references plus voice references |
| VRAM | 16 GB | 16 GB with int8 packs |
| Host RAM | 32 GB | 32 GB minimum, 64 GB recommended |
Render time#
The first take after a start loads the pack (a few minutes). After that, a 768-square H3 take of 124 frames renders in a few minutes on an RTX 5070 Ti class card; a 311-frame take takes about 40 minutes. Each extra face reference multiplies the time; keep it to two.
When it is slower than that#
Another process holding GPU memory is the usual cause. Voice assistants, local LLM servers and speech tools can squat several gigabytes; Maystro then spills into shared memory and every step slows down by a factor of three or four. Close them before rendering. See Troubleshooting.
Cancelling#
Cancel from the take card or the queue. The runtime stops at the next safe point, which can take a minute on a long take. The take is marked cancelled and no file is written.
Out of memory#
If a take fails with an out-of-memory error: shorten the take, drop to the 768-square preset, remove the second reference image, and make sure nothing else is on the GPU. Long takes at higher resolutions are outside the 16 GB envelope.