Datasets#
A dataset is a set of images, each with a caption, plus a trigger word that names the subject. Good datasets are small and clean: for a character, 10 to 25 sharp frames with a clearly visible face beat a hundred mixed ones.

From a character#
The fastest path. On a character card, Train character LoRA creates the project and a first dataset from the base reference and the angles. In the dataset dialog you can:
- pick which angles to include and whether to add the base image;
- add extra frames from the gallery (well-lit stills of the same person); photos you uploaded to the Assets tab are listed next to generated images, and the board menu counts both;
- choose the crop mode: auto centres a square crop on the detected face at the training resolution; none takes a centre crop;
- set the resolution (768 square is the default for video LoRAs), the trigger word, the description and the caption style.
The build report lists every frame with its crop and warnings: no face found, several faces, face too small. Fix or drop flagged frames before training.
From gallery images#
Create a dataset in the project and add images from the gallery. The picker shows generated images and uploaded photos alike, with a board filter; the default view spans every board. The same crop and caption path applies. Use this for styles, objects and for image LoRAs.
Captions and the trigger word#
Captions are generated from a template: a still-frame note, the trigger word,
the description, and a style line. Edit any caption in the sample grid. The
trigger word (chr_<name> for characters) appears once per caption; it is what
you will write in prompts to call the subject.
Keep captions honest. Describe what is in the frame, not what you wish the model would learn; the LoRA learns the difference between caption and image.
Caption detail. By default a character caption carries only the class
phrase of the description (Short): A still frame of chr_rbv2, a woman,
in a close-up portrait…. A full physical description in every caption lets
the model explain the frames with words instead of binding them to the trigger
word, and that identity then does not carry without a reference image. Choose
Full only when the words are the point (a style, an object).
Style tail. If the photos share a look (phone selfies, one studio light), name it in the Style field, for example in a casual phone snapshot. It is appended to every caption, so the look binds to that phrase instead of the trigger; leave the phrase out of your prompts and the character renders in the scene's own look.