MiniMax-H3 360° Orbit LoRA (first–last frame)
95
2 commits
1 linked in READMEs
updated Sep 27, 2026
A LoRA for MiniMax-H3 (FL2VA variant) that turns one photo into a frozen-time 360° camera orbit, and lands back on the exact frame it started from.
Give the model the same image as the first and the last keyframe, and the LoRA orbits the camera all the way around the subject while the scene stays frozen. The clip closes on its own first frame, so orbits can be chained or stitched into longer shots without a visible seam.
Each comparison shows three renders of the same input with the same seed, prompt, resolution and step count:
| left | middle | right |
|---|---|---|
| base model, first frame only | base model, first + last frame | base + this LoRA, first + last frame |




The previews are downscaled animated WebP. Full-resolution MP4s: comparisons skate · man2 · girl2 · man; LoRA-only clips skate · man2 · girl2 · man.
Reference-to-video (Ref2VA) treats its input images as loose appearance references, so clips don't end on a known frame and can't be joined cleanly. First–last-frame generation (FL2VA) pins both ends, but out of the box it freezes when both ends are the same image. This LoRA teaches the model a real, geometry-consistent orbit, while FL2VA keyframe pinning at inference keeps both ends exact.
Use this prompt verbatim.
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
| setting | value |
|---|---|
| base weights | MiniMax-H3 FL2VA, pruned (Comfy-Org/MiniMax-H3 int8 ConvRot repack) |
| keyframes | the same image as first and last frame, for a full 360° loop |
| resolution | 768 × 768 |
| frames | 73 (about 3 s at 24 fps) |
| steps | 28 |
| guidance | none. MiniMax-H3 is guidance-distilled, so there is no CFG or negative prompt |
| LoRA strength | 1.0 |
| audio | off |
These weights were trained and tested with ostris/ai-toolkit and its minimax_h3 model extension. The keys use ComfyUI naming (diffusion_model.*).
Hand-selected high-quality human Gaussian splats collected from the internet, with an orbit video rendered around each one. Because the splats are static 3D scenes, every frame is geometrically consistent by construction. The dataset shows exactly the motion the LoRA should learn: the camera moves and nothing else does.
| clips | 28 orbit renders |
| clip format | 768 × 768, 73 frames, 24 fps |
| captions | the single prompt above, for every clip; caption dropout 0.05 |
| audio | none |
| trainer | ostris/ai-toolkit, arch minimax_h3, partition fl2va_pruned |
| base model | MiniMax-H3 FL2VA pruned, int8 ConvRot (convrot8); Qwen3-VL-32B text encoder in nvfp4 |
| training adapter | ostris/minimax_h3_training_adapter v1, active during training only and not needed for inference |
| conditioning | first frame of each clip as the keyframe (i2v) |
| LoRA | rank 16, alpha 16, all transformer linear layers except adaln_proj (208 modules) |
| steps | 3000 (≈107 epochs) |
| batch size | 1, no gradient accumulation |
| optimizer | AdamW 8-bit, lr 1e-4, weight decay 1e-4 |
| precision | bf16, gradient checkpointing |
| noise schedule | flow matching, shifted timesteps |
| loss | MSE, guidance loss target 3.5 |
| resolution buckets | 512, 768 |
| hardware | 1 × NVIDIA A100 80 GB, 10 h 24 min (about 12.5 s/step) |
Pablo Dawson
MiniMax-H3 360° Orbit LoRA (first–last frame)
95
2 commits
1 linked in READMEs
updated Sep 27, 2026
A LoRA for MiniMax-H3 (FL2VA variant) that turns one photo into a frozen-time 360° camera orbit, and lands back on the exact frame it started from.
Give the model the same image as the first and the last keyframe, and the LoRA orbits the camera all the way around the subject while the scene stays frozen. The clip closes on its own first frame, so orbits can be chained or stitched into longer shots without a visible seam.
Each comparison shows three renders of the same input with the same seed, prompt, resolution and step count:
| left | middle | right |
|---|---|---|
| base model, first frame only | base model, first + last frame | base + this LoRA, first + last frame |




The previews are downscaled animated WebP. Full-resolution MP4s: comparisons skate · man2 · girl2 · man; LoRA-only clips skate · man2 · girl2 · man.
Reference-to-video (Ref2VA) treats its input images as loose appearance references, so clips don't end on a known frame and can't be joined cleanly. First–last-frame generation (FL2VA) pins both ends, but out of the box it freezes when both ends are the same image. This LoRA teaches the model a real, geometry-consistent orbit, while FL2VA keyframe pinning at inference keeps both ends exact.
Use this prompt verbatim.
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
| setting | value |
|---|---|
| base weights | MiniMax-H3 FL2VA, pruned (Comfy-Org/MiniMax-H3 int8 ConvRot repack) |
| keyframes | the same image as first and last frame, for a full 360° loop |
| resolution | 768 × 768 |
| frames | 73 (about 3 s at 24 fps) |
| steps | 28 |
| guidance | none. MiniMax-H3 is guidance-distilled, so there is no CFG or negative prompt |
| LoRA strength | 1.0 |
| audio | off |
These weights were trained and tested with ostris/ai-toolkit and its minimax_h3 model extension. The keys use ComfyUI naming (diffusion_model.*).
Hand-selected high-quality human Gaussian splats collected from the internet, with an orbit video rendered around each one. Because the splats are static 3D scenes, every frame is geometrically consistent by construction. The dataset shows exactly the motion the LoRA should learn: the camera moves and nothing else does.
| clips | 28 orbit renders |
| clip format | 768 × 768, 73 frames, 24 fps |
| captions | the single prompt above, for every clip; caption dropout 0.05 |
| audio | none |
| trainer | ostris/ai-toolkit, arch minimax_h3, partition fl2va_pruned |
| base model | MiniMax-H3 FL2VA pruned, int8 ConvRot (convrot8); Qwen3-VL-32B text encoder in nvfp4 |
| training adapter | ostris/minimax_h3_training_adapter v1, active during training only and not needed for inference |
| conditioning | first frame of each clip as the keyframe (i2v) |
| LoRA | rank 16, alpha 16, all transformer linear layers except adaln_proj (208 modules) |
| steps | 3000 (≈107 epochs) |
| batch size | 1, no gradient accumulation |
| optimizer | AdamW 8-bit, lr 1e-4, weight decay 1e-4 |
| precision | bf16, gradient checkpointing |
| noise schedule | flow matching, shifted timesteps |
| loss | MSE, guidance loss target 3.5 |
| resolution buckets | 512, 768 |
| hardware | 1 × NVIDIA A100 80 GB, 10 h 24 min (about 12.5 s/step) |
Pablo Dawson