pablodawson/MiniMax-H3-360-Orbit-LoRA

Model

MiniMax-H3 360° Orbit LoRA (first–last frame)

95

2 commits

1 linked in READMEs

updated Sep 27, 2026

See the code

README

MiniMax-H3 360° Orbit LoRA (first–last frame)

A LoRA for MiniMax-H3 (FL2VA variant) that turns one photo into a frozen-time 360° camera orbit, and lands back on the exact frame it started from.

Four 360° orbits generated by the LoRA from a single photo each

Give the model the same image as the first and the last keyframe, and the LoRA orbits the camera all the way around the subject while the scene stays frozen. The clip closes on its own first frame, so orbits can be chained or stitched into longer shots without a visible seam.

Results

Each comparison shows three renders of the same input with the same seed, prompt, resolution and step count:

leftmiddleright
base model, first frame onlybase model, first + last framebase + this LoRA, first + last frame

skate: base first-only vs base first+last vs LoRA first+last

man2: base first-only vs base first+last vs LoRA first+last

girl2: base first-only vs base first+last vs LoRA first+last

man: base first-only vs base first+last vs LoRA first+last

The previews are downscaled animated WebP. Full-resolution MP4s: comparisons skate · man2 · girl2 · man; LoRA-only clips skate · man2 · girl2 · man.

About the comparisons

  • Base, first + last frame: when both keyframes are the same image, the base model reads the clip as a still and barely moves.
  • Base, first frame only: the camera moves, but drifts away and never returns to the starting view.
  • With the LoRA: the camera travels the full orbit and returns to the starting frame, with smooth motion and no cuts.

Why this LoRA?

Reference-to-video (Ref2VA) treats its input images as loose appearance references, so clips don't end on a known frame and can't be joined cleanly. First–last-frame generation (FL2VA) pins both ends, but out of the box it freezes when both ends are the same image. This LoRA teaches the model a real, geometry-consistent orbit, while FL2VA keyframe pinning at inference keeps both ends exact.

Prompt

Use this prompt verbatim.

One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
settingvalue
base weightsMiniMax-H3 FL2VA, pruned (Comfy-Org/MiniMax-H3 int8 ConvRot repack)
keyframesthe same image as first and last frame, for a full 360° loop
resolution768 × 768
frames73 (about 3 s at 24 fps)
steps28
guidancenone. MiniMax-H3 is guidance-distilled, so there is no CFG or negative prompt
LoRA strength1.0
audiooff

Usage

These weights were trained and tested with ostris/ai-toolkit and its minimax_h3 model extension. The keys use ComfyUI naming (diffusion_model.*).

Training details

Data

Hand-selected high-quality human Gaussian splats collected from the internet, with an orbit video rendered around each one. Because the splats are static 3D scenes, every frame is geometrically consistent by construction. The dataset shows exactly the motion the LoRA should learn: the camera moves and nothing else does.

clips28 orbit renders
clip format768 × 768, 73 frames, 24 fps
captionsthe single prompt above, for every clip; caption dropout 0.05
audionone

Hyperparameters

trainerostris/ai-toolkit, arch minimax_h3, partition fl2va_pruned
base modelMiniMax-H3 FL2VA pruned, int8 ConvRot (convrot8); Qwen3-VL-32B text encoder in nvfp4
training adapterostris/minimax_h3_training_adapter v1, active during training only and not needed for inference
conditioningfirst frame of each clip as the keyframe (i2v)
LoRArank 16, alpha 16, all transformer linear layers except adaln_proj (208 modules)
steps3000 (≈107 epochs)
batch size1, no gradient accumulation
optimizerAdamW 8-bit, lr 1e-4, weight decay 1e-4
precisionbf16, gradient checkpointing
noise scheduleflow matching, shifted timesteps
lossMSE, guidance loss target 3.5
resolution buckets512, 768
hardware1 × NVIDIA A100 80 GB, 10 h 24 min (about 12.5 s/step)

Limitations

  • The domain is narrow. The training data is 28 human-centric square clips of 73 frames. Other subjects, aspect ratios and clip lengths are untested.
  • You may still get some subtle movements, blinks, etc.

Author

Pablo Dawson

3d-consistency
ai-toolkit
camera-control
first-last-frame
flf2v
gaussian-splatting
image-text-to-video
image-to-video
lora
minimax-h3
orbit

pablodawson/MiniMax-H3-360-Orbit-LoRA

Model

MiniMax-H3 360° Orbit LoRA (first–last frame)

95

2 commits

1 linked in READMEs

updated Sep 27, 2026

See the code

README

MiniMax-H3 360° Orbit LoRA (first–last frame)

A LoRA for MiniMax-H3 (FL2VA variant) that turns one photo into a frozen-time 360° camera orbit, and lands back on the exact frame it started from.

Four 360° orbits generated by the LoRA from a single photo each

Give the model the same image as the first and the last keyframe, and the LoRA orbits the camera all the way around the subject while the scene stays frozen. The clip closes on its own first frame, so orbits can be chained or stitched into longer shots without a visible seam.

Results

Each comparison shows three renders of the same input with the same seed, prompt, resolution and step count:

leftmiddleright
base model, first frame onlybase model, first + last framebase + this LoRA, first + last frame

skate: base first-only vs base first+last vs LoRA first+last

man2: base first-only vs base first+last vs LoRA first+last

girl2: base first-only vs base first+last vs LoRA first+last

man: base first-only vs base first+last vs LoRA first+last

The previews are downscaled animated WebP. Full-resolution MP4s: comparisons skate · man2 · girl2 · man; LoRA-only clips skate · man2 · girl2 · man.

About the comparisons

  • Base, first + last frame: when both keyframes are the same image, the base model reads the clip as a still and barely moves.
  • Base, first frame only: the camera moves, but drifts away and never returns to the starting view.
  • With the LoRA: the camera travels the full orbit and returns to the starting frame, with smooth motion and no cuts.

Why this LoRA?

Reference-to-video (Ref2VA) treats its input images as loose appearance references, so clips don't end on a known frame and can't be joined cleanly. First–last-frame generation (FL2VA) pins both ends, but out of the box it freezes when both ends are the same image. This LoRA teaches the model a real, geometry-consistent orbit, while FL2VA keyframe pinning at inference keeps both ends exact.

Prompt

Use this prompt verbatim.

One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
settingvalue
base weightsMiniMax-H3 FL2VA, pruned (Comfy-Org/MiniMax-H3 int8 ConvRot repack)
keyframesthe same image as first and last frame, for a full 360° loop
resolution768 × 768
frames73 (about 3 s at 24 fps)
steps28
guidancenone. MiniMax-H3 is guidance-distilled, so there is no CFG or negative prompt
LoRA strength1.0
audiooff

Usage

These weights were trained and tested with ostris/ai-toolkit and its minimax_h3 model extension. The keys use ComfyUI naming (diffusion_model.*).

Training details

Data

Hand-selected high-quality human Gaussian splats collected from the internet, with an orbit video rendered around each one. Because the splats are static 3D scenes, every frame is geometrically consistent by construction. The dataset shows exactly the motion the LoRA should learn: the camera moves and nothing else does.

clips28 orbit renders
clip format768 × 768, 73 frames, 24 fps
captionsthe single prompt above, for every clip; caption dropout 0.05
audionone

Hyperparameters

trainerostris/ai-toolkit, arch minimax_h3, partition fl2va_pruned
base modelMiniMax-H3 FL2VA pruned, int8 ConvRot (convrot8); Qwen3-VL-32B text encoder in nvfp4
training adapterostris/minimax_h3_training_adapter v1, active during training only and not needed for inference
conditioningfirst frame of each clip as the keyframe (i2v)
LoRArank 16, alpha 16, all transformer linear layers except adaln_proj (208 modules)
steps3000 (≈107 epochs)
batch size1, no gradient accumulation
optimizerAdamW 8-bit, lr 1e-4, weight decay 1e-4
precisionbf16, gradient checkpointing
noise scheduleflow matching, shifted timesteps
lossMSE, guidance loss target 3.5
resolution buckets512, 768
hardware1 × NVIDIA A100 80 GB, 10 h 24 min (about 12.5 s/step)

Limitations

  • The domain is narrow. The training data is 28 human-centric square clips of 73 frames. Other subjects, aspect ratios and clip lengths are untested.
  • You may still get some subtle movements, blinks, etc.

Author

Pablo Dawson

3d-consistency
ai-toolkit
camera-control
first-last-frame
flf2v
gaussian-splatting
image-text-to-video
image-to-video
lora
minimax-h3
orbit