pablodawson/MiniMax-H3-360-Orbit-LoRA is a
rank-16 ai-toolkit adapter for MiniMax-H3's first–last-frame (fl2va) partition that turns one photo
into a frozen-time 360° camera orbit: the camera travels all the way around the subject, nothing in the
scene moves, and the clip closes on the exact frame it started from — so orbits chain or stitch into longer
shots with no visible seam.
Upload a photo and press Orbit 360°. The photo is fed to MiniMax-H3 as both the first and the last keyframe — the wiring the LoRA was trained with, and what pins both ends. Every knob the card recommends is preset: 768×768, 73 frames (≈ 3 s at 24 fps), 28 steps, no CFG (MiniMax-H3 is guidance-distilled, so there is no negative prompt either), adapter strength 1.0, audio off, and the frozen-instant prompt used verbatim as the training captions had it:
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so this Space is the
denoising half only: the AdaLN-pruned transformer partition (40.24 GB) plus the two autoencoders,
loaded through MiniMaxH3GeneratorBlocks — MiniMaxH3Blocks with its text_encoder step removed. The
62.14 GiB Qwen3-VL conditioner runs in
multimodalart/qwen3vl-conditioner, which
this Space calls over the gradio API on every request; the wire format is one safetensors file holding
prompt_embeds + text_token_tags with the resolved height / width / num_frames in its metadata.
The pruned base is not a compromise here — it is the same checkpoint family the LoRA was trained against
(Comfy-Org's minimax_h3_fl2va_pruned_int8_convrot repack is an int8 quantization of these very weights),
and the adapter excludes every adaln_proj (network_kwargs.ignore_if_contains), the only place pruned and
released weights differ.
The LoRA ships in ComfyUI naming (diffusion_model.blocks.N.attn.qkv_proj, mlp.fc1/fc2, attn.out_proj,
plus the same four under token_refiner.blocks.N — 208 modules, 416 tensors), so app.py pushes it through
the same transforms the diffusers conversion script applied to the base weights:
blocks. → transformer_blocks., token_refiner.blocks. → token_refiner.refiner_blocks.,
attn.out_proj → attn.to_out.0, mlp.fc2 → ff.net.2 (pure renames),mlp.fc1 → ff.net.0.proj with its fused halves swapped — diffusers' SwiGLU reads [value; gate]
where the checkpoint stores [gate; value],attn.qkv_proj → to_q / to_k / to_v by splitting lora_B's 21504 rows into contiguous thirds
of 7168. ai-toolkit exports under diffusion_model. carry no peft .default. infix and are already
[q_all; k_all; v_all] — they must not be de-interleaved the way DiffSynth-Studio exports are.Every target is validated against the transformer's own parameter shapes and an unresolved one is fatal, so
the Space can never serve a half-applied adapter. It is attached as a runtime PEFT adapter
(load_lora_adapter + set_adapters) rather than folded, which is what keeps LoRA strength a live
per-request slider; no AoTI is used for the same reason — a compiled block package would bind the base
module's weights and run straight past the PEFT branch.
Fixed by the checkpoint: 24 fps, num_frames snapped to 17 * n + 5, no CFG and no negative prompt. The
model card places the tested range at 768×768, 73 frames — the whole training set is 28 human-centric square
clips — so those are the defaults, and the canvas / duration sliders are offered for other untested settings.
The audio latents are still generated and decoded (the pipeline is joint video+audio by construction), but per the card's "audio off" recommendation the soundtrack is not muxed into the output MP4.
| Variable | Default | Meaning |
|---|---|---|
H3_MODEL_REPO | multimodalart/MiniMax-H3-Pruned | The pruned diffusers-layout checkpoint. Public. |
H3_LORA_REPO | pablodawson/MiniMax-H3-360-Orbit-LoRA | The orbit adapter. Public. |
H3_LORA_FILE | minimax_h3_flf2v_lora_v1.safetensors | The adapter weights file. |
H3_LORA_SCALE | 1.0 | Default adapter strength (the card's recommendation). |
H3_CONDITIONER | multimodalart/qwen3vl-conditioner | The Space asked for embeddings. |
H3_PLACEMENT | lazy | lazy moves the partition onto the card on the first GPU call and leaves it there. |
H3_ATTENTION | _native_cudnn | cuDNN's fused kernel, 10–20% faster than the SDPA default. |
H3_GPU_SIZE | xlarge | ZeroGPU allocation size. large is what the sampler needs here. |
Two cards are booked per request — this Space's denoise loop and the conditioner's forward — and both are
billed to the requesting user: gradio_client attaches the caller's own x-ip-token to the conditioner
call, read off gradio's LocalContext inside the event listener, and ZeroGPU charges each booking to
whatever that token identifies. An unattributed caller (a script rather than a browser) leaves the
conditioner's booking on this Space's pod IP and its small shared quota.
Every request carries a photo of a person, so prompts pass the
hfmlsoc/ncii-light-guard-v01 NCII classifier before
the conditioner is called or a card is booked. The classifier runs in ncii_guard's spawned subprocess — in
the main process its torch activity poisons the CUDA forks and kills every @spaces.GPU worker. The prompt
here is fixed verbatim to the training caption, so the guard is a backstop rather than a filter users hit.
The three example photos ship with the Space under examples/, pre-cropped to the LoRA's own 768×768
square. They come from
linoyts/repo-to-space-example-inputs,
released CC0-1.0 (royalty-free sources, redistribution and modification permitted without attribution).
The live Space was verified twice through its gradio API with the bundled example photos. Both runs are the same default request and both landed correct output:
768×768 · 73 frames (3.042 s) · 28 steps · strength 1 · seed 904231 run 1: conditioner 1s (1276 tokens) · denoise + decode 107s (3.8 s/step) run 2: conditioner 13s · denoise + decode 139s (5.0 s/step)
The spread between the two (107 s → 139 s of GPU window) is the cold-start PIPE.to("cuda") placement
riding inside the window — run 2 was a fresh worker's very first call. The duration estimator carries a
45 s placement allowance for it and reserves 154 s for the default request, which covers the slower cold
run with ~10% margin; a warm worker finishes well inside.
The returned MP4 was checked frame-by-frame: 768×768, 73 frames at 24 fps, no audio track, and the loop genuinely closes — mean absolute pixel difference of 2.48 between the first and last frames against 82.95 between the first and a mid-orbit frame. That is the LoRA's signature behavior, not just a video that came back.
MiniMax-H3 is modular-only and not in a released diffusers, so requirements.txt installs it from the
canonical pull request, huggingface/diffusers#14371,
pinned to the commit 665f5782 (refs/pull/14371/head) rather than to the moving minimax-h3-refactor
branch. That PR is a WIP, so it needs re-pinning whenever it updates, and h3_split_blocks.py — which
subclasses its block classes to cut the pipeline in two — has to be re-checked against the new head at the
same time. The conditioner serves the same commit; the two halves must agree on the wire format.
The base model and this LoRA are under the MiniMax H3 Community License Agreement — community, non-commercial terms, not licensed in the US / EU / UK / Republic of Korea. The example photos are CC0-1.0.
pablodawson/MiniMax-H3-360-Orbit-LoRA is a
rank-16 ai-toolkit adapter for MiniMax-H3's first–last-frame (fl2va) partition that turns one photo
into a frozen-time 360° camera orbit: the camera travels all the way around the subject, nothing in the
scene moves, and the clip closes on the exact frame it started from — so orbits chain or stitch into longer
shots with no visible seam.
Upload a photo and press Orbit 360°. The photo is fed to MiniMax-H3 as both the first and the last keyframe — the wiring the LoRA was trained with, and what pins both ends. Every knob the card recommends is preset: 768×768, 73 frames (≈ 3 s at 24 fps), 28 steps, no CFG (MiniMax-H3 is guidance-distilled, so there is no negative prompt either), adapter strength 1.0, audio off, and the frozen-instant prompt used verbatim as the training captions had it:
One frozen instant. Only the camera moves. In a continuous 360 orbit. Preserve every person and object in exactly the same world position, orientation, shape and pose throughout the shot. Airborne objects remain suspended at the captured height and angle: no wobbling, shaking, spinning, drifting, falling or continued action. Keep faces, hands, clothing, liquids and the background motionless while retaining their natural appearance. Camera parallax is the only source of apparent movement. No cuts, zoom, morphing or added objects.
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so this Space is the
denoising half only: the AdaLN-pruned transformer partition (40.24 GB) plus the two autoencoders,
loaded through MiniMaxH3GeneratorBlocks — MiniMaxH3Blocks with its text_encoder step removed. The
62.14 GiB Qwen3-VL conditioner runs in
multimodalart/qwen3vl-conditioner, which
this Space calls over the gradio API on every request; the wire format is one safetensors file holding
prompt_embeds + text_token_tags with the resolved height / width / num_frames in its metadata.
The pruned base is not a compromise here — it is the same checkpoint family the LoRA was trained against
(Comfy-Org's minimax_h3_fl2va_pruned_int8_convrot repack is an int8 quantization of these very weights),
and the adapter excludes every adaln_proj (network_kwargs.ignore_if_contains), the only place pruned and
released weights differ.
The LoRA ships in ComfyUI naming (diffusion_model.blocks.N.attn.qkv_proj, mlp.fc1/fc2, attn.out_proj,
plus the same four under token_refiner.blocks.N — 208 modules, 416 tensors), so app.py pushes it through
the same transforms the diffusers conversion script applied to the base weights:
blocks. → transformer_blocks., token_refiner.blocks. → token_refiner.refiner_blocks.,
attn.out_proj → attn.to_out.0, mlp.fc2 → ff.net.2 (pure renames),mlp.fc1 → ff.net.0.proj with its fused halves swapped — diffusers' SwiGLU reads [value; gate]
where the checkpoint stores [gate; value],attn.qkv_proj → to_q / to_k / to_v by splitting lora_B's 21504 rows into contiguous thirds
of 7168. ai-toolkit exports under diffusion_model. carry no peft .default. infix and are already
[q_all; k_all; v_all] — they must not be de-interleaved the way DiffSynth-Studio exports are.Every target is validated against the transformer's own parameter shapes and an unresolved one is fatal, so
the Space can never serve a half-applied adapter. It is attached as a runtime PEFT adapter
(load_lora_adapter + set_adapters) rather than folded, which is what keeps LoRA strength a live
per-request slider; no AoTI is used for the same reason — a compiled block package would bind the base
module's weights and run straight past the PEFT branch.
Fixed by the checkpoint: 24 fps, num_frames snapped to 17 * n + 5, no CFG and no negative prompt. The
model card places the tested range at 768×768, 73 frames — the whole training set is 28 human-centric square
clips — so those are the defaults, and the canvas / duration sliders are offered for other untested settings.
The audio latents are still generated and decoded (the pipeline is joint video+audio by construction), but per the card's "audio off" recommendation the soundtrack is not muxed into the output MP4.
| Variable | Default | Meaning |
|---|---|---|
H3_MODEL_REPO | multimodalart/MiniMax-H3-Pruned | The pruned diffusers-layout checkpoint. Public. |
H3_LORA_REPO | pablodawson/MiniMax-H3-360-Orbit-LoRA | The orbit adapter. Public. |
H3_LORA_FILE | minimax_h3_flf2v_lora_v1.safetensors | The adapter weights file. |
H3_LORA_SCALE | 1.0 | Default adapter strength (the card's recommendation). |
H3_CONDITIONER | multimodalart/qwen3vl-conditioner | The Space asked for embeddings. |
H3_PLACEMENT | lazy | lazy moves the partition onto the card on the first GPU call and leaves it there. |
H3_ATTENTION | _native_cudnn | cuDNN's fused kernel, 10–20% faster than the SDPA default. |
H3_GPU_SIZE | xlarge | ZeroGPU allocation size. large is what the sampler needs here. |
Two cards are booked per request — this Space's denoise loop and the conditioner's forward — and both are
billed to the requesting user: gradio_client attaches the caller's own x-ip-token to the conditioner
call, read off gradio's LocalContext inside the event listener, and ZeroGPU charges each booking to
whatever that token identifies. An unattributed caller (a script rather than a browser) leaves the
conditioner's booking on this Space's pod IP and its small shared quota.
Every request carries a photo of a person, so prompts pass the
hfmlsoc/ncii-light-guard-v01 NCII classifier before
the conditioner is called or a card is booked. The classifier runs in ncii_guard's spawned subprocess — in
the main process its torch activity poisons the CUDA forks and kills every @spaces.GPU worker. The prompt
here is fixed verbatim to the training caption, so the guard is a backstop rather than a filter users hit.
The three example photos ship with the Space under examples/, pre-cropped to the LoRA's own 768×768
square. They come from
linoyts/repo-to-space-example-inputs,
released CC0-1.0 (royalty-free sources, redistribution and modification permitted without attribution).
The live Space was verified twice through its gradio API with the bundled example photos. Both runs are the same default request and both landed correct output:
768×768 · 73 frames (3.042 s) · 28 steps · strength 1 · seed 904231 run 1: conditioner 1s (1276 tokens) · denoise + decode 107s (3.8 s/step) run 2: conditioner 13s · denoise + decode 139s (5.0 s/step)
The spread between the two (107 s → 139 s of GPU window) is the cold-start PIPE.to("cuda") placement
riding inside the window — run 2 was a fresh worker's very first call. The duration estimator carries a
45 s placement allowance for it and reserves 154 s for the default request, which covers the slower cold
run with ~10% margin; a warm worker finishes well inside.
The returned MP4 was checked frame-by-frame: 768×768, 73 frames at 24 fps, no audio track, and the loop genuinely closes — mean absolute pixel difference of 2.48 between the first and last frames against 82.95 between the first and a mid-orbit frame. That is the LoRA's signature behavior, not just a video that came back.
MiniMax-H3 is modular-only and not in a released diffusers, so requirements.txt installs it from the
canonical pull request, huggingface/diffusers#14371,
pinned to the commit 665f5782 (refs/pull/14371/head) rather than to the moving minimax-h3-refactor
branch. That PR is a WIP, so it needs re-pinning whenever it updates, and h3_split_blocks.py — which
subclasses its block classes to cut the pipeline in two — has to be re-checked against the new head at the
same time. The conditioner serves the same commit; the two halves must agree on the wire format.
The base model and this LoRA are under the MiniMax H3 Community License Agreement — community, non-commercial terms, not licensed in the US / EU / UK / Republic of Korea. The example photos are CC0-1.0.