A demo of akatz-ai/MiniMax-H3-Character-Swap-LoRA,
Akatz Labs' experimental character-replacement adapter for
MiniMax-H3's ref2va partition. Hand it a scene clip and a character
reference, name who to replace, and it puts the reference character into the shot β identity, outfit and art style
carried over β while the background, camera, lighting and everyone else stay where they were. Video and its
synchronized soundtrack come out of a single denoising pass.
Two references, in the order the model reads them, and a short targeting instruction:
| Slot | Label in the prompt | What it is |
|---|---|---|
| reference 1 | <Picture 1> | the replacement character β a portrait or a full character sheet |
| reference 2 | <Video 1> | the clip to edit, 5 frames to 15 s |
Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.
That order is the one thing about this request worth being careful with, and it is not the order the prompt
names them in. MiniMax-H3 numbers a reference's label per modality, so the first image is <Picture 1> and the
first video is <Video 1> whichever way round they are packed β but the packed order still fixes the shared
audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained
with the character sheet first: control_path: [character_references, scene_videos] in
configs/trained-run-1000.json of akatz-ai/H3-Character-Swap-v1.
collect() builds the list that way and nothing else in the app reorders it.
No trigger word was trained. Strength 1.0 is the card's recommendation; 0 is the base ref2va model, which
is the comparison the adapter was judged against, so the slider doubles as an A/B.
The card is candid, and this demo does not oversell it. It is a 1,000-update experimental adapter:
Its 94 training edits targeted single still frames, with five-frame static clips standing in for <Video 1> β
which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40
preservation clips regularized, but it is the harder case.
1344x768 at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β the sample block of the LoRA's own
configs/trained-run-1000.json. The canvas dropdown follows the scene clip's aspect ratio on upload, because a
character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a
recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the
clip runs whenever that is a length MiniMax-H3 generates (2β14 s); the five-frame example clips are not, so they
fall through to the slider.
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so MiniMaxH3Blocks is
cut at its text_encoder step. This Space is the denoising half of ref2va β the transformer_ref partition
and the two autoencoders β and the 62 GiB Qwen3-VL conditioner runs in
multimodalart/qwen3vl-conditioner, which this
Space calls over the gradio API for every request. prompt_embeds + text_token_tags is the whole wire format;
reference_encoder stays here, next to the autoencoders it runs. h3_split_blocks.py is the subclass that removes
the step.
The DiT is multimodalart/MiniMax-H3-Pruned's
transformer_ref: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5
GiB instead of 61.7. That is the same checkpoint family the LoRA was trained against β ai-toolkit trained it on
Comfy-Org's minimax_h3_ref2va_pruned_int8_convrot, an int8 ConvRot quantization of these weights β and everything
the adapter touches is identical between the pruned and the released partition. The run's
network_kwargs.ignore_if_contains = ["adaln_proj"] kept it off the timestep path, which is the only place the two
differ.
Measured. The default request β 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a 41,408-row packed sequence β runs 228 s of denoise + decode on a warm worker, 10 of its 27 forwards served from the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.
GPU time is priced per request, not per Space. MiniMax-H3 attends over one packed sequence, and on this half the
references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the
generated ones. get_duration evaluates a fitted cost model over the sequence it is about to denoise β the
conditioner's exact token count plus the references measured from metadata β instead of reserving a flat ceiling for
everything, because the pool reserves whatever number it is given.
A first-block cache (h3_fbc.py, ported from duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache with an audio
exemption) skips blocks 1β49 on steps whose block-0 residual has barely moved. H3_FBC=0 restores the uncached
trajectory exactly.
There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by fully qualified name and would run straight past the PEFT branch the adapter lives in.
| Variable | Default | Meaning |
|---|---|---|
H3_LORA_SCALE | 1.0 | Default adapter strength. |
H3_CONDITIONER | multimodalart/qwen3vl-conditioner | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. |
H3_MODEL_REPO | multimodalart/MiniMax-H3-Pruned | The diffusers-layout DiT. |
H3_ATTENTION | _native_cudnn | cuDNN's fused kernel, 10β20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. |
H3_FBC / H3_FBC_THRESHOLD | 1 / 0.05 | First-block cache and its relative-L1 gate. |
H3_GPU_SIZE | xlarge | ZeroGPU allocation size. large does not fit. |
H3_PLACEMENT | lazy | Moves the partition onto the card on the first GPU call and leaves it there. |
Every request here carries an image and a video reference β the edit-on-a-real-photo case β so
hfmlsoc/ncii-light-guard-v01 screens the prompt on all of
them, before the conditioner call and before any GPU is booked. It runs in its own subprocess
(ncii_guard.py): loaded in the main process, its torch activity poisons every later ZeroGPU fork.
All five files in examples/ are from
akatz-ai/H3-Character-Swap-v1, the LoRA's own
training set β Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's CS001, CS051 and
CS090 edits, with the instructions their own captions carry:
| File | Dataset path |
|---|---|
cafe_scene.mp4 | checks/smoke-data/edits/train/scene_videos/CS001.mp4 (shared with CS051) |
character_hiker.png | .../character_references/CS001.png |
character_anime.png | .../character_references/CS051.png β a cross-style swap |
workshop_scene.mp4 | .../scene_videos/CS090.mp4 |
character_sheet_orin.png | .../character_references/CS090.png β a multi-view character sheet |
The two clips are re-encoded from the dataset's H.264 High 4:4:4 (yuv444p) to H.264 High yuv420p with the same
five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.
The adapter is distributed under the MiniMax H3 Community License Agreement, not Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets only and does not replace the model's terms.
A demo of akatz-ai/MiniMax-H3-Character-Swap-LoRA,
Akatz Labs' experimental character-replacement adapter for
MiniMax-H3's ref2va partition. Hand it a scene clip and a character
reference, name who to replace, and it puts the reference character into the shot β identity, outfit and art style
carried over β while the background, camera, lighting and everyone else stay where they were. Video and its
synchronized soundtrack come out of a single denoising pass.
Two references, in the order the model reads them, and a short targeting instruction:
| Slot | Label in the prompt | What it is |
|---|---|---|
| reference 1 | <Picture 1> | the replacement character β a portrait or a full character sheet |
| reference 2 | <Video 1> | the clip to edit, 5 frames to 15 s |
Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.
That order is the one thing about this request worth being careful with, and it is not the order the prompt
names them in. MiniMax-H3 numbers a reference's label per modality, so the first image is <Picture 1> and the
first video is <Video 1> whichever way round they are packed β but the packed order still fixes the shared
audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained
with the character sheet first: control_path: [character_references, scene_videos] in
configs/trained-run-1000.json of akatz-ai/H3-Character-Swap-v1.
collect() builds the list that way and nothing else in the app reorders it.
No trigger word was trained. Strength 1.0 is the card's recommendation; 0 is the base ref2va model, which
is the comparison the adapter was judged against, so the slider doubles as an A/B.
The card is candid, and this demo does not oversell it. It is a 1,000-update experimental adapter:
Its 94 training edits targeted single still frames, with five-frame static clips standing in for <Video 1> β
which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40
preservation clips regularized, but it is the harder case.
1344x768 at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β the sample block of the LoRA's own
configs/trained-run-1000.json. The canvas dropdown follows the scene clip's aspect ratio on upload, because a
character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a
recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the
clip runs whenever that is a length MiniMax-H3 generates (2β14 s); the five-frame example clips are not, so they
fall through to the slider.
MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so MiniMaxH3Blocks is
cut at its text_encoder step. This Space is the denoising half of ref2va β the transformer_ref partition
and the two autoencoders β and the 62 GiB Qwen3-VL conditioner runs in
multimodalart/qwen3vl-conditioner, which this
Space calls over the gradio API for every request. prompt_embeds + text_token_tags is the whole wire format;
reference_encoder stays here, next to the autoencoders it runs. h3_split_blocks.py is the subclass that removes
the step.
The DiT is multimodalart/MiniMax-H3-Pruned's
transformer_ref: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5
GiB instead of 61.7. That is the same checkpoint family the LoRA was trained against β ai-toolkit trained it on
Comfy-Org's minimax_h3_ref2va_pruned_int8_convrot, an int8 ConvRot quantization of these weights β and everything
the adapter touches is identical between the pruned and the released partition. The run's
network_kwargs.ignore_if_contains = ["adaln_proj"] kept it off the timestep path, which is the only place the two
differ.
Measured. The default request β 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a 41,408-row packed sequence β runs 228 s of denoise + decode on a warm worker, 10 of its 27 forwards served from the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.
GPU time is priced per request, not per Space. MiniMax-H3 attends over one packed sequence, and on this half the
references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the
generated ones. get_duration evaluates a fitted cost model over the sequence it is about to denoise β the
conditioner's exact token count plus the references measured from metadata β instead of reserving a flat ceiling for
everything, because the pool reserves whatever number it is given.
A first-block cache (h3_fbc.py, ported from duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache with an audio
exemption) skips blocks 1β49 on steps whose block-0 residual has barely moved. H3_FBC=0 restores the uncached
trajectory exactly.
There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by fully qualified name and would run straight past the PEFT branch the adapter lives in.
| Variable | Default | Meaning |
|---|---|---|
H3_LORA_SCALE | 1.0 | Default adapter strength. |
H3_CONDITIONER | multimodalart/qwen3vl-conditioner | The public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota. |
H3_MODEL_REPO | multimodalart/MiniMax-H3-Pruned | The diffusers-layout DiT. |
H3_ATTENTION | _native_cudnn | cuDNN's fused kernel, 10β20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for. |
H3_FBC / H3_FBC_THRESHOLD | 1 / 0.05 | First-block cache and its relative-L1 gate. |
H3_GPU_SIZE | xlarge | ZeroGPU allocation size. large does not fit. |
H3_PLACEMENT | lazy | Moves the partition onto the card on the first GPU call and leaves it there. |
Every request here carries an image and a video reference β the edit-on-a-real-photo case β so
hfmlsoc/ncii-light-guard-v01 screens the prompt on all of
them, before the conditioner call and before any GPU is booked. It runs in its own subprocess
(ncii_guard.py): loaded in the main process, its torch activity poisons every later ZeroGPU fork.
All five files in examples/ are from
akatz-ai/H3-Character-Swap-v1, the LoRA's own
training set β Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's CS001, CS051 and
CS090 edits, with the instructions their own captions carry:
| File | Dataset path |
|---|---|
cafe_scene.mp4 | checks/smoke-data/edits/train/scene_videos/CS001.mp4 (shared with CS051) |
character_hiker.png | .../character_references/CS001.png |
character_anime.png | .../character_references/CS051.png β a cross-style swap |
workshop_scene.mp4 | .../scene_videos/CS090.mp4 |
character_sheet_orin.png | .../character_references/CS090.png β a multi-view character sheet |
The two clips are re-encoded from the dataset's H.264 High 4:4:4 (yuv444p) to H.264 High yuv420p with the same
five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.
The adapter is distributed under the MiniMax H3 Community License Agreement, not Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets only and does not replace the model's terms.