hugging-apps/minimax-h3-character-swap-lora

Space

MiniMax-H3 Character Swap LoRA

18

7 commits

updated Sep 29, 2026

See the code

README

MiniMax-H3 Character Swap LoRA

A demo of akatz-ai/MiniMax-H3-Character-Swap-LoRA, Akatz Labs' experimental character-replacement adapter for MiniMax-H3's ref2va partition. Hand it a scene clip and a character reference, name who to replace, and it puts the reference character into the shot β€” identity, outfit and art style carried over β€” while the background, camera, lighting and everyone else stay where they were. Video and its synchronized soundtrack come out of a single denoising pass.

The request

Two references, in the order the model reads them, and a short targeting instruction:

SlotLabel in the promptWhat it is
reference 1<Picture 1>the replacement character β€” a portrait or a full character sheet
reference 2<Video 1>the clip to edit, 5 frames to 15 s

Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.

That order is the one thing about this request worth being careful with, and it is not the order the prompt names them in. MiniMax-H3 numbers a reference's label per modality, so the first image is <Picture 1> and the first video is <Video 1> whichever way round they are packed β€” but the packed order still fixes the shared audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained with the character sheet first: control_path: [character_references, scene_videos] in configs/trained-run-1000.json of akatz-ai/H3-Character-Swap-v1. collect() builds the list that way and nothing else in the app reorders it.

No trigger word was trained. Strength 1.0 is the card's recommendation; 0 is the base ref2va model, which is the comparison the adapter was judged against, so the slider doubles as an A/B.

What it is good at, and what it is not

The card is candid, and this demo does not oversell it. It is a 1,000-update experimental adapter:

  • background and scene preservation improved over the base model in the author's local comparisons β€” qualitative, not a benchmark,
  • motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual reposition,
  • long windows drift in framing and placement. Short continuous shots of roughly 3–5 s are the promising range,
  • two-character inference was tested but multi-character replacement was never supervised,
  • the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which does not repair lip-sync drift.

Its 94 training edits targeted single still frames, with five-frame static clips standing in for <Video 1> β€” which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40 preservation clips regularized, but it is the harder case.

Defaults, and where they come from

1344x768 at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β€” the sample block of the LoRA's own configs/trained-run-1000.json. The canvas dropdown follows the scene clip's aspect ratio on upload, because a character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they fall through to the slider.

How it is deployed

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so MiniMaxH3Blocks is cut at its text_encoder step. This Space is the denoising half of ref2va β€” the transformer_ref partition and the two autoencoders β€” and the 62 GiB Qwen3-VL conditioner runs in multimodalart/qwen3vl-conditioner, which this Space calls over the gradio API for every request. prompt_embeds + text_token_tags is the whole wire format; reference_encoder stays here, next to the autoencoders it runs. h3_split_blocks.py is the subclass that removes the step.

The DiT is multimodalart/MiniMax-H3-Pruned's transformer_ref: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5 GiB instead of 61.7. That is the same checkpoint family the LoRA was trained against β€” ai-toolkit trained it on Comfy-Org's minimax_h3_ref2va_pruned_int8_convrot, an int8 ConvRot quantization of these weights β€” and everything the adapter touches is identical between the pruned and the released partition. The run's network_kwargs.ignore_if_contains = ["adaln_proj"] kept it off the timestep path, which is the only place the two differ.

Measured. The default request β€” 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a 41,408-row packed sequence β€” runs 228 s of denoise + decode on a warm worker, 10 of its 27 forwards served from the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.

GPU time is priced per request, not per Space. MiniMax-H3 attends over one packed sequence, and on this half the references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the generated ones. get_duration evaluates a fitted cost model over the sequence it is about to denoise β€” the conditioner's exact token count plus the references measured from metadata β€” instead of reserving a flat ceiling for everything, because the pool reserves whatever number it is given.

A first-block cache (h3_fbc.py, ported from duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache with an audio exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. H3_FBC=0 restores the uncached trajectory exactly.

There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by fully qualified name and would run straight past the PEFT branch the adapter lives in.

Space variables

VariableDefaultMeaning
H3_LORA_SCALE1.0Default adapter strength.
H3_CONDITIONERmultimodalart/qwen3vl-conditionerThe public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota.
H3_MODEL_REPOmultimodalart/MiniMax-H3-PrunedThe diffusers-layout DiT.
H3_ATTENTION_native_cudnncuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for.
H3_FBC / H3_FBC_THRESHOLD1 / 0.05First-block cache and its relative-L1 gate.
H3_GPU_SIZExlargeZeroGPU allocation size. large does not fit.
H3_PLACEMENTlazyMoves the partition onto the card on the first GPU call and leaves it there.

Safety

Every request here carries an image and a video reference β€” the edit-on-a-real-photo case β€” so hfmlsoc/ncii-light-guard-v01 screens the prompt on all of them, before the conditioner call and before any GPU is booked. It runs in its own subprocess (ncii_guard.py): loaded in the main process, its torch activity poisons every later ZeroGPU fork.

Example assets

All five files in examples/ are from akatz-ai/H3-Character-Swap-v1, the LoRA's own training set β€” Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's CS001, CS051 and CS090 edits, with the instructions their own captions carry:

FileDataset path
cafe_scene.mp4checks/smoke-data/edits/train/scene_videos/CS001.mp4 (shared with CS051)
character_hiker.png.../character_references/CS001.png
character_anime.png.../character_references/CS051.png β€” a cross-style swap
workshop_scene.mp4.../scene_videos/CS090.mp4
character_sheet_orin.png.../character_references/CS090.png β€” a multi-view character sheet

The two clips are re-encoded from the dataset's H.264 High 4:4:4 (yuv444p) to H.264 High yuv420p with the same five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.

License

The adapter is distributed under the MiniMax H3 Community License Agreement, not Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets only and does not replace the model's terms.

gradio

hugging-apps/minimax-h3-character-swap-lora

Space

MiniMax-H3 Character Swap LoRA

18

7 commits

updated Sep 29, 2026

See the code

README

MiniMax-H3 Character Swap LoRA

A demo of akatz-ai/MiniMax-H3-Character-Swap-LoRA, Akatz Labs' experimental character-replacement adapter for MiniMax-H3's ref2va partition. Hand it a scene clip and a character reference, name who to replace, and it puts the reference character into the shot β€” identity, outfit and art style carried over β€” while the background, camera, lighting and everyone else stay where they were. Video and its synchronized soundtrack come out of a single denoising pass.

The request

Two references, in the order the model reads them, and a short targeting instruction:

SlotLabel in the promptWhat it is
reference 1<Picture 1>the replacement character β€” a portrait or a full character sheet
reference 2<Video 1>the clip to edit, 5 frames to 15 s

Swap the man in the purple shirt in <Video 1> with the character in <Picture 1>.

That order is the one thing about this request worth being careful with, and it is not the order the prompt names them in. MiniMax-H3 numbers a reference's label per modality, so the first image is <Picture 1> and the first video is <Video 1> whichever way round they are packed β€” but the packed order still fixes the shared audio/video rotary clock, so the same two references swapped round are a different request. The LoRA was trained with the character sheet first: control_path: [character_references, scene_videos] in configs/trained-run-1000.json of akatz-ai/H3-Character-Swap-v1. collect() builds the list that way and nothing else in the app reorders it.

No trigger word was trained. Strength 1.0 is the card's recommendation; 0 is the base ref2va model, which is the comparison the adapter was judged against, so the slider doubles as an A/B.

What it is good at, and what it is not

The card is candid, and this demo does not oversell it. It is a 1,000-update experimental adapter:

  • background and scene preservation improved over the base model in the author's local comparisons β€” qualitative, not a benchmark,
  • motion timing, facial expressions and hard cuts remain unreliable; a hard cut can become a zoom or a gradual reposition,
  • long windows drift in framing and placement. Short continuous shots of roughly 3–5 s are the promising range,
  • two-character inference was tested but multi-character replacement was never supervised,
  • the soundtrack is generated, not carried over. Audio preservation in the author's later review was a remux, which does not repair lip-sync drift.

Its 94 training edits targeted single still frames, with five-frame static clips standing in for <Video 1> β€” which is exactly what the examples below are. Real moving footage goes in the same slot and is what the 40 preservation clips regularized, but it is the harder case.

Defaults, and where they come from

1344x768 at 24 fps, 73 frames (3.04 s), 28 steps, seed 904231 β€” the sample block of the LoRA's own configs/trained-run-1000.json. The canvas dropdown follows the scene clip's aspect ratio on upload, because a character swap is asked to keep the source framing and a portrait clip generated on a landscape canvas is a recomposed shot before the model has done anything. "Match the scene clip's length" generates for as long as the clip runs whenever that is a length MiniMax-H3 generates (2–14 s); the five-frame example clips are not, so they fall through to the slider.

How it is deployed

MiniMax-H3 is 195.9 GiB in bfloat16 and a ZeroGPU Space is evicted at 150 GB of storage, so MiniMaxH3Blocks is cut at its text_encoder step. This Space is the denoising half of ref2va β€” the transformer_ref partition and the two autoencoders β€” and the 62 GiB Qwen3-VL conditioner runs in multimodalart/qwen3vl-conditioner, which this Space calls over the gradio API for every request. prompt_embeds + text_token_tags is the whole wire format; reference_encoder stays here, next to the autoencoders it runs. h3_split_blocks.py is the subclass that removes the step.

The DiT is multimodalart/MiniMax-H3-Pruned's transformer_ref: the released partition with its AdaLN input projections folded onto their reachable rank, 37.5 GiB instead of 61.7. That is the same checkpoint family the LoRA was trained against β€” ai-toolkit trained it on Comfy-Org's minimax_h3_ref2va_pruned_int8_convrot, an int8 ConvRot quantization of these weights β€” and everything the adapter touches is identical between the pruned and the released partition. The run's network_kwargs.ignore_if_contains = ["adaln_proj"] kept it off the timestep path, which is the only place the two differ.

Measured. The default request β€” 1344x768, 73 frames, 28 steps, one character sheet and one scene clip, a 41,408-row packed sequence β€” runs 228 s of denoise + decode on a warm worker, 10 of its 27 forwards served from the first-block cache, and 242 s end to end including the conditioner round trip. The reservation is larger than that on purpose: it has to cover a cold worker's placement and a request the cache skips nothing on.

GPU time is priced per request, not per Space. MiniMax-H3 attends over one packed sequence, and on this half the references dominate its length: a 2048-short-edge character sheet is thousands of conditioning rows on top of the generated ones. get_duration evaluates a fitted cost model over the sequence it is about to denoise β€” the conditioner's exact token count plus the references measured from metadata β€” instead of reserving a flat ceiling for everything, because the pool reserves whatever number it is given.

A first-block cache (h3_fbc.py, ported from duckyshell/ComfyUI-MiniMaxH3-FirstBlockCache with an audio exemption) skips blocks 1–49 on steps whose block-0 residual has barely moved. H3_FBC=0 restores the uncached trajectory exactly.

There is no AoTI on this Space, unlike its siblings: a compiled block package binds the base module's weights by fully qualified name and would run straight past the PEFT branch the adapter lives in.

Space variables

VariableDefaultMeaning
H3_LORA_SCALE1.0Default adapter strength.
H3_CONDITIONERmultimodalart/qwen3vl-conditionerThe public Space this one asks for embeddings; the client passes no token, so the call runs on the caller's own quota.
H3_MODEL_REPOmultimodalart/MiniMax-H3-PrunedThe diffusers-layout DiT.
H3_ATTENTION_native_cudnncuDNN's fused kernel, 10–20% faster than the SDPA default. The two float32 VAEs are pinned to torch SDPA, which cuDNN has no kernel for.
H3_FBC / H3_FBC_THRESHOLD1 / 0.05First-block cache and its relative-L1 gate.
H3_GPU_SIZExlargeZeroGPU allocation size. large does not fit.
H3_PLACEMENTlazyMoves the partition onto the card on the first GPU call and leaves it there.

Safety

Every request here carries an image and a video reference β€” the edit-on-a-real-photo case β€” so hfmlsoc/ncii-light-guard-v01 screens the prompt on all of them, before the conditioner call and before any GPU is booked. It runs in its own subprocess (ncii_guard.py): loaded in the main process, its torch activity poisons every later ZeroGPU fork.

Example assets

All five files in examples/ are from akatz-ai/H3-Character-Swap-v1, the LoRA's own training set β€” Apache-2.0 for Akatz Labs' synthetic contributions. They are the dataset's CS001, CS051 and CS090 edits, with the instructions their own captions carry:

FileDataset path
cafe_scene.mp4checks/smoke-data/edits/train/scene_videos/CS001.mp4 (shared with CS051)
character_hiker.png.../character_references/CS001.png
character_anime.png.../character_references/CS051.png β€” a cross-style swap
workshop_scene.mp4.../scene_videos/CS090.mp4
character_sheet_orin.png.../character_references/CS090.png β€” a multi-view character sheet

The two clips are re-encoded from the dataset's H.264 High 4:4:4 (yuv444p) to H.264 High yuv420p with the same five frames. Browsers cannot decode 4:4:4 H.264, so the originals would not play in the examples.

License

The adapter is distributed under the MiniMax H3 Community License Agreement, not Apache-2.0, and that agreement excludes the US, EU, UK and Republic of Korea from its standard territorial grant. Read the upstream terms; nothing here extends them. The dataset's own Apache-2.0 covers the example assets only and does not replace the model's terms.

gradio