This repository holds the LoRA adapter of a 2-step (2 NFE) student distilled from MiniMax-H3 with Projected Distribution Matching Distillation (PDMD). PDMD projects the DMD update onto the subspace orthogonal to the student–critic endpoint residual; it is a one-line change to DMD with no auxiliary loss, extra network, extra model pass, or extra training stage.
The adapter is applied to the transformer (MiniMaxH3Transformer3DModel) of the base model.
Every other component (VAE, audio VAE, schedulers, text encoder, processor) is unchanged and
comes from the base model.
| file | bytes | description |
|---|---|---|
lora_model_0.safetensors | 1,383,680,648 | 312 LoRA pairs (624 tensors), rank 128, alpha 128, bf16 |
lora_model_0.safetensors.json | 96,508 | metadata: tensor keys, shapes, base-weight check |
lora_model_0_fp32.safetensors | 2,767,276,568 | the same LoRA in fp32 |
lora_model_0_fp32.safetensors.json | 96,507 | metadata of the fp32 file |
Tensor keys have the form transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight,
where <module>.weight is the corresponding parameter of MiniMaxH3Transformer3DModel.
Fusing rule (also stored in the safetensors metadata):
W_base += lora_scale * (lora_B @ lora_A) # lora_scale = 1.0 (alpha / rank = 128 / 128)
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from diffusers.models.transformers.transformer_minimax_h3 import MiniMaxH3Transformer3DModel
# 1. Load the base MiniMax-H3 transformer (diffusers format, `transformer/` sub-folder).
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
"<path-or-repo-of-MiniMax-H3-diffusers>", subfolder="transformer", torch_dtype=torch.bfloat16
)
# 2. Fuse the PDMD 2-NFE LoRA into the base weights.
lora = load_file(hf_hub_download("pdmd2026/pdmd_2NFE_lora", "lora_model_0.safetensors"))
params = dict(transformer.named_parameters())
with torch.no_grad():
for k, A in lora.items():
if not k.endswith(".lora_A.weight"):
continue
B = lora[k.replace(".lora_A.", ".lora_B.")]
name = k[len("transformer."):].replace(".lora_A.weight", ".weight")
W = params[name]
W.add_((B.float() @ A.float()).to(W.dtype)) # lora_scale = 1.0
# 3. Build the MiniMax-H3 pipeline from the base checkpoint and swap in this transformer.
Sample with 2 denoising steps, using video time shift 12. For 2-NFE generation, we recommend audio time shift 6; use 3 for paper metrics. Credit to CALMDUST (@core_tan) on X for the suggestion. Example settings used for our reference renders: 1344×768, 345 frames at 24 fps, seed 42.
critic_endpoint_residual_perpendicular), 2 sampling segments (NFE 2), critic:student
update ratio 6:1, lr 5e-5 / critic lr 1e-5.@misc{wang2026pdmdprojecteddistributionmatching,
title={PDMD: Projected Distribution Matching Distillation for Video Diffusion Models},
author={Zimo Wang and Junkun Yuan and Angtian Wang and Haotian Yang and Canyu Zhang and Siyuan Yuan and Xingchang Huang and Bo Liu and Yizhi Wang and Yiding Yang and Chongyang Ma and Gordon Guocheng Qian},
year={2026},
eprint={2609.35768},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.35768},
}
This repository holds the LoRA adapter of a 2-step (2 NFE) student distilled from MiniMax-H3 with Projected Distribution Matching Distillation (PDMD). PDMD projects the DMD update onto the subspace orthogonal to the student–critic endpoint residual; it is a one-line change to DMD with no auxiliary loss, extra network, extra model pass, or extra training stage.
The adapter is applied to the transformer (MiniMaxH3Transformer3DModel) of the base model.
Every other component (VAE, audio VAE, schedulers, text encoder, processor) is unchanged and
comes from the base model.
| file | bytes | description |
|---|---|---|
lora_model_0.safetensors | 1,383,680,648 | 312 LoRA pairs (624 tensors), rank 128, alpha 128, bf16 |
lora_model_0.safetensors.json | 96,508 | metadata: tensor keys, shapes, base-weight check |
lora_model_0_fp32.safetensors | 2,767,276,568 | the same LoRA in fp32 |
lora_model_0_fp32.safetensors.json | 96,507 | metadata of the fp32 file |
Tensor keys have the form transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight,
where <module>.weight is the corresponding parameter of MiniMaxH3Transformer3DModel.
Fusing rule (also stored in the safetensors metadata):
W_base += lora_scale * (lora_B @ lora_A) # lora_scale = 1.0 (alpha / rank = 128 / 128)
import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from diffusers.models.transformers.transformer_minimax_h3 import MiniMaxH3Transformer3DModel
# 1. Load the base MiniMax-H3 transformer (diffusers format, `transformer/` sub-folder).
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
"<path-or-repo-of-MiniMax-H3-diffusers>", subfolder="transformer", torch_dtype=torch.bfloat16
)
# 2. Fuse the PDMD 2-NFE LoRA into the base weights.
lora = load_file(hf_hub_download("pdmd2026/pdmd_2NFE_lora", "lora_model_0.safetensors"))
params = dict(transformer.named_parameters())
with torch.no_grad():
for k, A in lora.items():
if not k.endswith(".lora_A.weight"):
continue
B = lora[k.replace(".lora_A.", ".lora_B.")]
name = k[len("transformer."):].replace(".lora_A.weight", ".weight")
W = params[name]
W.add_((B.float() @ A.float()).to(W.dtype)) # lora_scale = 1.0
# 3. Build the MiniMax-H3 pipeline from the base checkpoint and swap in this transformer.
Sample with 2 denoising steps, using video time shift 12. For 2-NFE generation, we recommend audio time shift 6; use 3 for paper metrics. Credit to CALMDUST (@core_tan) on X for the suggestion. Example settings used for our reference renders: 1344×768, 345 frames at 24 fps, seed 42.
critic_endpoint_residual_perpendicular), 2 sampling segments (NFE 2), critic:student
update ratio 6:1, lr 5e-5 / critic lr 1e-5.@misc{wang2026pdmdprojecteddistributionmatching,
title={PDMD: Projected Distribution Matching Distillation for Video Diffusion Models},
author={Zimo Wang and Junkun Yuan and Angtian Wang and Haotian Yang and Canyu Zhang and Siyuan Yuan and Xingchang Huang and Bo Liu and Yizhi Wang and Yiding Yang and Chongyang Ma and Gordon Guocheng Qian},
year={2026},
eprint={2609.35768},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.35768},
}