pdmd2026/pdmd_2NFE_lora

Model

PDMD 2-NFE LoRA for MiniMax-H3

4

15 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

PDMD 2-NFE LoRA for MiniMax-H3

Paper (arXiv) · Project page

This repository holds the LoRA adapter of a 2-step (2 NFE) student distilled from MiniMax-H3 with Projected Distribution Matching Distillation (PDMD). PDMD projects the DMD update onto the subspace orthogonal to the student–critic endpoint residual; it is a one-line change to DMD with no auxiliary loss, extra network, extra model pass, or extra training stage.

The adapter is applied to the transformer (MiniMaxH3Transformer3DModel) of the base model. Every other component (VAE, audio VAE, schedulers, text encoder, processor) is unchanged and comes from the base model.

Files

filebytesdescription
lora_model_0.safetensors1,383,680,648312 LoRA pairs (624 tensors), rank 128, alpha 128, bf16
lora_model_0.safetensors.json96,508metadata: tensor keys, shapes, base-weight check
lora_model_0_fp32.safetensors2,767,276,568the same LoRA in fp32
lora_model_0_fp32.safetensors.json96,507metadata of the fp32 file

Tensor keys have the form transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight, where <module>.weight is the corresponding parameter of MiniMaxH3Transformer3DModel. Fusing rule (also stored in the safetensors metadata):

W_base += lora_scale * (lora_B @ lora_A)      # lora_scale = 1.0 (alpha / rank = 128 / 128)

Usage

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from diffusers.models.transformers.transformer_minimax_h3 import MiniMaxH3Transformer3DModel

# 1. Load the base MiniMax-H3 transformer (diffusers format, `transformer/` sub-folder).
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
    "<path-or-repo-of-MiniMax-H3-diffusers>", subfolder="transformer", torch_dtype=torch.bfloat16
)

# 2. Fuse the PDMD 2-NFE LoRA into the base weights.
lora = load_file(hf_hub_download("pdmd2026/pdmd_2NFE_lora", "lora_model_0.safetensors"))
params = dict(transformer.named_parameters())
with torch.no_grad():
    for k, A in lora.items():
        if not k.endswith(".lora_A.weight"):
            continue
        B = lora[k.replace(".lora_A.", ".lora_B.")]
        name = k[len("transformer."):].replace(".lora_A.weight", ".weight")
        W = params[name]
        W.add_((B.float() @ A.float()).to(W.dtype))   # lora_scale = 1.0

# 3. Build the MiniMax-H3 pipeline from the base checkpoint and swap in this transformer.

Sample with 2 denoising steps, using video time shift 12. For 2-NFE generation, we recommend audio time shift 6; use 3 for paper metrics. Credit to CALMDUST (@core_tan) on X for the suggestion. Example settings used for our reference renders: 1344×768, 345 frames at 24 fps, seed 42.

Training summary

  • Base: MiniMax-H3 (33B DiT, joint video–audio).
  • Student: LoRA rank 128 on the transformer, trained with PDMD (projection mode critic_endpoint_residual_perpendicular), 2 sampling segments (NFE 2), critic:student update ratio 6:1, lr 5e-5 / critic lr 1e-5.
  • Checkpoint: training step 4000.

Citation

@misc{wang2026pdmdprojecteddistributionmatching,
      title={PDMD: Projected Distribution Matching Distillation for Video Diffusion Models},
      author={Zimo Wang and Junkun Yuan and Angtian Wang and Haotian Yang and Canyu Zhang and Siyuan Yuan and Xingchang Huang and Bo Liu and Yizhi Wang and Yiding Yang and Chongyang Ma and Gordon Guocheng Qian},
      year={2026},
      eprint={2609.35768},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.35768},
}
diffusers
distillation
distribution-matching-distillation
lora
pdmd
text-to-video
video-generation

pdmd2026/pdmd_2NFE_lora

Model

PDMD 2-NFE LoRA for MiniMax-H3

4

15 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

PDMD 2-NFE LoRA for MiniMax-H3

Paper (arXiv) · Project page

This repository holds the LoRA adapter of a 2-step (2 NFE) student distilled from MiniMax-H3 with Projected Distribution Matching Distillation (PDMD). PDMD projects the DMD update onto the subspace orthogonal to the student–critic endpoint residual; it is a one-line change to DMD with no auxiliary loss, extra network, extra model pass, or extra training stage.

The adapter is applied to the transformer (MiniMaxH3Transformer3DModel) of the base model. Every other component (VAE, audio VAE, schedulers, text encoder, processor) is unchanged and comes from the base model.

Files

filebytesdescription
lora_model_0.safetensors1,383,680,648312 LoRA pairs (624 tensors), rank 128, alpha 128, bf16
lora_model_0.safetensors.json96,508metadata: tensor keys, shapes, base-weight check
lora_model_0_fp32.safetensors2,767,276,568the same LoRA in fp32
lora_model_0_fp32.safetensors.json96,507metadata of the fp32 file

Tensor keys have the form transformer.<module>.lora_A.weight / transformer.<module>.lora_B.weight, where <module>.weight is the corresponding parameter of MiniMaxH3Transformer3DModel. Fusing rule (also stored in the safetensors metadata):

W_base += lora_scale * (lora_B @ lora_A)      # lora_scale = 1.0 (alpha / rank = 128 / 128)

Usage

import torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from diffusers.models.transformers.transformer_minimax_h3 import MiniMaxH3Transformer3DModel

# 1. Load the base MiniMax-H3 transformer (diffusers format, `transformer/` sub-folder).
transformer = MiniMaxH3Transformer3DModel.from_pretrained(
    "<path-or-repo-of-MiniMax-H3-diffusers>", subfolder="transformer", torch_dtype=torch.bfloat16
)

# 2. Fuse the PDMD 2-NFE LoRA into the base weights.
lora = load_file(hf_hub_download("pdmd2026/pdmd_2NFE_lora", "lora_model_0.safetensors"))
params = dict(transformer.named_parameters())
with torch.no_grad():
    for k, A in lora.items():
        if not k.endswith(".lora_A.weight"):
            continue
        B = lora[k.replace(".lora_A.", ".lora_B.")]
        name = k[len("transformer."):].replace(".lora_A.weight", ".weight")
        W = params[name]
        W.add_((B.float() @ A.float()).to(W.dtype))   # lora_scale = 1.0

# 3. Build the MiniMax-H3 pipeline from the base checkpoint and swap in this transformer.

Sample with 2 denoising steps, using video time shift 12. For 2-NFE generation, we recommend audio time shift 6; use 3 for paper metrics. Credit to CALMDUST (@core_tan) on X for the suggestion. Example settings used for our reference renders: 1344×768, 345 frames at 24 fps, seed 42.

Training summary

  • Base: MiniMax-H3 (33B DiT, joint video–audio).
  • Student: LoRA rank 128 on the transformer, trained with PDMD (projection mode critic_endpoint_residual_perpendicular), 2 sampling segments (NFE 2), critic:student update ratio 6:1, lr 5e-5 / critic lr 1e-5.
  • Checkpoint: training step 4000.

Citation

@misc{wang2026pdmdprojecteddistributionmatching,
      title={PDMD: Projected Distribution Matching Distillation for Video Diffusion Models},
      author={Zimo Wang and Junkun Yuan and Angtian Wang and Haotian Yang and Canyu Zhang and Siyuan Yuan and Xingchang Huang and Bo Liu and Yizhi Wang and Yiding Yang and Chongyang Ma and Gordon Guocheng Qian},
      year={2026},
      eprint={2609.35768},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.35768},
}
diffusers
distillation
distribution-matching-distillation
lora
pdmd
text-to-video
video-generation