Danzer1xxxxChan/H3-World

230

stars

23

commits

Python

primary language

Sep 3, 2026

updated

README

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. It generates action-controlled video from an initial frame by converting keyboard states into per-latent language instructions and binding each instruction to its corresponding future video latent through directed attention routing. Using 8,000 gameplay clips from ABot-World-Explorer-500h, H3-World learns 65.6M LoRA parameters, only 0.199% of the 33B backbone.

arXiv Hugging Face model Project page ModelScope

https://github.com/user-attachments/assets/1c862995-8809-447e-bade-2c47bfdb2738

H3-World: Turning Language Understanding into World Control
Danze Chen1,2, Zeqing Wang1,2, Ziyue Lin3, Xingyi Yang3, Yeying Jin1,2
1Tencent   2National University of Singapore   3The Hong Kong Polytechnic University

⚙️ Setup

Tested with Python 3.10 and CUDA 12.8.

conda create -n minimax_h3 python=3.10 -y
conda activate minimax_h3
pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

# Use the exact DiffSynth revision and the H3-World attention patch.
git clone https://github.com/modelscope/DiffSynth-Studio.git DiffSynth-Studio-h3-v2
git -C DiffSynth-Studio-h3-v2 checkout "$(cat code/diffsynth_base_commit.txt)"
git -C DiffSynth-Studio-h3-v2 apply ../code/diffsynth_h3_action.patch

# Keep Hugging Face, Torch, and Triton caches inside this repository.
source env.sh

Download the required weights into the following locations:

AssetRequired forTarget location
MiniMax-H3 base weights (about 135 GB)inference and trainingDiffSynth-Studio-h3-v2/models/MiniMax/MiniMax-H3/
H3-World LoRAinferencecheckpoints/H3-World/step-10000.safetensors
ABot-World-Explorer-500htraining onlyany local path, passed through ABOT_SRC_ROOT
python3 -c "
from huggingface_hub import snapshot_download
snapshot_download('MiniMax/MiniMax-H3', local_dir='DiffSynth-Studio-h3-v2/models/MiniMax/MiniMax-H3')"

python3 -c "
from huggingface_hub import hf_hub_download
hf_hub_download('DANNY621/H3-World', 'step-10000.safetensors', local_dir='checkpoints/H3-World')"

The patch is required for the released checkpoint. Do not install DiffSynth-Studio-h3-v2 in editable mode; the included training and inference scripts verify that they load the patched checkout.

🎬 Inference

The repository includes a held-out ABot test frame at examples/first_frame.png. The following fixed configuration generates a 5.2-second forward-motion video:

python3 code/abot/infer.py \
  --checkpoint checkpoints/H3-World/step-10000.safetensors \
  --first-frame examples/first_frame.png \
  --scene-prompt "A man in a yellow floral shirt stands in a dim, multi-level concrete parking garage." \
  --action-preset forward \
  --seed 2 \
  --steps 50 \
  --num-frames 124 \
  --cfg-scale 1.0 \
  --out outputs/example_forward.mp4

The included frame is sample d0b768c6 from the held-out test split. To use a custom image, replace examples/first_frame.png and describe its static scene and subject with --scene-prompt. Inputs are center-cropped to 832x480 when needed. The built-in presets are still, forward, back, strafe-left, strafe-right, tilt-up, tilt-down, pan-left, pan-right, pan-left-fast, and pan-right-fast. The full key-to-language mapping is defined in code/abot/action_script.py.

🏋️ Training

Prepare the 7,872-clip training split, cache its latents, inject the per-latent action instructions, then train LoRA:

# 1. Build clips and the fixed train/test split from ABot.
ABOT_SRC_ROOT=/path/to/ABot-World-Explorer-500h \
  python3 code/abot/build_abot_clips.py --num-clips 8000 --workers 48
ABOT_SRC_ROOT=/path/to/ABot-World-Explorer-500h \
  python3 code/abot/build_abot_clips.py --verify 8
python3 code/abot/split_abot_metadata.py \
  --input data/abot_meta_8000.jsonl \
  --train-output data/abot_meta_train_7872.jsonl \
  --test-output data/abot_meta_test_128.jsonl \
  --clips-dir data/clips

# 2. Cache the video, audio, and text latents, then add action text.
bash code/cache.sh
python3 code/abot/inject_abot_text.py \
  --meta data/abot_meta_train_7872.jsonl \
  --cache output/minimax_h3_abot/7872-cache \
  --device cuda:0

# 3. Train on four GPUs by default.
bash code/train.sh

code/train.sh uses rank-32 LoRA on qkv_proj and out_proj for 20 epochs, saving checkpoints every 2,000 steps. Override the visible devices with CUDA_VISIBLE_DEVICES=4,5,6,7 bash code/train.sh.

🙏 Acknowledgements

📚 Citation

@misc{chen2026h3worldturninglanguageunderstanding,
      title={H3-World: Turning Language Understanding into World Control},
      author={Danze Chen and Zeqing Wang and Ziyue Lin and Xingyi Yang and Yeying Jin},
      year={2026},
      eprint={2609.01560},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.01560},
}

Contributors

Danzer1xxxxChan

13 commits

danzerchan-png

10 commits

Danzer1xxxxChan/H3-World

230

stars

23

commits

Python

primary language

Sep 3, 2026

updated

README

H3-World: Turning Language Understanding into World Control

H3-World is the first interactive world model built on MiniMax-H3. It generates action-controlled video from an initial frame by converting keyboard states into per-latent language instructions and binding each instruction to its corresponding future video latent through directed attention routing. Using 8,000 gameplay clips from ABot-World-Explorer-500h, H3-World learns 65.6M LoRA parameters, only 0.199% of the 33B backbone.

arXiv Hugging Face model Project page ModelScope

https://github.com/user-attachments/assets/1c862995-8809-447e-bade-2c47bfdb2738

H3-World: Turning Language Understanding into World Control
Danze Chen1,2, Zeqing Wang1,2, Ziyue Lin3, Xingyi Yang3, Yeying Jin1,2
1Tencent   2National University of Singapore   3The Hong Kong Polytechnic University

⚙️ Setup

Tested with Python 3.10 and CUDA 12.8.

conda create -n minimax_h3 python=3.10 -y
conda activate minimax_h3
pip install torch==2.10.0 --index-url https://download.pytorch.org/whl/cu128
pip install -r requirements.txt

# Use the exact DiffSynth revision and the H3-World attention patch.
git clone https://github.com/modelscope/DiffSynth-Studio.git DiffSynth-Studio-h3-v2
git -C DiffSynth-Studio-h3-v2 checkout "$(cat code/diffsynth_base_commit.txt)"
git -C DiffSynth-Studio-h3-v2 apply ../code/diffsynth_h3_action.patch

# Keep Hugging Face, Torch, and Triton caches inside this repository.
source env.sh

Download the required weights into the following locations:

AssetRequired forTarget location
MiniMax-H3 base weights (about 135 GB)inference and trainingDiffSynth-Studio-h3-v2/models/MiniMax/MiniMax-H3/
H3-World LoRAinferencecheckpoints/H3-World/step-10000.safetensors
ABot-World-Explorer-500htraining onlyany local path, passed through ABOT_SRC_ROOT
python3 -c "
from huggingface_hub import snapshot_download
snapshot_download('MiniMax/MiniMax-H3', local_dir='DiffSynth-Studio-h3-v2/models/MiniMax/MiniMax-H3')"

python3 -c "
from huggingface_hub import hf_hub_download
hf_hub_download('DANNY621/H3-World', 'step-10000.safetensors', local_dir='checkpoints/H3-World')"

The patch is required for the released checkpoint. Do not install DiffSynth-Studio-h3-v2 in editable mode; the included training and inference scripts verify that they load the patched checkout.

🎬 Inference

The repository includes a held-out ABot test frame at examples/first_frame.png. The following fixed configuration generates a 5.2-second forward-motion video:

python3 code/abot/infer.py \
  --checkpoint checkpoints/H3-World/step-10000.safetensors \
  --first-frame examples/first_frame.png \
  --scene-prompt "A man in a yellow floral shirt stands in a dim, multi-level concrete parking garage." \
  --action-preset forward \
  --seed 2 \
  --steps 50 \
  --num-frames 124 \
  --cfg-scale 1.0 \
  --out outputs/example_forward.mp4

The included frame is sample d0b768c6 from the held-out test split. To use a custom image, replace examples/first_frame.png and describe its static scene and subject with --scene-prompt. Inputs are center-cropped to 832x480 when needed. The built-in presets are still, forward, back, strafe-left, strafe-right, tilt-up, tilt-down, pan-left, pan-right, pan-left-fast, and pan-right-fast. The full key-to-language mapping is defined in code/abot/action_script.py.

🏋️ Training

Prepare the 7,872-clip training split, cache its latents, inject the per-latent action instructions, then train LoRA:

# 1. Build clips and the fixed train/test split from ABot.
ABOT_SRC_ROOT=/path/to/ABot-World-Explorer-500h \
  python3 code/abot/build_abot_clips.py --num-clips 8000 --workers 48
ABOT_SRC_ROOT=/path/to/ABot-World-Explorer-500h \
  python3 code/abot/build_abot_clips.py --verify 8
python3 code/abot/split_abot_metadata.py \
  --input data/abot_meta_8000.jsonl \
  --train-output data/abot_meta_train_7872.jsonl \
  --test-output data/abot_meta_test_128.jsonl \
  --clips-dir data/clips

# 2. Cache the video, audio, and text latents, then add action text.
bash code/cache.sh
python3 code/abot/inject_abot_text.py \
  --meta data/abot_meta_train_7872.jsonl \
  --cache output/minimax_h3_abot/7872-cache \
  --device cuda:0

# 3. Train on four GPUs by default.
bash code/train.sh

code/train.sh uses rank-32 LoRA on qkv_proj and out_proj for 20 epochs, saving checkpoints every 2,000 steps. Override the visible devices with CUDA_VISIBLE_DEVICES=4,5,6,7 bash code/train.sh.

🙏 Acknowledgements

📚 Citation

@misc{chen2026h3worldturninglanguageunderstanding,
      title={H3-World: Turning Language Understanding into World Control},
      author={Danze Chen and Zeqing Wang and Ziyue Lin and Xingyi Yang and Yeying Jin},
      year={2026},
      eprint={2609.01560},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.01560},
}

Contributors

Danzer1xxxxChan

13 commits

danzerchan-png

10 commits

Languages

Python

94.9%

Shell

5.1%