kandinskylab/kandinsky-wm

Kandinsky WM 1.0 — a family of models for Physical AI. Image-to-video generation for autonomous driving, robotics & general physics.

Python

10

4 commits

updated Aug 5, 2026

See the code

README

Kandinsky WM

Kandinsky WM 1.0: A family of models for Physical AI

Image-to-Video generation for Physical AI: autonomous driving · robotics · general physics

🤗 Checkpoints collection · 🧩 Base model (Kandinsky 5.0) · 📝 Habr article


Contents


Overview

Kandinsky WM (World Model) 1.0 is a family of Image-to-Video models that adapt Kandinsky 5.0 Video Lite — a 2B-parameter latent video diffusion model (a Kandinsky5Transformer3DModel DiT paired with a HunyuanVideo VAE and Qwen2.5-VL + CLIP text encoders, trained with flow matching) — to Physical AI: video generation that is not only visually convincing but physically plausible — consistent scene geometry, object dynamics, interactions and cause-and-effect — so the model can serve as a source of synthetic training data and as a building block for world models and simulators.

The base model was domain-adapted on large corpora of 5-second scenes (first text-to-video, then a mixed text/image-to-video regime that preserves first-frame continuation), spanning autonomous driving, robotics, and a general domain (industrial processes, physical phenomena, and human–object / human–human interaction). Each checkpoint is then reinforcement-learning post-trained with GRPO against a reward model that scores the physical plausibility of the generated scene, steering the generator toward more faithful geometry, object behaviour and interactions than domain fine-tuning alone.

Each checkpoint generates 5-second, 121-frame clips at 768×512.

Model Zoo

Three domain checkpoints, same architecture — only the DiT weights differ (the VAE / text encoders are byte-identical across all three). Each is published as a one-line diffusers pipeline; a GitHub-code (DiT-only) copy lives in the same Hub repo used purely as file storage.

DomainCheckpoint
🚗 Autonomous Driving🤗 Kandinsky-WM-1.0-I2V-5s-AV
🤖 Robotics🤗 Kandinsky-WM-1.0-I2V-5s-RO
🌍 General Physics🤗 Kandinsky-WM-1.0-I2V-5s-PH

Examples

Four Kandinsky WM 1.0 image-to-video generations per domain (top-ranked by an internal visual review). Each grid shows the first frame of a clip — click it to play the generated video (hosted on the static_videos dataset). The first frame and prompt for every clip also live in assets/: under assets/<domain>/, frame_N.jpg and prompt_N.txt correspond to grid cell N (left→right, top→bottom).

🚗 Autonomous Driving

▶ AV clip 1 ▶ AV clip 2
▶ AV clip 3 ▶ AV clip 4

First frames + prompts: assets/av/.

🤖 Robotics

▶ Robotics clip 1 ▶ Robotics clip 2
▶ Robotics clip 3 ▶ Robotics clip 4

First frames + prompts: assets/robotics/.

🌍 General Physics

▶ General clip 1 ▶ General clip 2
▶ General clip 3 ▶ General clip 4

First frames + prompts: assets/general/.

Results

Kandinsky WM 1.0 on three physical-AI video benchmarks (our row in bold). RBench and Physics-IQ were run with Qwen3-VL prompt enhancers — noted above each table, with the scripts under prompt_enhancers/. Sizes are total parameter counts where publicly disclosed (MoE models note active params; proprietary/undisclosed left as —).

PAI-Bench-G (Physical AI Bench — Generation)

Leaderboard. Column abbreviations follow the PAI-Bench-G dimensions — see the leaderboard for exact definitions. Run without prompt enhancement.

RankModelSizeOverallDomainQualitySCBCMSAQIQOCISIBCSAVROINPHHU
1Cosmos3-Super64B83.989.578.292.794.199.252.770.820.597.798.194.477.590.090.795.087.6
2Cosmos3-Nano16B83.789.478.192.393.899.252.770.120.497.998.395.075.490.289.794.588.0
3Veo-3—82.186.777.691.493.199.251.969.821.797.096.994.468.786.989.791.684.4
4Kandinsky WM 1.02B81.786.077.491.894.199.053.264.921.697.297.793.672.682.188.090.786.0
5k5 Lite FT2B81.485.877.191.393.998.752.764.321.796.697.392.972.682.187.788.886.4
6Cosmos-Predict2.5-14B14B81.083.878.193.494.899.152.570.020.197.297.994.267.879.987.793.580.0
7Cosmos-Predict2.5-2B2B81.084.077.992.594.299.152.470.820.196.697.494.166.180.887.893.981.4
8Wan2.2-I2V-A14B27B (14B active)80.684.177.291.693.798.351.269.620.496.096.693.266.381.789.291.882.1
9K5 Lite2B80.583.077.991.794.499.354.165.621.798.198.689.266.377.386.387.284.6
10Wan2.2-TI2V-5B5B80.483.477.491.893.798.851.969.920.395.996.793.165.279.388.491.583.0

Physical-AI post-training lifts the base K5 Lite by +1.2 Overall (80.5 → 81.7) and +3.0 on the Domain axis (83.0 → 86.0) — a 2B model landing between Veo-3 and the Cosmos-Predict2.5 family.

RBench

Leaderboard, Qwen evaluator tab. RBench evaluates robot-oriented image-to-video generation across five task categories and four robot embodiments.

Prompt enhancer used: prompt_enhancers/rbench/enhance_rbench_prompts_qwen_v3.py — conservative Qwen3-VL image-grounded action canonicalization.

RankModelSizeAvg.Common ManipulationSpatial RelationshipMulti-entity CollaborationLong-horizon PlanningVisual ReasoningSingle ArmDual ArmQuadruped RobotHumanoid Robot
1Veo 3—0.7840.8970.7400.9240.8540.7500.7420.7080.7260.716
2Wan 2.5—0.7810.8880.8250.9160.7190.7380.7420.7440.7110.742
3Hailuo v2—0.7620.8430.8400.8920.7050.8200.6960.7060.6320.720
4Seedance 1.0—0.7550.8560.6650.8990.7160.7900.7080.7300.6740.759
5Wan2.2_A14B27B (14B active)0.6980.7090.6600.9210.6810.5500.6880.6780.6700.728
6Kandinsky-WM-1.02B0.6510.8360.6850.8460.6570.5240.5090.5680.5780.656
7Cosmos 2.514B0.6320.6880.5120.7680.5970.5070.6460.6470.6390.688
8LongCat-Video13.6B0.6090.6780.4650.8140.4900.3540.6980.6020.6660.710
9DreamGen(gr1)14B0.5750.5070.5000.8480.3530.4050.6600.6320.6110.656
10Wan2.2_5B5B0.5510.5980.4020.7220.4500.4200.5110.5340.6380.682
11Wan2.1_14B14B0.5420.6880.4000.7020.4650.2700.5190.5620.6040.664
12SkyReels13B0.5310.5460.4000.6870.3240.3580.6120.5740.6540.628
13DreamGen(droid)14B0.5140.4650.5050.5910.3020.3860.5890.5700.5840.633
14LTX-Video2B0.4860.4500.3820.7340.3580.2860.4870.4870.5880.603
15FramePack13B0.4530.4550.2400.6300.1970.3450.4030.5000.6700.635
16CogVideoX_5B5B0.3220.2900.2400.4260.0960.0300.3740.4220.4940.524
17Vidar—0.2070.1180.1400.0820.0190.0300.3440.3900.3800.364
18UnifoLM-WMA-0—0.1040.0290.0650.0250.0000.0000.2900.1060.2510.170

Physics-IQ Verified

Prompt enhancer used: prompt_enhancers/physics_iq/temporal_expansion.py — Qwen3-VL temporal caption expansion.

#ModelSizeInput typeScoreDate added
1Magi-1 24B + GeoPhys (BoN) (op)24Bmultiframe (v2v)58.2 ± 1.82026-06-19
2Magi-1 24B (op)24Bmultiframe (v2v)48.4 ± 1.12026-06-19
3Cosmos3-Super-Image2Video64Bi2v39.5 ± 0.82026-06-18
4Grok Imagine Video—i2v34.8 ± 0.62026-06-17
5Magi-1 24B + GeoPhys (BoN) (op)24Bi2v33.7 ± 1.42026-06-19
6Hunyuan Video 1.58.3Bi2v33.4 ± 0.82026-06-17
7Wan 2.227B (14B active)i2v32.2 ± 0.62026-06-17
8Kandinsky WM 1.02Bi2v30.8 ± 0.9internal
9Cosmos3-Nano16Bi2v30.3 ± 0.62026-06-18
10Magi-1 24B (op)24Bi2v30.2 ± 1.12026-06-19
11Sora 2—i2v26.5 ± 0.82026-06-17
12K5 Video Lite FT2Bi2v25.8 ± 1.7internal
13P-Video—i2v25.3 ± 1.82026-06-17
14K5 Video Lite2Bi2v16.0 ± 1.2internal

Quickstart

Two ways to run — the checkpoints ship in both formats. The GitHub-code path is primary; the diffusers path is a one-liner that already works off the Hub.

Path A — GitHub code

The base inference code is vendored as a git submodule at kandinsky-5/ (from kandinskylab/kandinsky-5).

# 0. get this repo with the submodule
git clone --recurse-submodules https://github.com/kandinskylab/kandinsky-wm.git
cd kandinsky-wm
# (already cloned without it? run: git submodule update --init --recursive)

# 1. install the base code
pip install -r kandinsky-5/requirements.txt

# 2. pull a checkpoint (repo-as-file-storage; DiT + shared VAE/text encoders)
python examples/download_checkpoint.py --domain robotics

# 3. generate (CLI) — run from inside the submodule
cd kandinsky-5
python test.py \
  --config ./configs/k5_lite_i2v_5s_sft_sd.yaml \
  --image  ../assets/robotics/frame_1.jpg \
  --prompt "$(cat ../assets/robotics/prompt_1.txt)" \
  --video_duration 5

Python variant: examples/run_i2v_github.py.

Path B — Diffusers

pip install -U diffusers transformers accelerate imageio-ffmpeg
import torch
from diffusers import Kandinsky5I2VPipeline
from diffusers.utils import export_to_video, load_image

pipe = Kandinsky5I2VPipeline.from_pretrained(
    "kandinskylab/Kandinsky-WM-1.0-I2V-5s-RO",
    torch_dtype=torch.bfloat16,
).to("cuda")

frames = pipe(
    image=load_image("assets/robotics/frame_1.jpg"),
    prompt=open("assets/robotics/prompt_1.txt").read().strip(),
    negative_prompt=("Static, 2D cartoon, cartoon, 2d animation, paintings, images, "
                     "worst quality, low quality, ugly, deformed, walking backwards"),
    height=512, width=768, num_frames=121,
    num_inference_steps=50, guidance_scale=5.0,
).frames[0]
export_to_video(frames, "robotics_out.mp4", fps=24, quality=9)

Script: examples/run_i2v_diffusers.py.

Repository layout

kandinsky-wm/
├── README.md
├── kandinsky-5/                    # git submodule → kandinskylab/kandinsky-5 (base inference code)
├── assets/                         # videos hosted on the static_videos HF dataset
│   ├── logo.png
│   ├── av/        frame_1..4.jpg · prompt_1..4.txt
│   ├── robotics/  frame_1..4.jpg · prompt_1..4.txt
│   └── general/   frame_1..4.jpg · prompt_1..4.txt
├── prompt_enhancers/               # Qwen3-VL benchmark prompt enhancers
│   ├── physics_iq/  temporal_expansion.py · grounding.py
│   └── rbench/      enhance_rbench_prompts_qwen_v3.py
└── examples/
    ├── download_checkpoint.py       # pull DiT + VAE + text encoders (download_models.py style)
    ├── run_i2v_github.py            # Path A (python)
    ├── run_i2v_github_cli.sh        # Path A (CLI)
    └── run_i2v_diffusers.py         # Path B (one-liner)

Acknowledgements

Built on Kandinsky 5.0 (HunyuanVideo VAE, Qwen2.5-VL, CLIP). Evaluated on PAI-Bench-G, Physics-IQ, and RBench.

diffusion-models
image-to-video
kandinsky
physical-ai
video-generation
world-model

kandinskylab/kandinsky-wm

Kandinsky WM 1.0 — a family of models for Physical AI. Image-to-video generation for autonomous driving, robotics & general physics.

Python

10

4 commits

updated Aug 5, 2026

See the code

README

Kandinsky WM

Kandinsky WM 1.0: A family of models for Physical AI

Image-to-Video generation for Physical AI: autonomous driving · robotics · general physics

🤗 Checkpoints collection · 🧩 Base model (Kandinsky 5.0) · 📝 Habr article


Contents


Overview

Kandinsky WM (World Model) 1.0 is a family of Image-to-Video models that adapt Kandinsky 5.0 Video Lite — a 2B-parameter latent video diffusion model (a Kandinsky5Transformer3DModel DiT paired with a HunyuanVideo VAE and Qwen2.5-VL + CLIP text encoders, trained with flow matching) — to Physical AI: video generation that is not only visually convincing but physically plausible — consistent scene geometry, object dynamics, interactions and cause-and-effect — so the model can serve as a source of synthetic training data and as a building block for world models and simulators.

The base model was domain-adapted on large corpora of 5-second scenes (first text-to-video, then a mixed text/image-to-video regime that preserves first-frame continuation), spanning autonomous driving, robotics, and a general domain (industrial processes, physical phenomena, and human–object / human–human interaction). Each checkpoint is then reinforcement-learning post-trained with GRPO against a reward model that scores the physical plausibility of the generated scene, steering the generator toward more faithful geometry, object behaviour and interactions than domain fine-tuning alone.

Each checkpoint generates 5-second, 121-frame clips at 768×512.

Model Zoo

Three domain checkpoints, same architecture — only the DiT weights differ (the VAE / text encoders are byte-identical across all three). Each is published as a one-line diffusers pipeline; a GitHub-code (DiT-only) copy lives in the same Hub repo used purely as file storage.

DomainCheckpoint
🚗 Autonomous Driving🤗 Kandinsky-WM-1.0-I2V-5s-AV
🤖 Robotics🤗 Kandinsky-WM-1.0-I2V-5s-RO
🌍 General Physics🤗 Kandinsky-WM-1.0-I2V-5s-PH

Examples

Four Kandinsky WM 1.0 image-to-video generations per domain (top-ranked by an internal visual review). Each grid shows the first frame of a clip — click it to play the generated video (hosted on the static_videos dataset). The first frame and prompt for every clip also live in assets/: under assets/<domain>/, frame_N.jpg and prompt_N.txt correspond to grid cell N (left→right, top→bottom).

🚗 Autonomous Driving

▶ AV clip 1 ▶ AV clip 2
▶ AV clip 3 ▶ AV clip 4

First frames + prompts: assets/av/.

🤖 Robotics

▶ Robotics clip 1 ▶ Robotics clip 2
▶ Robotics clip 3 ▶ Robotics clip 4

First frames + prompts: assets/robotics/.

🌍 General Physics

▶ General clip 1 ▶ General clip 2
▶ General clip 3 ▶ General clip 4

First frames + prompts: assets/general/.

Results

Kandinsky WM 1.0 on three physical-AI video benchmarks (our row in bold). RBench and Physics-IQ were run with Qwen3-VL prompt enhancers — noted above each table, with the scripts under prompt_enhancers/. Sizes are total parameter counts where publicly disclosed (MoE models note active params; proprietary/undisclosed left as —).

PAI-Bench-G (Physical AI Bench — Generation)

Leaderboard. Column abbreviations follow the PAI-Bench-G dimensions — see the leaderboard for exact definitions. Run without prompt enhancement.

RankModelSizeOverallDomainQualitySCBCMSAQIQOCISIBCSAVROINPHHU
1Cosmos3-Super64B83.989.578.292.794.199.252.770.820.597.798.194.477.590.090.795.087.6
2Cosmos3-Nano16B83.789.478.192.393.899.252.770.120.497.998.395.075.490.289.794.588.0
3Veo-3—82.186.777.691.493.199.251.969.821.797.096.994.468.786.989.791.684.4
4Kandinsky WM 1.02B81.786.077.491.894.199.053.264.921.697.297.793.672.682.188.090.786.0
5k5 Lite FT2B81.485.877.191.393.998.752.764.321.796.697.392.972.682.187.788.886.4
6Cosmos-Predict2.5-14B14B81.083.878.193.494.899.152.570.020.197.297.994.267.879.987.793.580.0
7Cosmos-Predict2.5-2B2B81.084.077.992.594.299.152.470.820.196.697.494.166.180.887.893.981.4
8Wan2.2-I2V-A14B27B (14B active)80.684.177.291.693.798.351.269.620.496.096.693.266.381.789.291.882.1
9K5 Lite2B80.583.077.991.794.499.354.165.621.798.198.689.266.377.386.387.284.6
10Wan2.2-TI2V-5B5B80.483.477.491.893.798.851.969.920.395.996.793.165.279.388.491.583.0

Physical-AI post-training lifts the base K5 Lite by +1.2 Overall (80.5 → 81.7) and +3.0 on the Domain axis (83.0 → 86.0) — a 2B model landing between Veo-3 and the Cosmos-Predict2.5 family.

RBench

Leaderboard, Qwen evaluator tab. RBench evaluates robot-oriented image-to-video generation across five task categories and four robot embodiments.

Prompt enhancer used: prompt_enhancers/rbench/enhance_rbench_prompts_qwen_v3.py — conservative Qwen3-VL image-grounded action canonicalization.

RankModelSizeAvg.Common ManipulationSpatial RelationshipMulti-entity CollaborationLong-horizon PlanningVisual ReasoningSingle ArmDual ArmQuadruped RobotHumanoid Robot
1Veo 3—0.7840.8970.7400.9240.8540.7500.7420.7080.7260.716
2Wan 2.5—0.7810.8880.8250.9160.7190.7380.7420.7440.7110.742
3Hailuo v2—0.7620.8430.8400.8920.7050.8200.6960.7060.6320.720
4Seedance 1.0—0.7550.8560.6650.8990.7160.7900.7080.7300.6740.759
5Wan2.2_A14B27B (14B active)0.6980.7090.6600.9210.6810.5500.6880.6780.6700.728
6Kandinsky-WM-1.02B0.6510.8360.6850.8460.6570.5240.5090.5680.5780.656
7Cosmos 2.514B0.6320.6880.5120.7680.5970.5070.6460.6470.6390.688
8LongCat-Video13.6B0.6090.6780.4650.8140.4900.3540.6980.6020.6660.710
9DreamGen(gr1)14B0.5750.5070.5000.8480.3530.4050.6600.6320.6110.656
10Wan2.2_5B5B0.5510.5980.4020.7220.4500.4200.5110.5340.6380.682
11Wan2.1_14B14B0.5420.6880.4000.7020.4650.2700.5190.5620.6040.664
12SkyReels13B0.5310.5460.4000.6870.3240.3580.6120.5740.6540.628
13DreamGen(droid)14B0.5140.4650.5050.5910.3020.3860.5890.5700.5840.633
14LTX-Video2B0.4860.4500.3820.7340.3580.2860.4870.4870.5880.603
15FramePack13B0.4530.4550.2400.6300.1970.3450.4030.5000.6700.635
16CogVideoX_5B5B0.3220.2900.2400.4260.0960.0300.3740.4220.4940.524
17Vidar—0.2070.1180.1400.0820.0190.0300.3440.3900.3800.364
18UnifoLM-WMA-0—0.1040.0290.0650.0250.0000.0000.2900.1060.2510.170

Physics-IQ Verified

Prompt enhancer used: prompt_enhancers/physics_iq/temporal_expansion.py — Qwen3-VL temporal caption expansion.

#ModelSizeInput typeScoreDate added
1Magi-1 24B + GeoPhys (BoN) (op)24Bmultiframe (v2v)58.2 ± 1.82026-06-19
2Magi-1 24B (op)24Bmultiframe (v2v)48.4 ± 1.12026-06-19
3Cosmos3-Super-Image2Video64Bi2v39.5 ± 0.82026-06-18
4Grok Imagine Video—i2v34.8 ± 0.62026-06-17
5Magi-1 24B + GeoPhys (BoN) (op)24Bi2v33.7 ± 1.42026-06-19
6Hunyuan Video 1.58.3Bi2v33.4 ± 0.82026-06-17
7Wan 2.227B (14B active)i2v32.2 ± 0.62026-06-17
8Kandinsky WM 1.02Bi2v30.8 ± 0.9internal
9Cosmos3-Nano16Bi2v30.3 ± 0.62026-06-18
10Magi-1 24B (op)24Bi2v30.2 ± 1.12026-06-19
11Sora 2—i2v26.5 ± 0.82026-06-17
12K5 Video Lite FT2Bi2v25.8 ± 1.7internal
13P-Video—i2v25.3 ± 1.82026-06-17
14K5 Video Lite2Bi2v16.0 ± 1.2internal

Quickstart

Two ways to run — the checkpoints ship in both formats. The GitHub-code path is primary; the diffusers path is a one-liner that already works off the Hub.

Path A — GitHub code

The base inference code is vendored as a git submodule at kandinsky-5/ (from kandinskylab/kandinsky-5).

# 0. get this repo with the submodule
git clone --recurse-submodules https://github.com/kandinskylab/kandinsky-wm.git
cd kandinsky-wm
# (already cloned without it? run: git submodule update --init --recursive)

# 1. install the base code
pip install -r kandinsky-5/requirements.txt

# 2. pull a checkpoint (repo-as-file-storage; DiT + shared VAE/text encoders)
python examples/download_checkpoint.py --domain robotics

# 3. generate (CLI) — run from inside the submodule
cd kandinsky-5
python test.py \
  --config ./configs/k5_lite_i2v_5s_sft_sd.yaml \
  --image  ../assets/robotics/frame_1.jpg \
  --prompt "$(cat ../assets/robotics/prompt_1.txt)" \
  --video_duration 5

Python variant: examples/run_i2v_github.py.

Path B — Diffusers

pip install -U diffusers transformers accelerate imageio-ffmpeg
import torch
from diffusers import Kandinsky5I2VPipeline
from diffusers.utils import export_to_video, load_image

pipe = Kandinsky5I2VPipeline.from_pretrained(
    "kandinskylab/Kandinsky-WM-1.0-I2V-5s-RO",
    torch_dtype=torch.bfloat16,
).to("cuda")

frames = pipe(
    image=load_image("assets/robotics/frame_1.jpg"),
    prompt=open("assets/robotics/prompt_1.txt").read().strip(),
    negative_prompt=("Static, 2D cartoon, cartoon, 2d animation, paintings, images, "
                     "worst quality, low quality, ugly, deformed, walking backwards"),
    height=512, width=768, num_frames=121,
    num_inference_steps=50, guidance_scale=5.0,
).frames[0]
export_to_video(frames, "robotics_out.mp4", fps=24, quality=9)

Script: examples/run_i2v_diffusers.py.

Repository layout

kandinsky-wm/
├── README.md
├── kandinsky-5/                    # git submodule → kandinskylab/kandinsky-5 (base inference code)
├── assets/                         # videos hosted on the static_videos HF dataset
│   ├── logo.png
│   ├── av/        frame_1..4.jpg · prompt_1..4.txt
│   ├── robotics/  frame_1..4.jpg · prompt_1..4.txt
│   └── general/   frame_1..4.jpg · prompt_1..4.txt
├── prompt_enhancers/               # Qwen3-VL benchmark prompt enhancers
│   ├── physics_iq/  temporal_expansion.py · grounding.py
│   └── rbench/      enhance_rbench_prompts_qwen_v3.py
└── examples/
    ├── download_checkpoint.py       # pull DiT + VAE + text encoders (download_models.py style)
    ├── run_i2v_github.py            # Path A (python)
    ├── run_i2v_github_cli.sh        # Path A (CLI)
    └── run_i2v_diffusers.py         # Path B (one-liner)

Acknowledgements

Built on Kandinsky 5.0 (HunyuanVideo VAE, Qwen2.5-VL, CLIP). Evaluated on PAI-Bench-G, Physics-IQ, and RBench.

diffusion-models
image-to-video
kandinsky
physical-ai
video-generation
world-model