kandinskylab/kandinsky-6

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Python

165

22 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

Kandinsky 6.0 video gen + video upscaler! (r/LocalLLaMA)

Kandinsky 6.0 Pro (29B) and Lite (3B). Already support for ComfyUI and Diffusers. https://github.com/kandinskylab/kandinsky-6

2

Oct 7, 2026

README



KandinskyLab Report Diffusers HF Demo ComfyUI vLL-Omni SGLang FastVideo

Kandinsky 6.0: A family of diffusion models for Video + Audio generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes. A plug-in super-resolution model raises the output resolution to Full HD (1920×1080).

Project Updates

Quick start

An NVIDIA GPU and Python 3.13 or 3.14.

The first run downloads Kandinsky-6.0-Pro-distill-5s into $KANDINSKY_HOME/weights. When $KANDINSKY_HOME is unset, that directory is ~/.cache/kandinsky. The clip is written to $KANDINSKY_HOME/outputs/generate_<YYYY-MM-DDTHH-MM-SS>/generations/output.mp4, with the expanded prompt in expanded_prompt.txt next to launch.json, logs/, and profiles/. The generate command downloads the checkpoint named in the config if it is not already on disk. Other catalog names are pro, pro-pretrain, lite, lite-distill, and lite-pretrain.

Clone the repository

You need uv and just.

git clone https://github.com/kandinskylab/kandinsky-6.git
cd k6_video
just setup

just setup reads the GPU and installs one PyTorch build:

GPUsCompute capabilityInstalled
Hopper (H100, H200)9.0PyTorch for CUDA 13.0, FlashAttention 3
Ampere, Ada, consumer Blackwell8.0, 8.6, 8.9, 12.0, 12.1PyTorch for CUDA 13.0, then SageAttention 2.2.0

Ampere, Ada, and consumer Blackwell compile SageAttention, so nvcc has to be on PATH. CUDA_HOME defaults to /usr/local/cuda.

just download pro-distill
just generate "a cat on a mat"

Another GPU uses its preset:

just generate "a cat on a mat" --config kandinsky/configs/devices/rtx-5090.yaml --out clip.mp4

Presets live in kandinsky/configs/devices/.

ComfyUI

For ComfyUI, install kandinsky6 and kandinsky6-sr through ComfyUI Manager, then restart ComfyUI. The comfyui/ directory contains extension source code; no manual copying is needed — see the setup guide and manual model downloads.

vLLM-Omni

Kandinsky 6 is on vLLM-Omni main. Install vLLM 0.31.0 and current vLLM-Omni from source. Use a separate environment from just setup.

uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm==0.31.0 --torch-backend=auto \
  --extra-index-url https://wheels.vllm.ai/db9527a46873454610df6dbedf79a36d6bf1a7f6
git clone https://github.com/vllm-project/vllm-omni.git
cd vllm-omni
uv pip install -e .

Pro defaults are 864×480, 125 frames at 24 fps, 50 steps, and guidance 5.0. Audio is on. On an 80 GB GPU, --enable-cpu-offload is required.

Text-to-video-and-audio:

python examples/offline_inference/text_to_video/text_to_video.py \
  --model kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers \
  --prompt "A golden retriever runs along a sunny beach, waves crashing, cinematic footage" \
  --seed 42 \
  --enable-cpu-offload \
  --output kandinsky6_t2va.mp4

Image-to-video-and-audio, with the image used as a masked tail frame:

python examples/offline_inference/image_to_video/image_to_video.py \
  --model kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers \
  --image first_frame.png \
  --prompt "A golden retriever runs along a sunny beach, waves crashing, cinematic footage" \
  --seed 42 \
  --enable-cpu-offload \
  --output kandinsky6_i2va.mp4

Online serving:

vllm serve kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers --omni \
  --host 127.0.0.1 --port 8091 \
  --num-gpus 1 --enable-cpu-offload
VID=$(curl -s -X POST http://127.0.0.1:8091/v1/videos \
  -F prompt="A golden retriever runs along a sunny beach, waves crashing" \
  -F size=864x480 -F num_frames=125 -F num_inference_steps=50 -F seed=42 \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['id'])")
curl -s http://127.0.0.1:8091/v1/videos/$VID
curl -s -o kandinsky6.mp4 http://127.0.0.1:8091/v1/videos/$VID/content

The checkpoint is kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers. The API reference is the Kandinsky 6 module page.

Performance

Working time (s) for a 5-second clip on the non-distilled model, after warmup. Weight loading and MP4 encoding are excluded.

RTX PRO 6000 (96 GB), A100 80 GB, and H100 80 GB use module offload. The consumer cards use block offload.

RTX 4090, RTX 5060 Ti, RTX 5080, RTX 5090, RTX PRO 6000, and A100 80 GB use SageAttention2++ for visual self-attention. The H100 uses FlashAttention 3.

RTX 4090RTX 5060 TiRTX 5080RTX 5090RTX PRO 6000A100 80 GBH100
Lite SD4371310577309242532239
Lite HD4221328579296274472203
Lite Full HD5781774770406387664284
Pro SD93630801336754621972356
Pro HD119427161189638561791292
Pro Full HD1247353015468547651106402

kandinskylab/kandinsky-6

Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation

Python

165

22 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

Kandinsky 6.0 video gen + video upscaler! (r/LocalLLaMA)

Kandinsky 6.0 Pro (29B) and Lite (3B). Already support for ComfyUI and Diffusers. https://github.com/kandinskylab/kandinsky-6

2

Oct 7, 2026

README



KandinskyLab Report Diffusers HF Demo ComfyUI vLL-Omni SGLang FastVideo

Kandinsky 6.0: A family of diffusion models for Video + Audio generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes. A plug-in super-resolution model raises the output resolution to Full HD (1920×1080).

Project Updates

Quick start

An NVIDIA GPU and Python 3.13 or 3.14.

The first run downloads Kandinsky-6.0-Pro-distill-5s into $KANDINSKY_HOME/weights. When $KANDINSKY_HOME is unset, that directory is ~/.cache/kandinsky. The clip is written to $KANDINSKY_HOME/outputs/generate_<YYYY-MM-DDTHH-MM-SS>/generations/output.mp4, with the expanded prompt in expanded_prompt.txt next to launch.json, logs/, and profiles/. The generate command downloads the checkpoint named in the config if it is not already on disk. Other catalog names are pro, pro-pretrain, lite, lite-distill, and lite-pretrain.

Clone the repository

You need uv and just.

git clone https://github.com/kandinskylab/kandinsky-6.git
cd k6_video
just setup

just setup reads the GPU and installs one PyTorch build:

GPUsCompute capabilityInstalled
Hopper (H100, H200)9.0PyTorch for CUDA 13.0, FlashAttention 3
Ampere, Ada, consumer Blackwell8.0, 8.6, 8.9, 12.0, 12.1PyTorch for CUDA 13.0, then SageAttention 2.2.0

Ampere, Ada, and consumer Blackwell compile SageAttention, so nvcc has to be on PATH. CUDA_HOME defaults to /usr/local/cuda.

just download pro-distill
just generate "a cat on a mat"

Another GPU uses its preset:

just generate "a cat on a mat" --config kandinsky/configs/devices/rtx-5090.yaml --out clip.mp4

Presets live in kandinsky/configs/devices/.

ComfyUI

For ComfyUI, install kandinsky6 and kandinsky6-sr through ComfyUI Manager, then restart ComfyUI. The comfyui/ directory contains extension source code; no manual copying is needed — see the setup guide and manual model downloads.

vLLM-Omni

Kandinsky 6 is on vLLM-Omni main. Install vLLM 0.31.0 and current vLLM-Omni from source. Use a separate environment from just setup.

uv venv --python 3.12
source .venv/bin/activate
uv pip install vllm==0.31.0 --torch-backend=auto \
  --extra-index-url https://wheels.vllm.ai/db9527a46873454610df6dbedf79a36d6bf1a7f6
git clone https://github.com/vllm-project/vllm-omni.git
cd vllm-omni
uv pip install -e .

Pro defaults are 864×480, 125 frames at 24 fps, 50 steps, and guidance 5.0. Audio is on. On an 80 GB GPU, --enable-cpu-offload is required.

Text-to-video-and-audio:

python examples/offline_inference/text_to_video/text_to_video.py \
  --model kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers \
  --prompt "A golden retriever runs along a sunny beach, waves crashing, cinematic footage" \
  --seed 42 \
  --enable-cpu-offload \
  --output kandinsky6_t2va.mp4

Image-to-video-and-audio, with the image used as a masked tail frame:

python examples/offline_inference/image_to_video/image_to_video.py \
  --model kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers \
  --image first_frame.png \
  --prompt "A golden retriever runs along a sunny beach, waves crashing, cinematic footage" \
  --seed 42 \
  --enable-cpu-offload \
  --output kandinsky6_i2va.mp4

Online serving:

vllm serve kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers --omni \
  --host 127.0.0.1 --port 8091 \
  --num-gpus 1 --enable-cpu-offload
VID=$(curl -s -X POST http://127.0.0.1:8091/v1/videos \
  -F prompt="A golden retriever runs along a sunny beach, waves crashing" \
  -F size=864x480 -F num_frames=125 -F num_inference_steps=50 -F seed=42 \
  | python3 -c "import sys,json; print(json.load(sys.stdin)['id'])")
curl -s http://127.0.0.1:8091/v1/videos/$VID
curl -s -o kandinsky6.mp4 http://127.0.0.1:8091/v1/videos/$VID/content

The checkpoint is kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers. The API reference is the Kandinsky 6 module page.

Performance

Working time (s) for a 5-second clip on the non-distilled model, after warmup. Weight loading and MP4 encoding are excluded.

RTX PRO 6000 (96 GB), A100 80 GB, and H100 80 GB use module offload. The consumer cards use block offload.

RTX 4090, RTX 5060 Ti, RTX 5080, RTX 5090, RTX PRO 6000, and A100 80 GB use SageAttention2++ for visual self-attention. The H100 uses FlashAttention 3.

RTX 4090RTX 5060 TiRTX 5080RTX 5090RTX PRO 6000A100 80 GBH100
Lite SD4371310577309242532239
Lite HD4221328579296274472203
Lite Full HD5781774770406387664284
Pro SD93630801336754621972356
Pro HD119427161189638561791292
Pro Full HD1247353015468547651106402