We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).
22
33 commits
1 linked in READMEs
updated Oct 7, 2026
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).
This repository contains Kandinsky 6.0 Pro, distilled: a distilled Pro checkpoint that samples in 10 steps with the PiFlow scheduler.
import torch
from diffusers import Kandinsky6TI2VAPipeline
from diffusers.utils import encode_video
DEFAULT_PROMPT = (
"cinematic shot: a giant stone samurai on a stormy cliff above a neon city opens glowing golden eyes and "
"raises a katana. Blue lightning strikes the blade, creating a massive shockwave through the clouds. The "
"camera rapidly pulls back from a low angle. Photorealistic, epic scale, dark blue and gold lighting, rain, "
"sparks, volumetric lightning, blockbuster quality. Audio: heavy rain, deep thunder, metallic sword hum, "
"rising brass and choir, electrical crackle, perfectly synchronized lightning impact, sub-bass shockwave. "
"No dialogue, text, or logos."
)
pipe = Kandinsky6TI2VAPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
output = pipe(
prompt=DEFAULT_PROMPT,
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
Use 10 steps and guidance 1.0 with this checkpoint. The 50-step, guidance 5.0 settings of the other checkpoints don't apply.
Pass a reference image with image=. It conditions the first frame, and the pipeline resizes and center-crops it to height × width. The prompt should describe what happens next in that image, so the example uses its own prompt rather than DEFAULT_PROMPT. It reuses pipe from the T2AV example.
from diffusers.utils import load_image
IMAGE_PROMPT = (
"cinematic shot: a red dragon wakes on a mound of gold coins in a sunlit cave, slowly unfolds its wings "
"and lowers its head toward the glowing stream below. Coins slide and clink down the pile as the dragon "
"shifts its weight. The camera slowly pushes in toward the dragon's eyes. Photorealistic, fantasy, warm "
"golden light, drifting dust motes, lush cave flowers. Audio: deep rumbling breath, leathery wing flaps, "
"clinking coins, trickling water, distant echoing roar. No dialogue, text, or logos."
)
image = load_image(
"https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers/resolve/main/assets/i2va_input.jpg"
)
output = pipe(
prompt=IMAGE_PROMPT,
image=image,
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output_ti2av.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
Pass sample_audio=False to generate video only, and expand_prompts=True to let the Qwen2.5-VL text encoder rewrite a short prompt into a detailed one before encoding. expand_prompts adds latency but no extra model weights.
Kandinsky 6.0 video super-resolution is a separate pipeline, Kandinsky6SRPipeline. It takes the frames produced above and upscales them by ×2, ×2.25 or ×4 with tiled diffusion in K-VAE latent space, blending overlapping tiles back with Hann windows. The example uses the flow-matching checkpoint, which runs in 4 steps per tile.
import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import encode_video
torch._inductor.config.max_autotune = True # required: lets inductor pick flex-attention tiles that fit the SR block mask
sr_pipe = Kandinsky6SRPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", torch_dtype=torch.bfloat16
)
sr_pipe.transformer.set_attention_backend("flex")
sr_pipe.enable_model_cpu_offload()
sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)
upscaled = sr_pipe(
video=output.frames[0],
resolution_scale=2.25, # 2, 2.25 or 4
num_inference_steps=4, # 2 with Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
).frames[0]
encode_video(
upscaled,
fps=24,
output_path="output_sr.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
The example continues from the Usage guide: pipe and output come from it. The SR pipeline works on frames only, so the audio is passed to encode_video from the generation step.
resolution_scale=2.25 (the default, and the route used after Kandinsky 6 generation) is a ×1.125 bilinear pre-upscale followed by the ×2 path. 2 and 4 are the other options.min_overlap (default 0.2) sets the minimum tile overlap. tiles_batch_size (default 1) sets how many tiles go through the model per call, trading VRAM for speed.| Folder | Class | Role |
|---|---|---|
transformer/ | Kandinsky6Transformer3DModel | Multimodal video+audio diffusion transformer (Pro) |
vae/ | AutoencoderKLHunyuanVideo | Video VAE: encodes the reference image, decodes the generated video |
text_encoder/ + tokenizer/ | Qwen2_5_VLForConditionalGeneration + Qwen2_5_VLProcessor | Token-level text embeddings; also used for expand_prompts=True |
text_encoder_2/ + tokenizer_2/ | CLIPTextModel + CLIPTokenizer | Pooled text embedding |
audio_vae/ | MMAudioVAE | Audio latents ↔ mel spectrogram |
vocoder/ | MMAudioVocoder (BigVGAN-style) | Mel spectrogram → waveform |
scheduler/ | PiflowScheduler | PiFlow scheduler for the distilled checkpoint (shift = 5.0, 10 grid points) |
model_index.json names the pipeline class Kandinsky6TI2VAPipeline.
| Checkpoint | Repository | Use |
|---|---|---|
| Pro | kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers | generation |
| Pro, distilled, 10 steps (this repository) | kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers | generation |
| Pro, pretrained | kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers | generation |
| Lite | kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers | generation |
| Lite, distilled, 10 steps | kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers | generation |
| Lite, pretrained | kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers | generation |
| VSR, flow matching, 4 steps/tile | kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers | upscale |
| VSR, distilled, 2 steps/tile | kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers | upscale |
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).
22
33 commits
1 linked in READMEs
updated Oct 7, 2026
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).
This repository contains Kandinsky 6.0 Pro, distilled: a distilled Pro checkpoint that samples in 10 steps with the PiFlow scheduler.
import torch
from diffusers import Kandinsky6TI2VAPipeline
from diffusers.utils import encode_video
DEFAULT_PROMPT = (
"cinematic shot: a giant stone samurai on a stormy cliff above a neon city opens glowing golden eyes and "
"raises a katana. Blue lightning strikes the blade, creating a massive shockwave through the clouds. The "
"camera rapidly pulls back from a low angle. Photorealistic, epic scale, dark blue and gold lighting, rain, "
"sparks, volumetric lightning, blockbuster quality. Audio: heavy rain, deep thunder, metallic sword hum, "
"rising brass and choir, electrical crackle, perfectly synchronized lightning impact, sub-bass shockwave. "
"No dialogue, text, or logos."
)
pipe = Kandinsky6TI2VAPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
output = pipe(
prompt=DEFAULT_PROMPT,
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
Use 10 steps and guidance 1.0 with this checkpoint. The 50-step, guidance 5.0 settings of the other checkpoints don't apply.
Pass a reference image with image=. It conditions the first frame, and the pipeline resizes and center-crops it to height × width. The prompt should describe what happens next in that image, so the example uses its own prompt rather than DEFAULT_PROMPT. It reuses pipe from the T2AV example.
from diffusers.utils import load_image
IMAGE_PROMPT = (
"cinematic shot: a red dragon wakes on a mound of gold coins in a sunlit cave, slowly unfolds its wings "
"and lowers its head toward the glowing stream below. Coins slide and clink down the pile as the dragon "
"shifts its weight. The camera slowly pushes in toward the dragon's eyes. Photorealistic, fantasy, warm "
"golden light, drifting dust motes, lush cave flowers. Audio: deep rumbling breath, leathery wing flaps, "
"clinking coins, trickling water, distant echoing roar. No dialogue, text, or logos."
)
image = load_image(
"https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers/resolve/main/assets/i2va_input.jpg"
)
output = pipe(
prompt=IMAGE_PROMPT,
image=image,
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output_ti2av.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
Pass sample_audio=False to generate video only, and expand_prompts=True to let the Qwen2.5-VL text encoder rewrite a short prompt into a detailed one before encoding. expand_prompts adds latency but no extra model weights.
Kandinsky 6.0 video super-resolution is a separate pipeline, Kandinsky6SRPipeline. It takes the frames produced above and upscales them by ×2, ×2.25 or ×4 with tiled diffusion in K-VAE latent space, blending overlapping tiles back with Hann windows. The example uses the flow-matching checkpoint, which runs in 4 steps per tile.
import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import encode_video
torch._inductor.config.max_autotune = True # required: lets inductor pick flex-attention tiles that fit the SR block mask
sr_pipe = Kandinsky6SRPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", torch_dtype=torch.bfloat16
)
sr_pipe.transformer.set_attention_backend("flex")
sr_pipe.enable_model_cpu_offload()
sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)
upscaled = sr_pipe(
video=output.frames[0],
resolution_scale=2.25, # 2, 2.25 or 4
num_inference_steps=4, # 2 with Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
).frames[0]
encode_video(
upscaled,
fps=24,
output_path="output_sr.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
The example continues from the Usage guide: pipe and output come from it. The SR pipeline works on frames only, so the audio is passed to encode_video from the generation step.
resolution_scale=2.25 (the default, and the route used after Kandinsky 6 generation) is a ×1.125 bilinear pre-upscale followed by the ×2 path. 2 and 4 are the other options.min_overlap (default 0.2) sets the minimum tile overlap. tiles_batch_size (default 1) sets how many tiles go through the model per call, trading VRAM for speed.| Folder | Class | Role |
|---|---|---|
transformer/ | Kandinsky6Transformer3DModel | Multimodal video+audio diffusion transformer (Pro) |
vae/ | AutoencoderKLHunyuanVideo | Video VAE: encodes the reference image, decodes the generated video |
text_encoder/ + tokenizer/ | Qwen2_5_VLForConditionalGeneration + Qwen2_5_VLProcessor | Token-level text embeddings; also used for expand_prompts=True |
text_encoder_2/ + tokenizer_2/ | CLIPTextModel + CLIPTokenizer | Pooled text embedding |
audio_vae/ | MMAudioVAE | Audio latents ↔ mel spectrogram |
vocoder/ | MMAudioVocoder (BigVGAN-style) | Mel spectrogram → waveform |
scheduler/ | PiflowScheduler | PiFlow scheduler for the distilled checkpoint (shift = 5.0, 10 grid points) |
model_index.json names the pipeline class Kandinsky6TI2VAPipeline.
| Checkpoint | Repository | Use |
|---|---|---|
| Pro | kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers | generation |
| Pro, distilled, 10 steps (this repository) | kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers | generation |
| Pro, pretrained | kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers | generation |
| Lite | kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers | generation |
| Lite, distilled, 10 steps | kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers | generation |
| Lite, pretrained | kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers | generation |
| VSR, flow matching, 4 steps/tile | kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers | upscale |
| VSR, distilled, 2 steps/tile | kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers | upscale |