KandinskyLab GitHub Report PyPI HF Demo Diffusers

ComfyUI

Kandinsky 6.0: A family of diffusion models for Video + Audio generation

We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).

The complete Kandinsky 6.0 video super-resolution pipeline as a single 🧨 Diffusers bundle: the 1.4B SR DiT distilled with π-Flow DX down to 2 model evaluations per tile, the KVAE video VAE and the x2 / x4 latent-upscaler bank. It upscales a video by x2, x2.25 or x4 with tiled diffusion in KVAE latent space. The Diffusers pipeline is video-only; mux the source audio back in if you need it in the output.

This repository contains all components needed for inference with from_pretrained. For the flow-matching teacher with 4 Euler steps per tile, see the flow-matching Diffusers bundle.

Usage with Diffusers

import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import export_to_video, load_video

torch._inductor.config.max_autotune = True  # required: lets inductor pick flex-attention tiles that fit the SR block mask

pipe = Kandinsky6SRPipeline.from_pretrained("kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()
pipe.transformer.set_attention_backend("flex")
pipe.transformer.compile_repeated_blocks(fullgraph=True)

video = load_video("lq.mp4")  # -> list[PIL.Image.Image]
output = pipe(
    video=video,
    resolution_scale=2.25,  # 2, 2.25 or 4
    num_inference_steps=2,  # must be set explicitly -- the pipeline's own default (4) is for the flow-matching checkpoint
    generator=torch.Generator("cuda").manual_seed(1137),
)
export_to_video(output.frames[0], "lq_sr.mp4", fps=24)

Video Super Resolution Models

Component Repository
Diffusers bundle, π-Flow distilled (2 evaluations per tile) (this repository) Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
Diffusers bundle, flow matching Kandinsky-6.0-VSR-5s-Diffusers

Components

Folder Class Parameters Contents
transformer/ Kandinsky6SRTransformer3DModel 1.41B π-Flow DX distilled SR DiT (2 evaluations per tile)
vae/ Kandinsky6SRVAE 1.74B Video VAE (included in this bundle)
latent_upscaler/ Kandinsky6SRLatentUpscalerBank (x2 + x4) 3.65B x2 and x4 latent upscalers (included in this bundle)
scheduler/ PiflowScheduler — π-Flow DX policy rollout: nfe = 2, shift = 3.5, 10 grid points, 128 policy substeps

model_index.json names the pipeline class Kandinsky6SRPipeline.

Downloads last month
180
Safetensors
Model size
1B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers

Finetuned
(1)
this model

Space using kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers 1

Collection including kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers

Paper for kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers