Instructions to use kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
Kandinsky 6.0: A family of diffusion models for Video + Audio generation
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (TI2AV) modes; a plug-in super-resolution model raises the output resolution to Full-HD (1920×1080).
The complete Kandinsky 6.0 video super-resolution pipeline as a single 🧨 Diffusers bundle: the 1.4B SR DiT distilled with π-Flow DX down to 2 model evaluations per tile, the KVAE video VAE and the x2 / x4 latent-upscaler bank. It upscales a video by x2, x2.25 or x4 with tiled diffusion in KVAE latent space. The Diffusers pipeline is video-only; mux the source audio back in if you need it in the output.
This repository contains all components needed for inference with from_pretrained. For the flow-matching teacher with 4 Euler steps per tile, see the flow-matching Diffusers bundle.
Usage with Diffusers
import torch
from diffusers import Kandinsky6SRPipeline
from diffusers.utils import export_to_video, load_video
torch._inductor.config.max_autotune = True # required: lets inductor pick flex-attention tiles that fit the SR block mask
pipe = Kandinsky6SRPipeline.from_pretrained("kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", torch_dtype=torch.bfloat16)
pipe.enable_model_cpu_offload()
pipe.transformer.set_attention_backend("flex")
pipe.transformer.compile_repeated_blocks(fullgraph=True)
video = load_video("lq.mp4") # -> list[PIL.Image.Image]
output = pipe(
video=video,
resolution_scale=2.25, # 2, 2.25 or 4
num_inference_steps=2, # must be set explicitly -- the pipeline's own default (4) is for the flow-matching checkpoint
generator=torch.Generator("cuda").manual_seed(1137),
)
export_to_video(output.frames[0], "lq_sr.mp4", fps=24)
Video Super Resolution Models
| Component | Repository |
|---|---|
| Diffusers bundle, π-Flow distilled (2 evaluations per tile) (this repository) | Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers |
| Diffusers bundle, flow matching | Kandinsky-6.0-VSR-5s-Diffusers |
Components
| Folder | Class | Parameters | Contents |
|---|---|---|---|
transformer/ |
Kandinsky6SRTransformer3DModel |
1.41B | π-Flow DX distilled SR DiT (2 evaluations per tile) |
vae/ |
Kandinsky6SRVAE |
1.74B | Video VAE (included in this bundle) |
latent_upscaler/ |
Kandinsky6SRLatentUpscalerBank (x2 + x4) |
3.65B | x2 and x4 latent upscalers (included in this bundle) |
scheduler/ |
PiflowScheduler |
— | π-Flow DX policy rollout: nfe = 2, shift = 3.5, 10 grid points, 128 policy substeps |
model_index.json names the pipeline class Kandinsky6SRPipeline.
- Downloads last month
- 180
Model tree for kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers
Base model
kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers