Kandinsky Lab open-sourced Kandinsky 6.0 Video on Oct 6: Lite (3B) and Pro (29B) text/image-to-audio-video diffusion models that generate 5-second clips with synchronized 44 kHz audio including lip-sync, plus built-in super-resolution to 1080p. Dual-stream CrossDiT; code, weights and Diffusers integration under MIT. The report claims it beats open LTX 2.5 on most VABench metrics and is competitive with proprietary systems on speech quality.
Key Takeaways
- ✓Scale: Lite 3B / Pro 29B; Pro = 19B video stream + 5B audio stream + 5B cross-attention (report)
- ✓Output: 5 s clips with synchronized 44 kHz audio incl. lip-sync, T2AV and I2AV, built-in SR to 1920×1080
- ✓Evals: beats open LTX 2.5 on most VABench metrics; trades wins with Veo 3.1 Fast (better artifacts/camera motion, behind on speech/sync); behind MiniMax H3 and Seedance 2.0 on visuals
- ✓Speed: distilled to 10 NFE with 51% vs 49% human preference vs full model; non-distilled Pro Full HD ≈402 s on H100, ≈1247 s on RTX 4090
- ✓Deploy: block offload cuts Pro SD peak memory 72.8 → 21.7 GiB, 16 GB preset; Diffusers, ComfyUI and vLLM-Omni support; MIT license

Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
Background
Joint audio-video generation (dialogue, sound effects and lip-sync in one pass) has been dominated by closed systems such as Veo 3.1 and Sora 2. Kandinsky Lab's Kandinsky 6.0 Video open-sources the full stack under MIT: Lite (3B) and Pro (29B) models generating 5-second clips with synchronized 44 kHz audio in T2AV and I2AV modes, with a built-in super-resolution model up to 1920×1080.
Architecture
A dual-stream CrossDiT reuses the Kandinsky 5.0 video stream and adds a newly trained audio stream, connected blockwise by bidirectional cross-attention (3D RoPE for video, 1D RoPE for audio). Pro splits into 19B video + 5B audio + 5B cross-attention. Training: audio stream pretrained alone, then joint audio-video training, SFT, RL post-training and two-stage distillation to 10 function evaluations.
Benchmarks
Per the technical report: beats open LTX 2.5 on most VABench metrics; clearly beats Kandinsky 5.0 Pro in human SBS; trades wins with Veo 3.1 Fast (significantly better artifacts and camera motion, behind on speech and sync); behind MiniMax H3 and Seedance 2.0 on visuals while competitive on speech. Distilled vs full model preference is 51% vs 49%. All numbers are self-reported.
Getting started
NVIDIA GPU + Python 3.13/3.14, then just setup, just download pro-distill, just generate "a cat on a mat". Block offload cuts Pro SD peak memory from 72.8 to 21.7 GiB, with a 16 GB preset; non-distilled Pro Full HD takes ~402 s on H100. Try the HF Space, the ComfyUI node or the Diffusers collection.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.