Research
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Overview Research area: Computer vision and generative multimedia, specifically diffusion transformers for joint (synchronized) video and audio generation. Technical level: Advanced. The paper assumes

- arXiv
- 2610.05608
- Published
- 2026-10-04
- Authors
- Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin, Grigorii Alekseenko, Anastasia Aliaskina, Olga Androsova, Vladimir Arkhipkin, Anna Averchenkova, Alexander Belykh, Serafima Bocharova, Sofiya Bogakovskaya, Anton Bukashkin, Mark Bulygin, Kirill Buzygin, Irina Cheremnykh, Kirill Chernyshev, Mikhail Chernyshov, Vladimir Chernyy, David Chikovani, Georgy Daniltsev, Denis Dimitrov, Anna Dmitrienko, Vladimir Dokholyan, Sergey Emelyanov, Dmitry Ermilov, Georgii Fedorov, Polina Gavrilova, Nikolai Gerasimenko, Aleksandr Gordeev, Andrey Inozemtsev, Andrei Ivaniuta, Alexander Ivanov, Mikhail Karaev, Anastasiia Kargapoltseva, Ivan Kirillov, Nikita Kiselev, Valeria Kobenko, Yury Kolabushin, Denis Koposov, Anatoly Korobov, Vladimir Korviakov, Kirill Kozlov, Denis Krzhivokolskiy, Konstantin Kuklev, Alexander Kunitsyn, Sergey Kuzin, Vladislav Lakhtionov, Alexey Letunovskiy, Maxim Litvinov, Alexander Lyulkov, Georgy Makarov, Kirill Malakhov, Egor Malykh, Mikhail Mamaev, Dmitrii Mikhailov, Polina Mikhailova, Ivan Mikheev, Elizaveta Muromtseva, Nikolai Nazarkin, Tatiana Nikulina, Lev Novitskiy, Stanislav Onuchin, Nikita Osterov, Denis Parkhomenko, Anatoliy Parpara, Vladimir Polovnikov, Konstantin Reznikov, Azat Saginbaev, Nikita Samsonov, Alexander Sentsov, Nikita Shaimov, Artem Sherstyuk, Andrey Shutkin, Egor Silvestrov, Bulat Suleimanov, Matvey Suprunov, Sergey Taranov, Irina Tolstykh, Tatiana Trofimuk, Ilya Trushkin, Aleksandra Tsybina, Olga Varlashina, Viacheslav Vasilev, Ilya Vasiliev, Eugeny Vilisov, Sergey Yakubson, Konstantin Zakharov
AI summary
Overview
- Research area: Computer vision and generative multimedia, specifically diffusion transformers for joint (synchronized) video and audio generation.
- Technical level: Advanced. The paper assumes familiarity with diffusion and flow-matching models, DiT/CrossDiT backbones, VAEs, RoPE, cross-attention, reinforcement learning post-training, and distillation.
- Scope (one sentence): The paper introduces and describes Kandinsky 6.0 Video Lite (3B parameters) and Pro (29B parameters), two foundation diffusion models that generate 5-second video clips with synchronized 44 kHz audio (including lip-sync) in text-to-audio-video and image-to-audio-video modes, and are released open-source under the MIT license.
What This Paper Is About
Generating video with matching, synchronized sound is considerably harder than generating either modality alone, because the two streams must agree semantically, temporally, and emotionally. Closed commercial systems (Veo 3.1, Sora 2, Wan 2.6) produce strong results but do not release source code, which limits reproducibility. Kandinsky 6.0 Video addresses this gap by building an open family of models that jointly generates video and 44 kHz audio, extending the earlier Kandinsky 5.0 video generator with a newly trained audio stream connected through bidirectional cross-attention.
Key Contributions
- Two models at different scales. Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both generate 5-second clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes, in SD resolution with a built-in super-resolution model that upscales to Full-HD.
- A continuous pretraining strategy. The audio stream is first pretrained from scratch on large-scale audio corpora, then connected to the already-pretrained video stream through newly initialized cross-attention layers, and finally both streams are trained jointly on paired audio-video data.
- A multi-stage post-training pipeline. Domain-specific supervised fine-tuning combined via model soup, reinforcement learning adapted from OmniNFT, and two-stage distillation (π-Flow followed by adversarial refinement) down to 10 function evaluations.
- An open release plus evaluation. Weights, source code, and diffusers integration are released under the MIT license, with evaluation on VABench and in side-by-side human study.
Main Findings
- Two-scale family with shared architecture: Both versions share the same architecture and are trained with flow matching; they differ in block counts, hidden dimensions, batch size, step counts, and learning rates. Pro uses 60 visual and 60 audio blocks versus 32 for Lite, with video_model_dim 4096 versus 1792 and audio_model_dim 2048 versus 896.
- Parameter distribution: In Lite (3B total), the video stream has 2B parameters, the audio stream 0.6B, and the cross-attention layers 0.4B. In Pro (29B total), the video stream has 19B, the audio stream 5B, and the cross-attention layers 5B.
- Human evaluation result: In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor Kandinsky 5.0 Video Pro and remains competitive with leading audio-video generation models, particularly in speech quality.
- VABench result: It outperforms the open-source LTX 2.5 on most VABench metrics. (Aggregate numerical VABench scores are not present in the available text.)
- Reference-frame conditioning works mainly through RoPE and masking, not token role embeddings: Inference-time ablations showed that the Token Role Embeddings have a relatively weak effect; setting the channel mask to one for all tokens almost completely blocked denoising, while an all-zero mask still allowed denoising but produced artifacts, mainly in the first latent frame. Changing the temporal RoPE position of the reference frame made the model use it at that position.
- Super-resolution details: The SR-DiT is text-free, initialized from a text-free Kandinsky 5.0 Video Lite SFT checkpoint, has 32 transformer blocks, hidden dimension 1792, 28 attention heads of dimension 64, FFN dimension 7168, and approximately 1.41B backbone parameters. The ×4 Latent Upscaler has approximately 1.44B parameters (about 79% in the five input-resolution ResBlocks); the ×2 branch has approximately 0.763B parameters.
- Data scale: Pretraining used a 50-million-image T2I subset of the Kandinsky 5.0 collection, 20 million video scenes, 40 million audio tracks (5–20 seconds), 7 million audio-video segments, and 116,000 I2AV samples. The Russian Cultural Code dataset contains 1.2M images and 530K video scenes.
- Lip-sync filtering rule: Clips with detected people were retained only if their Lip-Sync alignment score was at least 3 on a scale from 0 to 5; clips with audio-visual events were retained only if Desync was below 0.1 on a scale from 0 to 1. About 25% of the T2AV dataset has no visible faces.
Methodology in Plain English
The model is a dual-stream transformer. One stream handles video using a pretrained Kandinsky 5.0 backbone, the other handles audio; the two are wired together with bidirectional cross-attention blocks so that video tokens can look at audio tokens and vice versa. Video is compressed with the Hunyuan VAE, audio with the MMAudio autoencoder at 44 kHz, and text is encoded with Qwen2.5-VL-7B-Instruct plus CLIP ViT-L/14. Positions are encoded with RoPE: video tokens use the frame index normalized against 24 fps, and audio tokens use their temporal position.
Training proceeds in ordered phases. First the audio branch is learned from scratch on audio-only data while the video branch continues text-to-video training on audio-video captions. Then both streams are fused and trained jointly on paired data, with the diffusion time sampled independently for each modality so one modality can condition on the other. A quarter of the final pretraining samples are in image-to-audio-video mode, where the reference frame latent is concatenated along the temporal axis, marked by a binary channel mask, and given its own RoPE position.
After pretraining comes supervised fine-tuning in two stages across an 11-domain taxonomy, where a separate model is trained per domain and the domain models are merged by uniform weight averaging into a model soup. Reinforcement learning adapted from OmniNFT follows, and finally two-stage distillation (π-Flow, then adversarial refinement) compresses sampling to 10 function evaluations. Separately, a super-resolution pipeline takes the generated video and upscales it: a convolutional Latent Upscaler resizes in latent space, and an SR-DiT diffusion transformer removes degradations, supporting ×2 and ×4 routes, with a ×2.25 route for Full-HD output that combines bilinear ×1.125 pixel-space pre-upscaling with the ×2 route.
Why This Matters
Impact on research. Most strong synchronized audio-video generators are closed. Releasing weights, code, and diffusers integration under the MIT license gives the community a reproducible, inspectable baseline for T2AV and I2AV research, and the paper documents its data filtering, training schedule, and architectural ablations in enough detail to be built upon.
Real-world applications.
- Rapid prototyping of short audiovisual content with speech, music, and environmental sound, including lip-sync, from a text prompt or a single reference image.
- Dubbing and localization work, given the model's reported strength in speech quality and its bilingual (English and Russian) captioning pipeline.
- Previsualization for film, advertising, or game storyboards, where a 5-second synchronized clip can be upscaled to Full-HD.
- Culturally aware media generation, supported by the expanded Russian Cultural Code dataset (1.2M images, 530K video scenes).
Industry relevance. The paper's comparison targets include closed systems (Veo 3.1, Sora 2, Wan 2.6) and open ones (LTX 2.5, HunyuanVideo, Wan, Kandinsky 5.0), so it positions an open release directly against proprietary products. The engineering choices—sparse NABLA attention, latent-space super-resolution to avoid re-encoding each tile, and distillation to 10 function evaluations—are explicitly aimed at deployment cost, not only quality.
Future Directions
- Replace inference-time ablations with training-time tests. The authors note that inference-time ablations cannot fully determine how useful the Token Role Embeddings are during training, which would require a separate training run without them.
- Extend beyond 5 seconds. Both models generate 5-second clips (up to 121 frames at 23–25 fps); longer temporal coherence is not addressed in the available text.
- Close the gap with proprietary systems. The paper reports being competitive with proprietary systems and ahead of open LTX 2.5 on most VABench metrics, but states that open-source solutions still lag closed ones in clip duration, resolution, and synchronization accuracy.
- Broaden language and cultural coverage. Captions are produced in Russian and English, and the Russian Cultural Code dataset is specifically curated; other languages and cultural contexts are not covered in the available text.
Target Audience
Generative-AI and computer-vision researchers working on diffusion transformers, joint audio-video synthesis, or multimodal alignment; engineers building media-generation products who need an open, MIT-licensed model with published checkpoints and diffusers support; and practitioners interested in super-resolution, distillation, or reinforcement-learning post-training recipes for large generative models. Readers without a background in diffusion models, transformer architectures, and latent autoencoders will find the architecture and training sections difficult.
Authors’ abstract
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920$\times$1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.