Skip to content
AI.info

Research

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation

Overview Research area: Generative computer vision — large-scale text-to-image and text-to-video synthesis. Technical level: Advanced. The report assumes familiarity with diffusion/flow-matching model

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation
arXiv
2511.14993
Published
2025-11-19
Authors
Vladimir Arkhipkin, Vladimir Korviakov, Nikolai Gerasimenko, Denis Parkhomenko, Viacheslav Vasilev, Alexey Letunovskiy, Nikolai Vaulin, Maria Kovaleva, Ivan Kirillov, Lev Novitskiy, Denis Koposov, Nikita Kiselev, Alexander Varlamov, Dmitrii Mikhailov, Vladimir Polovnikov, Andrey Shutkin, Julia Agafonova, Ilya Vasiliev, Anastasiia Kargapoltseva, Anna Dmitrienko, Anastasia Maltseva, Anna Averchenkova, Olga Kim, Tatiana Nikulina, Denis Dimitrov

AI summary

Overview

Research area: Generative computer vision — large-scale text-to-image and text-to-video synthesis.

Technical level: Advanced. The report assumes familiarity with diffusion/flow-matching models, Diffusion Transformers (DiT), attention mechanisms, distillation, and distributed training.

Scope in one sentence: This technical report describes the Kandinsky 5.0 family — a 6B-parameter image model and 2B and 19B parameter video models for high-resolution image generation and up to 10-second video generation — covering its data pipeline, architecture, multi-stage training, optimizations, and human evaluation.

What This Paper Is About

Building high-quality video generation systems is hard: the data must be filtered at massive scale, and the computational cost of attention grows sharply with resolution and clip length. Kandinsky 5.0 addresses both problems by pairing a curated multi-hundred-million-item data pipeline with a Cross-Attention Diffusion Transformer (CrossDiT) and a sparse attention method called NABLA. The goal is a publicly released family of models that reaches state-of-the-art quality while remaining fast enough to train and run in practice.

Key Contributions

  1. A documented data curation lifecycle. The report describes collection, processing, filtering, deduplication, classification, captioning, and clustering for text-to-image, text-to-video, image-editing instruction, Russian-culture-specific, and supervised fine-tuning (SFT) datasets, including data preparation for instructive image editing tuning and SFT for both video and image modalities.

  2. A multi-stage training pipeline for all six models. The pipeline covers pretraining to learn general patterns of the visual world, quality-enhancement SFT, and an RLHF post-training adversarial method that compares generated images against images from the SFT dataset, reported to achieve superior realism, visual quality, and prompt alignment.

  3. The CrossDiT architecture and the NABLA attention method. NABLA targets high-resolution video (exceeding 512×512 px) with durations longer than 5 seconds, overcoming the quadratic complexity of standard spatio-temporal attention with a 2.7 times reduction in training and inference time at a 90% sparsity ratio.

  4. System-level optimizations plus distillation and open release. These include VAE optimization, text encoder quantization, Fully or Hybrid Sharded Data Parallel (F/HSDP) training, activation checkpointing, and a video distillation recipe that reduces Number of Function Evaluations (NFE) from 100 to 16. Code and weights from various training stages are released through the diffusers library.

Main Findings

  • The model family has three line-ups. Kandinsky 5.0 Image Lite is a 6B-parameter line-up for high-resolution text-to-image generation and image editing; Kandinsky 5.0 Video Lite is a 2B-parameter model for text-to-video and image-to-video; Kandinsky 5.0 Video Pro is a 19B-parameter model for superior video generation quality. Video models produce up to 10-second clips.

  • NABLA delivers large speedups with preserved quality. The method achieves a 2.7 times reduction in training and inference time with a 90% sparsity ratio on high-resolution, longer-than-5-second video, with quality confirmed by FVD, VBench, and CLIP-score, and by side-by-side human evaluation.

  • Distillation cuts inference cost dramatically. Combining Classifier-Free Guidance Distillation, Trajectory Segmented Consistency Distillation (TSCD), and subsequent adversarial post-training reduces NFE from 100 to 16 while preserving visual quality, as evidenced by side-by-side human evaluation.

  • Human evaluation favors Kandinsky 5.0 video. The authors report superior video generation quality in human evaluation on a prompt set from MovieGen; the specific win rates and per-model comparisons are not reported in the available content.

  • The video datasets are very large. Kandinsky T2V comprises more than 250 million video scenes; scenes are segmented with PySceneDetect at durations between 2 and 60 seconds, filtered by resolution (shorter side under 256 pixels removed), percentual-hash deduplication, MS-SSIM structural dynamics sampled at 2 FPS, DOVER and Q-Align quality scoring, CRAFT text detection averaged over three frames, YOLOv8 object detection over five frames, and a VideoMAE-based camera/object dynamics model. Captioning uses Tarsier2-7B, and InternVideo2-1B embeddings are clustered with K-Means into 10,000 clusters.

  • The image dataset is also very large. Kandinsky T2I contains more than 500 million general-domain images drawn from sources including LAION and COYO. Filtering uses a watermark classifier (watermark_resnext101_32x8d-large) plus a YOLO-based detector, TOPIQ and Q-Align for technical and aesthetic quality, CRAFT for text regions, and SAM 2 masks with a Sobel filter for complexity. Captions come from InternVL2-26B (with refinements by InternLM3-8B) and Qwen2.5VL-32B for Russian captions on images where width×height ≥ 512². Data is stored grouped by shortest side at 256, 512, and 1024.

  • The image-editing instruction dataset is built by matching image pairs. From roughly 240 million collected images, the pipeline applies CLIP and DINO similarity plus face recognition, clusters images into 10,000 groups for adaptive thresholding with T=0.15, verifies geometry with LoFTR and RANSAC (minimum 20 points per group), and enforces thresholds of DINO similarity above 0.8 with more than 300 inliers, CLIP similarity above 0.8 with more than 200 inliers, and face similarity above 0.7. Near-duplicates are excluded at DINO similarity above 0.97. This yields approximately 150 million curated image pairs with instructions.

  • Captioning model selection was done by human comparison. Side-by-side evaluation of GPT-4o, GPT-4 Mini with reasoning, Gemini 2.5 Pro, and Qwen2.5-VL-32B ranked Gemini 2.5 Pro first, with GLM 4.5 competitive; the authors chose GLM 4.5 without reasoning and fine-tuned it with LoRA for cost-effectiveness. For the editing SFT subset, filters required a Q-Align score above 4 and a Q-Align aesthetic score above 2, yielding approximately 600k candidate pairs for manual selection.

  • The SFT dataset is small and heavily vetted. The initial pool was 93,296 high-resolution video scenes and approximately 10 million images. Strict criteria (consensus of at least 3/2 agreement on "good" ratings) selected roughly 3% of the video pool and 5% of images, producing v1 with 2,833 video scenes and 45,000 images, and a relaxed v2 with 12,461 video scenes and 153,000 images. Data was classified into 9 domains using Qwen2.5-VL-Instruct-32B, with a category list covering animals, architecture, art, cartoons, food, interiors, nature, people, tech, and other.

  • Culturally specific data is included separately. The Kandinsky RCC dataset contains 229,504 video scenes and 768,555 images focused on the Russian cultural code, manually curated with manually written Russian descriptions that were machine-translated into English.

  • Supported resolutions are tabulated. The paper's resolution table lists 512×512, 512×768, 768×512, 768×1280, and 896×1152 for the Lite version, and 1024×1024, 640×1408, 1408×640, 1280×768, and 1152×896 for the Pro version.

  • This is the first Flow Matching generation in the family. Earlier Kandinsky versions used autoregressive or diffusion approaches, and Kandinsky 5.0 introduces Flow Matching.

Methodology in Plain English

The authors first assemble data at scale. Raw images and videos are collected from open datasets and large repositories, then passed through a chain of automated filters: resolution cutoffs, perceptual hashes to remove duplicates, watermark detectors, quality and aesthetic scorers, text detectors, segmentation-based complexity checks, and object/scene classifiers. Surviving items get synthetic captions from large multimodal models, and videos are additionally clustered for balanced sampling.

For training, the pipeline runs in stages: large-scale pretraining on the curated data, then supervised fine-tuning on a small, expert-filtered, high-quality subset, then reinforcement-learning-style adversarial post-training that compares generated images against SFT examples. Video models go through an additional distillation step that merges guidance distillation and consistency distillation so that far fewer sampling steps are needed.

Architecturally, the models use a Cross-Attention Diffusion Transformer (CrossDiT) built from CrossDiT blocks. To keep attention affordable on high-resolution, longer-than-5-second video, the NABLA mechanism computes attention sparsely — reported at a 90% sparsity ratio — which the authors say cuts training and inference time by 2.7 times without hurting quality.

Rounding the system out are engineering optimizations: faster VAE encoding, quantized text encoders, sharded data parallelism (F/HSDP), and activation checkpointing to fit large models into memory. Evaluation is done through human side-by-side comparisons, including on a prompt set from MovieGen.

Why This Matters

Research impact. The report combines a large-scale data curation description, a sparse attention method for video, and a distillation recipe into one openly released family. Releasing code, weights from multiple training stages, and diffusers integration lowers the barrier for other groups to study and build on video foundation models, similar to the role played by open projects such as HunyuanVideo, Mochi, CogVideoX, Wan, and VACE.

Real-world applications (as described in the paper):

  • Text-to-image synthesis and image editing, including precise instruction-based editing trained on the Kandinsky I2I dataset.
  • Text-to-video and image-to-video generation of up to 10-second, high-resolution clips.
  • Culturally specific generation, supported by the Russian-cultural-code dataset of 229,504 video scenes and 768,555 images.
  • Building blocks for multimedia generation systems and "world models," which the paper positions as analogous in significance to Large Language Models in NLP.

Industry relevance. Inference cost is a primary barrier to deploying video models. The reported reduction of NFE from 100 to 16, plus the 2.7 times attention speedup at a 90% sparsity ratio, directly targets the latency and compute budgets that determine whether video generation is commercially viable. Public release of the weights also gives industry a baseline for fine-tuning on domain-specific content.

Future Directions

  • **Scaling and the quality

Authors’ abstract

This report introduces Kandinsky 5.0, a family of state-of-the-art foundation models for high-resolution image and 10-second video synthesis. The framework comprises three core line-up of models: Kandinsky 5.0 Image Lite - a line-up of 6B parameter image generation models, Kandinsky 5.0 Video Lite - a fast and lightweight 2B parameter text-to-video and image-to-video models, and Kandinsky 5.0 Video Pro - 19B parameter models that achieves superior video generation quality. We provide a comprehensive review of the data curation lifecycle - including collection, processing, filtering and clustering - for the multi-stage training pipeline that involves extensive pre-training and incorporates quality-enhancement techniques such as self-supervised fine-tuning (SFT) and reinforcement learning (RL)-based post-training. We also present novel architectural, training, and inference optimizations that enable Kandinsky 5.0 to achieve high generation speeds and state-of-the-art performance across various tasks, as demonstrated by human evaluation. As a large-scale, publicly available generative framework, Kandinsky 5.0 leverages the full potential of its pre-training and subsequent stages to be adapted for a wide range of generative applications. We hope that this report, together with the release of our open-source code and training checkpoints, will substantially advance the development and accessibility of high-quality generative models for the research community.

Read the original paper