Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation
Authors: Team Kandinsky, Julia Agafonova, Bulat Akhmatov, Mikhail Aksyutin, Grigorii Alekseenko, Anastasia Aliaskina, Olga Androsova, Vladimir Arkhipkin, +80 more
We present Kandinsky 6.0 Video, a family of foundation diffusion models for synchronized text-to-audio-video generation, comprising Kandinsky 6.0 Video Lite (3B parameters) and Kandinsky 6.0 Video Pro (29B parameters). Both models generate 5-second video clips with synchronized 44 kHz audio, including lip-sync, in text-to-audio-video (T2AV) and image-to-audio-video (I2AV) modes; a built-in super-resolution model raises the output resolution to Full-HD (1920×1080). Building on the video generation capabilities of Kandinsky 5.0, Kandinsky 6.0 Video employs a dual-stream CrossDiT architecture that connects a pretrained video stream and a newly trained audio stream through bidirectional cross-attention for temporal and semantic alignment. Our continuous pretraining strategy first trains the audio stream from scratch on large-scale audio corpora and then trains both streams jointly on paired audio-video data while preserving unimodal fidelity; pretraining is followed by supervised fine-tuning, reinforcement-learning-based post-training, and distillation. In side-by-side human evaluation, Kandinsky 6.0 Video Pro clearly outperforms its predecessor, Kandinsky 5.0 Video Pro, and remains competitive with leading audio-video generation models, particularly in speech quality. To accelerate open research and deployment in multimedia generation, we release the code, model checkpoints, and diffusers integration under the MIT license.
Figures & tables
Figure 1
Domain
Stage 1 (Visual)
Stage 2 (Sync)
people action
1,410
1,460
animals
591
840
cartoons
550
998
people speech
493
1,069
music
466
466
food
347
662
Table 1 : Domain breakdown of SFT datasets for the audio-visual model.
Figure 1 : Simplified scheme of a single Kandinsky 6.0 Video transformer block. The video (top) and audio (bottom) streams each apply self-attention, cross-attention to the text embeddings (T2V and T2A), cross-modal attention, and a feed-forward layer. In the V2A block, video queries attend to audio keys and values; in the A2V block, audio queries attend to video keys and values.
Figure 2 : Detailed scheme of the Kandinsky 6.0 Video transformer block. Each stream has its own text encoder block and timestep embedding; the timestep embedding modulates every sub-layer through scale, shift, and gate parameters. In the cross-modal attention layers, video tokens use 3D RoPE and audio tokens use 1D RoPE.
Parameter
T2V (Lite/Pro)
T2A (Lite/Pro)
T2AV (Lite/Pro)
I2AV (Lite/Pro)
CrossDiT
video_model_dim
1792 / 4096
–
1792 / 4096
1792 / 4096
audio_model_dim
–
896 / 2048
896 / 2048
896 / 2048
ff_dim
7168 / 16384
3584 / 7168
3584 (audio),
3584 (audio),
7168 (video) /
7168 (video) /
7168 (audio),
7168 (audio),
Table 2 : Architectural hyperparameters of Kandinsky 6.0 Video Lite and Pro (values given as Lite / Pro) at each training stage: video-only (T2V) and audio-only (T2A) pre-training, joint text-to-audio-video (T2AV) training, and image-to-audio-video (I2AV) training. The Qwen2.5-VL and CLIP text encoders are the same in all configurations.
Figure 3 : SR-DiT architecture. The noisy video latent, first-frame HQ anchor, and binary anchor mask are concatenated along the channel dimension and processed by timestep-modulated DiT blocks. The output head predicts the flow-matching velocity. The model operates without text conditioning.
Figure 4 : Latent Upscaler architecture: the ×4 branch (left), the ×2 branch (middle), and their building blocks (right). Both branches follow the K-VAE decoder geometry and map 64-channel LQ latents to upsampled 64-channel latents, preserving the temporal resolution. The outputs labeled ‘‘HQ latent’’ refer to the higher-resolution latent grid; they retain degradations, which are subsequently removed by SR-DiT.
Figure 5 : Training pipeline of Kandinsky 6.0 Video. The video stream (initialized from the pretrained Kandinsky 5.0 Video model) and the audio stream are first pre-trained separately, then fused and trained jointly in T2AV mode and in mixed T2AV/I2AV mode. Joint pre-training is followed by supervised fine-tuning with model soup, RL-based post-training, and distillation.
Parameter
T2V
T2A,
T2A,
T2AV
T2AV +
(Lite/Pro)
Audio captions
AV captions
(Lite/Pro)
I2AV (25 % )
(Lite/Pro)
(Lite/Pro)
(Lite/Pro)
num_steps
45k (T2V captions)
40k / 30k
2k / 4k
57.5k / 49k
4k / 5k
20k/14k (T2AV captions)
Optimizer
learning_rate
3e-05
0.0001
3e-05
5e-05 / 5e-05 (46k steps)
1e-05 / 2.5e-05
Table 3 : Training hyperparameters and data sampling settings for the continuous pre-training stages of Kandinsky 6.0 Video Lite and Pro (values given as Lite / Pro where they differ): separate T2V and T2A pre-training (the latter first on audio captions and then on audio-video captions), joint T2AV training, and the final T2AV stage with 25% of I2AV samples.
Figure 6 : Input construction in T2AV (top) and I2AV (bottom) training modes. In I2AV mode, the latent of the reference frame is concatenated with the noisy video latents along the temporal axis and marked with mask value 1 (0 for the frames to be generated); a learnable Token Role Embedding ( emb_ref / emb_video ) is added to the patch-embedded tokens. In both modes, the text prompt is encoded by CLIP and a VLM, and audio latents are produced by the audio VAE.
Figure 7 : Reward phase of RL post-training. For each prompt, a group of videos with audio is generated with different random seeds and scored by eight reward terms. Each score is centered within the group and scaled by the batch-wide standard deviation, routed into separate video and audio advantages (sync terms enter both), clamped to [−5,5] , broadcast to the video and audio tokens, and rescaled to per-token rewards rtoken∈[0,1] , where 1 corresponds to a rollout far above the group mean and 0 far below it.
Figure 8 : Training phase of RL post-training. The rollout latent is re-noised and denoised by the current LoRA, the old LoRA, and the frozen base (SFT) model. Predictions of the current and old LoRA are combined into positive and negative predictions; their per-token MSE losses to the rollout latent are interpolated with the rewards rtoken from the reward phase and aggregated with the region-wise attention weights attn_w derived from A2V cross-attention. A KL term with respect to the base-model prediction regularizes training; the old LoRA is an exponential moving average of the current one.
Figure 9 : Rollout strategies for adversarial training. Top: standard LADD, where the intermediate state xt is obtained by forward noising of the clean sample x0 and denoised in a single step with gradients. Middle: backward student rollout with gradients propagated through all denoising steps. Bottom: backward student rollout with gradients propagated only through the final steps.
Figure 10 : Discriminator architecture used in Sim-LADD. Newly initialized Kandinsky 6.0 Video transformer blocks are attached as heads to intermediate blocks of the pretrained backbone; each head outputs separate video and audio scores.
Figure 11 : SR-DiT training pipeline for restoring HQ video from degraded LQ latents with first-frame HQ anchor conditioning. Training combines a latent-space flow-matching loss with pixel-space reconstruction and anti-grid losses computed through the frozen K-VAE decoder. The diagram illustrates Stages 1-2; timestep sampling uses u∼U[0,1] followed by t=5u/(1+4u) .
Scale
Stage
MUSIQ ↑
DOVER ↑
CLIP-IQA ↑
Lapl. artifacts ↓
Warp error ↓
×4
Stage 1
48.92
0.5467
0.4426
0.4366
0.6141
Stage 2
50.38
0.5589
0.4569
0.5337
0.6478
Stage 3
50.41
0.5553
0.4587
0.4227
0.6781
×2
Stage 1
66.05
0.7329
0.5950
0.6062
0.6551
Stage 2
65.29
0.7282
0.5932
0.6072
0.6531
Stage 3
65.81
0.7348
0.5906
0.5710
0.6831
Table 4 : Validation metrics of the SR-DiT after each training stage on the tiled validation sets (912 videos each). Stage 1: dense attention; Stage 2: NABLA fine-tuning; Stage 3: π -Flow distillation. Values are comparable only within one scale factor.
×2
×4
Group
n
LU
Bilinear
Δ
LU
Bilinear
Δ
Kandinsky 5.0 T2V, weak
25
41.60
42.97
−1.37
39.19
39.16
+0.03
4K collection, weak Real-ESRGAN
19
41.42
42.37
−0.95
39.91
39.13
+0.78
Camera-motion clips (images)
25
38.62
38.26
+0.36
37.85
36.83
+1.02
4K collection, synthetic
25
45.08
42.92
+2.16
40.25
36.98
+3.26
All
94
41.70
41.58
+0.11
39.26
37.96
+1.30
Table 5 : PSNR (dB) of the Latent Upscaler on the held-out validation set, computed per frame on pixels decoded by K-VAE and averaged per clip. The baseline decodes the low-resolution latent and upsamples it bilinearly in pixel space; Δ is the difference between the LU and the baseline.
Figure 12 : Output resolutions of Kandinsky 6.0 Video at inference. In SD mode, the base-model output is returned without changes; in Full HD mode, it is upscaled × 2.25 by the super-resolution model and then downscaled to the target Full HD resolution.
Figure 13 : Tiled SR inference with the Latent Upscaler. The LQ video is encoded once with K-VAE and split into overlapping latent tiles. Each tile is upscaled by the ×2 or ×4 LU, refined by SR-DiT, decoded, and merged using Hann-window blending. The denoising loop shown corresponds to the pre-distillation model; the final distilled model replaces this loop with π -Flow DX policy sampling using two network evaluations (NFE =2 ).
Target
Source
Scale
SR output
Final
Tiles
Time, s
HD
768×432
×2
1536×864
1280×720
6
26.2
HD
512×512
×2
1024×1024
720×720
9
24.7
Full HD
864×480
×2.25
1952×1088
1920×1080
9
39.6
Full HD
512×512
×2.25
1152×1152
1080×1080
9
24.7
2K
864×480
×4
3456×1920
2560×1440
30
130.8
2K
768×432
×4
3072×1728
2560×1440
25
108.8
Table 6 : Super-resolution routes and processing time per 5-second clip on a single H100 GPU, using the Latent Upscaler, the distilled SR-DiT (NFE =2 ), and the Magi compiler [ 71 ] for K-VAE decoding. Times include K-VAE encoding and exclude the final resize. Spatial dimensions are given as width × height.
Resident
Block offload
Pro, SD
72.8
21.7
Pro, HD / Full HD
72.8
39.6–39.8
Lite, SD
43.5
21.7
Lite, HD / Full HD
44.5–44.8
39.5–39.8
Table 7 : Peak PyTorch allocated memory (GiB) on the H100 with and without block offloading, for both model sizes and all three operating modes. The two values in the HD/Full HD row correspond to the two base geometries ( 768×432 and 864×480 ), not to a range across repeats.
32 / 24 GB
16 GB
DiT block offloading
yes
yes
SR decode segment length
4 frames
4 frames
SR spatial split
none
2 strips
Text encoder
full
NF4, released after encoding
Table 8 : Settings of the VRAM budget presets for 32/24 GB and 16 GB GPUs.
32 / 24 GB
16 GB
SD
21.7
7.9
HD
17.7
13.7
Full HD
21.7
12.3
Table 9 : Peak PyTorch allocated memory (GiB) at the 32/24 GB and 16 GB presets. Peak memory is the same for Lite and Pro, because only two DiT blocks remain resident. The peak follows the base resolution, which is why HD peaks lower than SD and Full HD.
RTX 4090
RTX 5060 Ti
RTX 5080
RTX 5090
RTX PRO 6000
A100 80 GB
H100
Lite SD
437
1310
577
309
242
532
239
Lite HD
422
1328
579
296
274
472
203
Lite Full HD
578
1774
770
406
387
664
284
Pro SD
936
3080
1336
754
621
972
356
Pro HD
1194
2716
1189
638
561
791
292
Pro Full HD
1247
3530
1546
854
765
1106
402
Table 10 : Generation times in seconds for a 5-second clip with the full, non-distilled base model, after warmup, excluding weight loading and MP4 encoding. Both model sizes (Lite and Pro) are evaluated at three operating modes each: SD, HD, and Full HD, giving six configurations per GPU.
Figure 14 : VABench radar chart. All metrics are normalized, with desync and word error rate inverted so that a larger radius always indicates better performance.
LTX 2.5
Kandinsky 6.0 Video Lite
Kandinsky 6.0 Video Pro
speech_qn
1.427
1.477
1.496
audio_aes
3.307
3.323
3.345
tv_align
0.185
0.227
0.229
ta_align
0.307
0.374
0.367
av_align
0.258
0.245
0.268
lipsync
1.567
1.761
2.072
Table 11 : VABench results. Higher is better for all metrics except desync and word error rate. Best per row in bold, second best underlined.
Figure 15 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs Kandinsky 5.0 Video Pro . In all figures of this section, the light green segment denotes preference for Kandinsky 6.0 Video Pro , gray denotes ties, and the dark segment denotes preference for the opponent model.
Figure 16 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs Kling 2.6 .
Figure 17 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs LTX 2.5 .
Figure 18 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs Veo 3.1 Fast . Light green — Kandinsky 6.0 Video Pro, gray — ties, dark — Veo 3.1 Fast.
Figure 19 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs MiniMax H3 . Light green — Kandinsky 6.0 Video Pro, gray — ties, dark — MiniMax H3.
Figure 20 : Side-by-side comparison of Kandinsky 6.0 Video Pro vs Seedance 2.0 . Light green — Kandinsky 6.0 Video Pro, gray — ties, dark — Seedance 2.0.
Figure 21 : Side-by-side comparison of Kandinsky 6.0 Video Pro (RL) vs Kandinsky 6.0 Video Pro (SFT) on three benchmarks (T2AV — Text-to-Audio-Video, I2AV — Image-to-Audio-Video). Light green — RL, gray — ties, dark — SFT.
Figure 22 : Side-by-side comparison of Kandinsky 6.0 Video Pro (distilled) vs Kandinsky 6.0 Video Pro (full) . Light green — distilled, gray — ties, dark — full.
Figure 23 : A 3D animated astronaut in a bulky white spacesuit walks across the surface of an alien planet with glowing purple crystals and two distant planets in the sky, turns to the camera, and speaks in awe: <S>I have definitely never seen anything like this on Earth.<E> <AUDCAP> Soft, muffled footsteps on dusty alien soil; a gentle ethereal hum with subtle crystalline chimes and a soft cinematic score. <ENDAUDCAP>
Figure 24 : A cartoon dog detective in a trench coat and fedora examines a strange print through a magnifying glass, then looks up at the camera and says: <S>I know who it is!<E> <AUDCAP> Ambient night city sounds, distant traffic, a suspenseful underscore, and a sharp gasp of realization. <ENDAUDCAP>
Figure 25 : A cartoon mad scientist in a neon-lit laboratory mixes two glowing liquids in a flask; the mixture foams up, and he turns to the camera in panic: <S>I mixed up the formula! How did this happen?!<E> <AUDCAP> Hum of laboratory equipment, bubbling liquid, faint electrical zaps, vigorous fizzing, and a sharp gasp of panic. <ENDAUDCAP>
Figure 26 : A street musician plays acoustic guitar on a cobblestone square in warm late-afternoon light while a small crowd claps along. <AUDCAP> Melodic fingerpicked acoustic guitar with rhythmic synchronized clapping and faint ambient sounds of a public square. <ENDAUDCAP>
Figure 27 : A gloved hand drives a metal nail into a wooden plank with a claw hammer in a sunlit forest. Each strike sends small wood chips flying. <AUDCAP> Loud metallic clinks and heavy resonant thuds at each strike, with faint forest ambience and birds in the distance. <ENDAUDCAP>
Figure 28 : A young girl in a white medical coat with a stethoscope stands in a bright hospital room and speaks to the camera: <S>When I grow up, I want to be a doctor!<E> <AUDCAP> Quiet hospital room ambience, soft cheerful background music, and a clear, sweet, enthusiastic young voice. <ENDAUDCAP>
Figure 29 : A pirate captain on the deck of a wooden sailing ship in rough seas lowers his spyglass and shouts toward the camera: <S>Land ahead!<E> <AUDCAP> Crashing waves against the hull, creaking planks and rigging, strong wind, and a loud triumphant shout. <ENDAUDCAP>
Figure 30 : An animated painting based on Valentin Serov’s Girl with Peaches (1887); the original painting is in the public domain. The crops show the face and hair, the bow on the dress, and a fruit on the table. LU outputs remain visually close to the enlarged input. SR sharpens hair strands and adds texture to the bow, fruit, and tablecloth, with finer detail at ×4 than at ×2 .
Figure 31 : A tennis scene, frame 111. The crops show the face, the racket, and the shoe with the surrounding court surface. LU outputs remain visually close to the enlarged input. SR makes facial features, racket strings, and shoe and court textures more distinct, particularly at ×4 .
Figure 32 : An interview scene, frame 121. The crops show the face and glasses, the hand, and the fabric of the jacket. LU outputs remain visually close to the enlarged input. SR sharpens facial features and glasses contours and enhances hair and fabric detail, with more pronounced texture at ×4 than at ×2 .