Kandinsky 6.0 Video: Foundation Models Achieve Synchronized Video‑Audio Generation
| Source: HF Papers | Original article
A new family of foundation diffusion models called Kandinsky 6.0 Video can generate synchronized 5‑second video clips with 44 kHz audio, offered in Lite (3 B parameters) and Pro (29 B parameters).
Kandinsky Lab has unveiled Kandinsky 6.0 Video, a new family of diffusion‑based foundation models that generate short, synchronized video‑audio clips from text or images. The lineup comprises a lightweight 3 billion‑parameter version (Video Lite) and a larger 29 billion‑parameter variant (Video Pro). Both models produce five‑second clips with 44 kHz audio that is automatically aligned to the visual content, including lip‑sync, and support two generation modes: text‑to‑audio‑video (T2AV) and image‑to‑audio‑video (I2AV).
The announcement marks the first time a publicly described diffusion system delivers fully synchronized sound and picture without separate post‑processing steps. By embedding audio generation directly into the diffusion pipeline, Kandinsky 6.0 Video promises tighter audiovisual coherence than workflows that stitch together independently created tracks. The dual‑size offering also gives creators a choice between faster, lower‑cost inference (Lite) and higher‑fidelity output (Pro), echoing the lab’s recent rollout of Kandinsky 6 Image Pro, which introduced a Mixture‑of‑Experts architecture for faster, more nuanced image generation.
Industry observers will be watching how the models perform on benchmark tasks such as lip‑sync accuracy, temporal consistency, and audio quality, as well as whether the five‑second limit expands in future releases. Integration into existing content‑creation pipelines—social media, marketing, and small‑scale production—could accelerate adoption, especially if the models are made accessible via APIs or open‑source licenses. The next steps are likely to involve larger‑scale evaluations, potential extensions to longer durations, and comparisons with competing multimodal generators that have emerged in the past year.
Sources
Back to AIPULSEN