AuK unveils open-source foundational model for speech generation and editing
open-source speech
| Source: HF Papers | Original article
AuK, an open‑source foundational model, unifies speech generation and editing via natural‑language instructions and audio context, built on roughly 3.03 B instruction‑audio instances and 1.95 M examples.
A new technical report from Tencent‑Hunyuan unveils AuK, a 1.5 billion‑parameter open‑source foundation model that merges speech synthesis, editing and enhancement under a single natural‑language interface. The authors built the model on roughly 3.03 billion instruction‑audio pairs derived from 1.95 million hours of diverse audio, covering five task families that include zero‑shot text‑to‑speech, content and acoustic editing, paralinguistic manipulation, speech enhancement and source separation. By accepting plain language commands together with an audio context, AuK can generate new utterances, modify existing recordings, or clean noisy inputs without switching between specialised systems.
The release matters because it consolidates capabilities that have traditionally required separate, often proprietary, models. An open‑source, instruction‑following speech engine lowers the entry barrier for developers, researchers and startups seeking to embed advanced voice functionalities in applications ranging from virtual assistants to media production tools. Moreover, the scale of supervision—nearly two million hours of audio—places AuK among the most extensively trained speech models, promising higher fidelity and broader language coverage than many existing open solutions.
The community will now watch how quickly AuK is adopted in downstream projects and whether benchmark results confirm its claimed versatility. Early indicators to monitor include integration into open‑source toolkits, contributions to the GitHub repository, and performance comparisons on standard TTS and audio‑editing datasets. Equally important will be any follow‑up releases that expand the model’s size or add multilingual support, as well as potential collaborations that leverage AuK for multimodal AI systems. If the model lives up to its promise, it could become a cornerstone for the next generation of voice‑centric AI services.
Sources
Back to AIPULSEN