Xiaomi AI Lab's next-gen Kaldi team has open-sourced OmniVoice, a zero-shot voice cloning TTS model that supports 646 languages. Give it a few seconds of reference audio and it can replicate that voice — even across languages. Hand it a Mandarin clip, and the model can speak Japanese, Korean, or any other language in the same voice. Code, weights, and training data are all out in the open under the Apache-2.0 license.
Architecturally, OmniVoice keeps things simple. The whole model is a single bidirectional Transformer that directly maps text to multi-codebook acoustic tokens — no two-stage pipeline of semantic-to-acoustic. Two key design choices make this work: a full-codebook random masking strategy to improve training efficiency, and initialization with pre-trained LLM parameters to boost pronunciation accuracy. Inference runs at 40x real-time in plain PyTorch, no extra optimizations needed.
Training data comes entirely from 50 open-source speech datasets, totaling 580,000 hours after noise reduction and quality filtering. Low-resource languages get dynamic upsampling to ensure effective training. In tests across 24 languages, OmniVoice beat multiple commercial systems in both voice similarity and intelligibility. On a broader test of 102 languages, its intelligibility matched or even exceeded real recordings. It can even synthesize languages with less than 10 hours of training data.
Beyond voice cloning, the model also supports text descriptions for custom voices (like "male, middle-aged, very low pitch" or "female, young, Sichuan dialect"), automatic denoising of noisy reference audio, insertion of emotional markers like laughter and sighs, and pronunciation correction for polyphonic characters and proper nouns in Chinese and English.