Menu

Categories

Tags

Xiaomi's new TTS model sings, clones voices, and takes mood requests

April 24, 2026 | Source: xiaomimimo | AI, Developer | 275 views 0 comments

Xiaomi has released the MiMo-V2.5-TTS family of speech synthesis models, now available via the MiMo open platform API. During the public beta, it's free — for a limited time. The series includes three models, each targeting a different use case.

MiMo-V2.5-TTS comes with a set of premium built-in voices and supports a singing mode that accurately captures pitch and rhythm. MiMo-V2.5-TTS-VoiceDesign lets you generate a brand-new voice from a single natural language description — no reference audio needed. You can define attributes like age, gender, accent, and temperament. MiMo-V2.5-TTS-VoiceClone handles voice cloning: give it a few seconds of reference audio, and it replicates the target speaker's voice, preserving breath, rhythm, and pause patterns — no training or fine-tuning required.

All three models support controlling speech style through natural language commands. Want a voice that's "gentle but tired" or "calm in a frenzy"? Just describe it. They also accept audio tags like "inhale," "laugh," or "choke up" for precise control. Language support includes Mandarin Chinese, English, and Chinese dialects such as Northeastern, Sichuanese, Henan, and Cantonese. Audio output is 24000 Hz, with PCM16 recommended for streaming.

Xiaomi has also been making moves in large language models — its MiMo V2.5 series recently entered public testing, with the flagship V2.5-Pro aiming to compete with Anthropic's Claude Opus 4.6 and OpenAI's GPT-5.4.

Tags: #Xiaomi

Leave a Reply

Your email address will not be published. Required fields are marked *