Inworld AI has released Realtime TTS-2, a voice synthesis model built specifically for real-time conversation. Its predecessor, TTS 1.5, already topped the third-party Artificial Analysis Speech Arena, beating offerings from Google and ElevenLabs. But TTS-2 isn't just about sounding good — it adds four core capabilities that shift the focus from 'reading nicely' to 'speaking like a human.'
First, dialogue awareness. The model takes raw audio from previous conversation turns, not text transcripts, so it can tell the difference between a resigned 'okay' and a sarcastic one. The same line delivered after a joke sounds completely different from after bad news.
Second, natural language voice guidance. Developers can now describe the desired tone in plain English — think 'tired but gentle, like you just got home from work' — instead of picking from a fixed set of labels like 'happy' or 'sad.' It's prompt engineering for voice.
Third, cross-language consistency. The same voice character can switch seamlessly between over 100 languages, even mid-sentence, without losing its identity.
Fourth, text-to-voice generation. You can create a reusable voice character with nothing more than a text description — no recording samples needed.
The first audio chunk from the TTS layer comes in under 200 milliseconds. The model is available as a research preview through Inworld's API and Realtime API, supporting 15 official languages and over 90 experimental ones. It's already live on Cloudflare, LiveKit, and DeepInfra.
CEO Kylan Gibbs told Business Insider that Inworld is sticking to models and APIs, not consumer products, to avoid competing with its customers. The company has raised over $100 million from investors including Founders Fund, Intel, and Microsoft.