Resemble AI today open-sourced DramaBox on Hugging Face, a voice generation model that the company calls the first "directable" speech engine — meaning you can finally make AI voices sound like something other than a monotone robot.
The key feature is what they call "separated prompt control." You put the dialogue in double quotes, and outside the quotes you write stage directions like "sighs," "long pause," "whispers," or even "voice raspy from sadness." The model doesn't read those instructions aloud; instead, it physically renders the vocal emotion, turning text-to-speech into actual character performance. That replaces workflows that previously required human voice actors or painstaking post-production.
On the technical side, DramaBox offers zero-shot voice cloning — just 10 seconds of reference audio is enough to lock onto a target voice. You can also use natural language prompts to set a character's age, accent, and mood. The model outputs native 48kHz stereo studio-quality audio. To prevent deepfakes, all generated audio is watermarked by default with an invisible Perth watermark that survives MP3 compression and typical audio editing.
Under the hood, DramaBox is fine-tuned from Lightricks' 3.3 billion-parameter LTX-2.3 audio foundation model, combining diffusion Transformer (DiT) and flow matching architectures, with text embeddings handled by Google's Gemma 3 12B model.