Chinese AI startup StepAudio has released StepAudio 2.5 Realtime, an end-to-end real-time voice model that aims for genuinely human-like conversation. It supports full persona customization and can perceive paralinguistic cues — tone, pauses, sighs, you name it. The model is now available via API on the company's open platform.
The company ran five benchmark tests (as of April 2026) and claims a clean sweep. The most telling metric — subjective evaluation where real people rated conversations through a mobile app — gave StepAudio 2.5 Realtime a score of 80.41, versus GPT-Realtime-1.5's 68.01 and Gemini Live's 67.16. On a voice QA benchmark, it scored 79.80, nearly 1.5 times GPT-Realtime-1.5's 53.20. Other scores: paralinguistic understanding 82.18, general conversation 86.36, and in-car scenario 84.80.
Technically, StepAudio says three design choices matter. First, it built a matrix of millions of persona features by algorithmically expanding over 10,000 original personas, trained on real conversational data to handle niche topics. Second, it applied specialized RLHF (reinforcement learning from human feedback) for role-playing scenarios to prevent the AI from breaking character mid-conversation. Third, deep integration of understanding and generation inherits expressiveness from its StepAudio 2.5 TTS sibling, allowing both global scene setting and fine-grained sentence-level nuance.
The API is compatible with OpenAI's Realtime API protocol over WebSocket, making migration cheap for developers. Pricing: input at 10 yuan per million tokens (2 yuan with cache hit), output at 70 yuan per million tokens. StepAudio estimates continuous voice calls cost about 3.8 yuan per hour — roughly $0.53.