OpenAI just dropped three new voice models into its Realtime API: GPT-Realtime-2 (voice conversation plus reasoning), GPT-Realtime-Translate (real-time translation), and GPT-Realtime-Whisper (streaming transcription). The standout is GPT-Realtime-2, OpenAI's first speech model with GPT-5-level reasoning, boasting a context window that jumps from 32K to 128K — enough for one to two hours of dense conversation.
GPT-Realtime-2 focuses on three areas. For reasoning, it offers five levels of intensity, from minimal to xhigh, letting developers choose based on their latency tolerance. For tool calling, the model can invoke multiple tools simultaneously, using casual spoken phrases like "Let me check your calendar" to keep users updated. If something goes wrong, it might say "Sorry, I'm having trouble with that right now" instead of just dropping the call. And the context window is quadrupled to 128K.
On benchmarks, GPT-Realtime-2 scores 15.2% higher on Big Bench Audio (an audio reasoning benchmark) and 13.8% higher on Audio MultiChallenge (a multi-turn instruction following benchmark) compared to its predecessor GPT-Realtime-1.5. Josh Weisberg, head of AI at Zillow, says that with prompt optimization, call completion rates on the hardest adversarial benchmarks jumped from 69% to 95%, with big improvements in fair housing compliance.
GPT-Realtime-Translate handles real-time speech translation across more than 70 input languages into 13 output languages. Deutsche Telekom is testing it for multilingual customer support. BolnaAI, an Indian voice AI company, reports that the model achieves a word error rate (WER) that is 12.5% lower than any previously tested model on Hindi, Tamil, and Telugu.
GPT-Realtime-Whisper is a streaming speech-to-text model that transcribes as you speak, meant for live captions, meeting notes, and similar use cases.
Pricing: GPT-Realtime-2 input is $32 per million tokens (cached input: $0.40), output $64 per million tokens; GPT-Realtime-Translate is $0.034 per minute; GPT-Realtime-Whisper is $0.017 per minute.