
OpenAI has released two new speech-to-text models: GPT Transcribe for recorded files and batch jobs, and GPT Live Transcribe for real-time scenarios like live captions and phone calls.
Both models leverage context from the recording's topic, keywords, and language cues to improve recognition of short phrases, numbers, technical jargon, multiple accents, and noisy environments — areas where previous models often struggled.
According to Artificial Analysis, GPT Transcribe achieves a word error rate of just 3.31%, an improvement of 0.7 percentage points over the previous GPT-4o Transcribe. OpenAI also cut the price by 25%, to $4.50 per 1,000 minutes of audio.
https://twitter.com/OpenAIDevs/status/2082201169443905798