OpenAI's new transcription models are faster and 25 percent cheaper, but still not the most accurate
GPT Transcribe and GPT Live Transcribe cut the price to $0.0045 per minute and hit a 3.31 percent word error rate. ElevenLabs, Google and Mistral all score better.
OpenAI has released two new speech recognition models through its API: GPT Transcribe for pre-recorded audio files and GPT Live Transcribe for real-time streaming. Both turn spoken audio into text, the technology behind meeting notes, subtitles, podcast transcripts and voice assistants. The file version processes audio about 34 times faster than real time, so an hour-long recording is done in under two minutes.
The numbers are solid but not spectacular. According to Artificial Analysis, an independent benchmarking outfit, GPT Transcribe reaches a word error rate of 3.31 percent. Word error rate, usually shortened to WER, is simply the share of words the system gets wrong, so lower is better. That is a 0.7 percentage point improvement over GPT-4o Transcribe, the year-old predecessor. Pricing drops by 25 percent at the same time, to $0.0045 per minute of audio. Both models accept extra context: you can hand them a text, a keyword list, or several input languages to improve accuracy on names and jargon.
Where OpenAI does not lead is accuracy. In the same ranking, ElevenLabs Scribe v2 sits at 2.3 percent, Google’s Gemini 3 Pro at 2.9 percent, and Mistral’s Voxtral Small at 3 percent. Mistral has also been undercutting everyone on price with Voxtral Transcribe V2, starting at $0.003 per minute, a third cheaper than OpenAI’s new rate.
Transcription used to be a specialist product you bought from a specialist company. It is now a commodity that every large lab ships as a line item, and the competition has moved from “does it work” to fractions of a percent and fractions of a cent. That is good news for buyers and uncomfortable news for anyone whose whole business was transcription. It also explains why OpenAI bundles the new models with its recently announced Realtime generation, which includes the streaming transcription model GPT-Realtime-Whisper: the value is shifting from the transcript itself to what sits around it.
What this means for you: If you occasionally need a transcript of an interview, a lecture or a family video, any of these models will do a fine job, and price differences of a tenth of a cent per minute will never show up on your bill. If you are building something on top of transcription, or processing thousands of hours, the ranking matters and OpenAI is currently not the best pick on either accuracy or price. A fair caveat: benchmark error rates are averages over test sets, and your audio, with its accents, background noise and specialist vocabulary, may rank the models differently. Test with your own recordings before you commit.
Sources
1,200 AI lab employees ask Washington for a brake pedal they admit nobody knows how to build
The Pacing the Frontier statement is signed by Dario Amodei, Jakub Pachocki and John Schulman. It does not call for a pause. It asks for the ability to slow down, and the signatories openly disagree on how.