GPT-4o answers spoken input in about 320 milliseconds — close to human conversational latency — but voice shipped months after the announcement.
Sub-second spoken latency is what makes live roleplay practice tolerable rather than stilted — but if you tried it the week of the announcement, you could not, and that gap between demo and availability is now a standing pattern worth pricing in.
OpenAI announced GPT-4o on 13 May 2024. The model accepts any combination of text, audio, image and video input and returns text, audio and image output. OpenAI reported audio response times as low as 232 milliseconds, averaging 320 milliseconds.
Real-time voice did not ship on announcement day. Advanced Voice Mode began rolling out to paying users in late September 2024, more than four months later.
Two minutes, once a week. What changed in AI, and what to run because of it.