Home
AI news
Platform Updates
OpenAI Releases GPT-4o with Voice, Vision and Text in One Model
Platform Updates
Confirmed
Brief

OpenAI Releases GPT-4o with Voice, Vision and Text in One Model

GPT-4o answers spoken input in about 320 milliseconds — close to human conversational latency — but voice shipped months after the announcement.

The Promptifi Team
·
May 13, 2024
·
·
OpenAI
Prompts in the library for this workflow
Browse the library →
Why this matters for sellers

Sub-second spoken latency is what makes live roleplay practice tolerable rather than stilted — but if you tried it the week of the announcement, you could not, and that gap between demo and availability is now a standing pattern worth pricing in.

In plain English

OpenAI announced GPT-4o on 13 May 2024. The model accepts any combination of text, audio, image and video input and returns text, audio and image output. OpenAI reported audio response times as low as 232 milliseconds, averaging 320 milliseconds.

Real-time voice did not ship on announcement day. Advanced Voice Mode began rolling out to paying users in late September 2024, more than four months later.

What it means for you
What to run
Copy the prompt →
Free · No account needed

The brief for people who carry a number

Two minutes, once a week. What changed in AI, and what to run because of it.

One email a week · No product pitches · Unsubscribe any time
P

The Promptifi Team

Promptifi's AI desk is written and fact-checked by people who carry a number. Drafting is AI-assisted; every item is checked against its primary source by a human before it publishes.