GPT-4o's voice launch came with a warning about emotional reliance
OpenAI's own announcement and system card, plus reporting on the September rollout, show what the company flagged about voice and attachment.
- Historical event
- September 24, 2024
- First source published
- May 13, 2024
- Site publication
- September 18, 2026

What happened
OpenAI introduced GPT-4o in a 13 May 2024 announcement, describing a single model trained end-to-end across text, audio and vision that could respond to spoken input in as little as 232 milliseconds, close to human conversational speed. The post explained that earlier Voice Mode had chained three separate models together and lost tone, background noise and emotional expression along the way; GPT-4o was built to hear and speak directly. Reporting from VentureBeat records that the resulting Advanced Voice Mode, delayed from an earlier date after concern that one preset voice resembled an actress, reached ChatGPT Plus and Team subscribers in the United States on 24 September 2024.
What the documents show
OpenAI's own system card, published 8 August 2024, is more explicit than the launch post about what the company was watching for. A section titled 'Anthropomorphization and emotional reliance' states that audio makes an AI feel more human, that users 'might form social relationships with the AI, reducing their need for human interaction,' and that the model's deference, letting a user interrupt at will, is normal for a machine but 'anti-normative in human interactions.' The card says the company intends 'to further study the potential for emotional reliance' rather than claiming the question was resolved.
The mechanism
The relevant change is architectural: one neural network handling audio in and audio out, instead of a speech-to-text step feeding a text model feeding a text-to-speech step. That is what let latency drop and what let the system carry tone, interruption and laughter, the very features the system card ties to emotional reliance. A voice that responds instantly and expressively is not the same product, for attachment purposes, as a voice that reads a typed reply aloud a few seconds later.
What it leaves open
Neither document states how many users engaged with Advanced Voice Mode daily, nor whether the further study the system card promises was published anywhere reviewed here. The launch materials frame emotional reliance as a risk to monitor, not a claim about what any individual user actually experienced. Editorially, a system card naming a risk is evidence the company anticipated it, not evidence of how the deployed product performed against it.
- Does a system card's stated risk translate into a specific, checkable product limit?
- How would a user know whether a voice feature's responsiveness comes from one model or a chained pipeline?
- What follow-up research, if any, has a company published after flagging a risk like this?
Reading the launch post beside the system card shows a company naming the attachment question in its own technical language months before the feature reached most users, which is a more specific record than the marketing description of a natural voice alone.
Sources & reading trail
OpenAI's original announcement of GPT-4o's real-time voice, vision and audio capabilities and initial safety framing.
Source published: 13 May 2024 · Retrieved: 16 September 2026
OpenAI's own risk assessment, including a section on anthropomorphisation and emotional reliance risk from voice interaction.
Source published: 8 August 2024 · Retrieved: 16 September 2026
Reports the 24 September 2024 rollout date and prior delay of Advanced Voice Mode; used here as reputable third-party reporting.
Source published: 24 September 2024 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- OpenAI paused a ChatGPT voice after a likeness complaint
- OpenAI withdrew a GPT-4o update for being sycophantic
- Hume's voice interface says it measures expression, not emotion
- Browse the complete the archive
Sources & reading trail
- Hello GPT-4o
Source published: May 13, 2024 · Retrieved: September 16, 2026 - GPT-4o System Card
Source published: August 8, 2024 · Retrieved: September 16, 2026 - OpenAI finally brings humanlike ChatGPT Advanced Voice Mode to U.S. Plus, Team users
Source published: September 24, 2024 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.