A 2025 study pair found heavy chatbot use linked to worse outcomes
OpenAI's analysis of millions of chats and an MIT trial of 981 users found affective use rare, but heavier voluntary use tracked worse outcomes.
- Historical event
- March 21, 2025
- First source published
- March 21, 2025
- Site publication
- September 18, 2026

What happened
On 21 March 2025, OpenAI and the MIT Media Lab jointly published 'Early methods for studying affective use and emotional well-being on ChatGPT', describing two linked studies. OpenAI ran an automated, privacy-preserving classifier analysis of nearly 40 million ChatGPT conversations, paired with targeted user surveys. Separately, MIT Media Lab researchers with OpenAI co-authors ran a four-week, IRB-approved, pre-registered randomised controlled trial, described in an accompanying paper, with 981 participants exchanging over 300,000 messages, randomly assigned to a voice mode (engaging, neutral or text) and a conversation type (personal, non-personal or open-ended).
What the documents show
OpenAI's post reports affective cues were absent from most sampled conversations, and that even among heavy users of Advanced Voice Mode, defined as the platform's top daily users by message count, high affective use was concentrated in a small subgroup, more likely to describe ChatGPT as a friend. The trial paper, corroborated by an MIT Media Lab summary, states random assignment to modality or conversation type produced no significant differences in outcomes, but participants who voluntarily used the chatbot more, whatever condition assigned, showed consistently worse outcomes, and greater trust and 'social attraction' toward the chatbot correlated with more emotional dependence and problematic use.
The mechanism
Random assignment is what lets a trial claim cause and effect for the thing actually randomised: here, voice mode and conversation type, where the comparison found no significant effect. The widely quoted 'heavy user' finding is different: it comes from observing how much people chose to use the chatbot once inside the trial, a behaviour nobody assigned. That is a correlation, not a randomised test, so it cannot establish that heavier use caused worse outcomes rather than people already prone to worse outcomes using the tool more. OpenAI's separate conversation analysis is observational in the same way, describing existing patterns rather than testing an intervention.
What it leaves open
Both organisations' own stated limitations note the findings are not yet peer-reviewed, are specific to ChatGPT rather than chatbots generally, rest on self-report and imperfect automated classifiers, cover only English-language conversations from United States participants, and exclude users under 18. A four-week trial may also be too short to detect slower-building effects. This is an editorial reading: these studies examine a general-purpose assistant occasionally used affectively, not a dedicated companion app built around a persistent persona, so how far the findings transfer to that category remains an open question.
- Was a reported relationship established by random assignment, or by observing behaviour participants chose for themselves?
- Does a 'heavy user' finding show the chatbot caused worse outcomes, or that people already prone to worse outcomes used it more?
- Do findings about a general assistant's voice and text features transfer to an app built specifically around a persistent companion persona?
Together the documents are useful for separating a rare behaviour, affective use of a general assistant, from a more specific correlation between heavier voluntary use and worse self-reported outcomes. Neither document claims the second finding is causal, and both name that as a limit of their own design.
Sources & reading trail
OpenAI's own description of both studies, the affective-cue and heavy-user findings, the Advanced Voice Mode heavy-user definition, and the stated limitations.
Source published: 21 March 2025 · Retrieved: 16 September 2026
The trial paper's abstract: 981-participant, four-week randomised design, and the voluntary-use/worse-outcome and trust/dependence correlations.
Source published: 21 March 2025 · Retrieved: 16 September 2026
Independently confirms the sample size, message count, and that no significant effect was detected from the randomised conditions themselves.
Source published: 21 March 2025 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- Anthropic found companionship use under half a percent of chats
- GPT-4o's voice launch came with a warning about emotional reliance
- A lab measured how models flatter user beliefs
- Browse the complete the archive
Sources & reading trail
- Early methods for studying affective use and emotional well-being on ChatGPT
Source published: March 21, 2025 · Retrieved: September 16, 2026 - How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Controlled Study
Source published: March 21, 2025 · Retrieved: September 16, 2026 - How AI and Human Behaviors Shape Psychosocial Effects of Extended Chatbot Use: A Longitudinal Controlled Study (MIT Media Lab publication page)
Source published: March 21, 2025 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.