RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
Lovebot Journaljournal
← The archive

A 2025 study pair found heavy chatbot use linked to worse outcomes

OpenAI's analysis of millions of chats and an MIT trial of 981 users found affective use rare, but heavier voluntary use tracked worse outcomes.

Historical event
March 21, 2025
First source published
March 21, 2025
Site publication
September 18, 2026
Visual for this record: A 2025 study pair found heavy chatbot use linked to worse outcomes
Visual published by openaimpact.com, shown for identification of the record. Credit: openaimpact.com · source page ↗ Rights: owner-review-pending. Source

What happened

On 21 March 2025, OpenAI and the MIT Media Lab jointly published 'Early methods for studying affective use and emotional well-being on ChatGPT', describing two linked studies. OpenAI ran an automated, privacy-preserving classifier analysis of nearly 40 million ChatGPT conversations, paired with targeted user surveys. Separately, MIT Media Lab researchers with OpenAI co-authors ran a four-week, IRB-approved, pre-registered randomised controlled trial, described in an accompanying paper, with 981 participants exchanging over 300,000 messages, randomly assigned to a voice mode (engaging, neutral or text) and a conversation type (personal, non-personal or open-ended).

What the documents show

OpenAI's post reports affective cues were absent from most sampled conversations, and that even among heavy users of Advanced Voice Mode, defined as the platform's top daily users by message count, high affective use was concentrated in a small subgroup, more likely to describe ChatGPT as a friend. The trial paper, corroborated by an MIT Media Lab summary, states random assignment to modality or conversation type produced no significant differences in outcomes, but participants who voluntarily used the chatbot more, whatever condition assigned, showed consistently worse outcomes, and greater trust and 'social attraction' toward the chatbot correlated with more emotional dependence and problematic use.

The mechanism

Random assignment is what lets a trial claim cause and effect for the thing actually randomised: here, voice mode and conversation type, where the comparison found no significant effect. The widely quoted 'heavy user' finding is different: it comes from observing how much people chose to use the chatbot once inside the trial, a behaviour nobody assigned. That is a correlation, not a randomised test, so it cannot establish that heavier use caused worse outcomes rather than people already prone to worse outcomes using the tool more. OpenAI's separate conversation analysis is observational in the same way, describing existing patterns rather than testing an intervention.

What it leaves open

Both organisations' own stated limitations note the findings are not yet peer-reviewed, are specific to ChatGPT rather than chatbots generally, rest on self-report and imperfect automated classifiers, cover only English-language conversations from United States participants, and exclude users under 18. A four-week trial may also be too short to detect slower-building effects. This is an editorial reading: these studies examine a general-purpose assistant occasionally used affectively, not a dedicated companion app built around a persistent persona, so how far the findings transfer to that category remains an open question.

Together the documents are useful for separating a rare behaviour, affective use of a general assistant, from a more specific correlation between heavier voluntary use and worse self-reported outcomes. Neither document claims the second finding is causal, and both name that as a limit of their own design.

Sources & reading trail

OpenAI's own description of both studies, the affective-cue and heavy-user findings, the Advanced Voice Mode heavy-user definition, and the stated limitations.

Source published: 21 March 2025 · Retrieved: 16 September 2026

The trial paper's abstract: 981-participant, four-week randomised design, and the voluntary-use/worse-outcome and trust/dependence correlations.

Source published: 21 March 2025 · Retrieved: 16 September 2026

Independently confirms the sample size, message count, and that no significant effect was detected from the randomised conditions themselves.

Source published: 21 March 2025 · Retrieved: 16 September 2026

Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.

Continue reading

Sources & reading trail

The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.

Published September 18, 2026, not on the date of the event described.