Wellbeing studies on chatbots differ in size and funding
A 2025 Therabot trial, an OpenAI-MIT study and a 2017 Woebot trial show how sample size and disclosed funding change a claim's weight.
- Site publication
- September 18, 2026

What happened
Three published studies illustrate how differently a chatbot wellbeing claim can be built. A Dartmouth-led randomised trial, published 27 March 2025 in NEJM AI, enrolled 210 adults with diagnosed depression, anxiety or eating-disorder risk, split into a 106-person Therabot group and a 104-person waitlist control, reporting a '51%' average reduction in depression symptoms. OpenAI and MIT Media Lab's joint study, published 21 March 2025, combined an automated analysis of 'nearly 40 million' ChatGPT conversations with a roughly 1,000-person randomised trial. A 2017 trial in the Journal of Medical Internet Research Mental Health tested the chatbot Woebot on 70 college-age participants over two weeks.
What the documents show
The three differ sharply on funding and disclosed interest. The Woebot paper states plainly that its second author 'is the founder' of 'Woebot Labs Inc', the company that 'covered the cost of participant incentives', a direct conflict of interest disclosed in the paper itself. The OpenAI-MIT study is described as 'a research collaboration' between the company whose product was studied and an outside academic lab. The Dartmouth-led Therabot trial, run through a university, does not carry the same company-authorship pattern, and its published abstract lists ten academic co-authors rather than a founder-operator. Sample sizes also differ by an order of magnitude between the small Woebot trial and the larger Therabot and OpenAI studies.
The mechanism
A randomised controlled trial, like Therabot's or Woebot's, assigns people to a treatment or control group to isolate a causal effect; an observational study, like OpenAI's 40-million-conversation analysis, instead describes correlations in real usage without assigning anyone to a condition. The two designs answer different questions, and OpenAI's own study explicitly combined both for this reason. A conflict of interest, such as Woebot Labs funding a trial of its own product, does not by itself invalidate a result, but it is a fact a reader needs in order to weigh the finding, and the Woebot paper's own disclosure is a model of stating it plainly.
What it leaves open
None of these three studies claims to generalise beyond its own population; the OpenAI-MIT study states directly that it advises 'against generalizing the results'. Treating any single trial as proof about companion apps generally goes beyond what these documents themselves claim.
- Was the study a randomised trial or an observational analysis, and does the claim match the design?
- Who funded the study, and does a study author or funder have a financial stake in the product?
- What outcome measure was used, and over what time period?
This is an editorial checklist, not a verdict on any product: a wellbeing claim is only as strong as the sample, design and disclosed funding behind the study a company chooses to cite.
Sources & reading trail
Reports the Therabot randomised trial's 210-person sample, group sizes and symptom-reduction percentages.
Source published: 27 March 2025 · Retrieved: 16 September 2026
Confirms the trial's design, national recruitment, N=210, group allocation and primary/secondary outcomes as published in NEJM AI.
Source published: 27 March 2025 · Retrieved: 16 September 2026
Describes a joint OpenAI-MIT Media Lab observational and randomised study design and an explicit caution against generalising the results.
Source published: 21 March 2025 · Retrieved: 16 September 2026
Reports a 70-participant, two-week trial with a disclosed conflict of interest: the study's co-author founded the company that made and funded testing of Woebot.
Source published: 6 June 2017 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- A 210-person trial tested a generative AI therapy chatbot
- A 2025 study pair found heavy chatbot use linked to worse outcomes
- A small trial gave Woebot's chatbot claims their evidence template
- Browse the complete the archive
Sources & reading trail
- First therapy chatbot trial yields mental health benefits
Source published: March 27, 2025 · Retrieved: September 16, 2026 - Randomized Trial of a Generative AI Chatbot for Mental Health Treatment (abstract)
Source published: March 27, 2025 · Retrieved: September 16, 2026 - Early methods for studying affective use and emotional well-being on ChatGPT
Source published: March 21, 2025 · Retrieved: September 16, 2026 - Delivering Cognitive Behavior Therapy to Young Adults With Symptoms of Depression and Anxiety Using a Fully Automated Conversational Agent (Woebot)
Source published: June 6, 2017 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.