A 210-person trial tested a generative AI therapy chatbot
A randomised, waitlist-controlled trial of Therabot reported reduced depression, anxiety and eating-disorder symptoms at eight weeks.
- Historical event
- March 27, 2025
- First source published
- March 27, 2025
- Site publication
- September 18, 2026

What happened
On 27 March 2025, NEJM AI published 'Randomized Trial of a Generative AI Chatbot for Mental Health Treatment' by Michael Heinz, Nicholas Jacobson and eight co-authors at Dartmouth. The trial tested Therabot, described in the paper's abstract as an 'expert-fine-tuned' generative AI system built for mental health treatment, not a companion product. It enrolled 210 adults nationally, screened into three groups: major depressive disorder, generalized anxiety disorder, or high risk for a feeding or eating disorder. Participants were randomly assigned to a four-week Therabot intervention (106 people) or a waitlist control (104 people); the waitlist group received no access until after the study's eight-week endpoint.
What the documents show
A Dartmouth press release reports a 51 percent average reduction in depression symptoms, a 31 percent reduction in anxiety symptoms, and a 19 percent reduction in body-image and weight concerns, measured against the waitlist group. The abstract states primary outcomes were symptom change from baseline to the 4-week postintervention point and an 8-week follow-up, with secondary outcomes covering engagement, acceptability and therapeutic alliance, the collaborative relationship between a user and a therapeutic tool. The analysis used mixed models and effect sizes from a log-odds ratio, comparing differential change between groups rather than a simple before-and-after count.
The mechanism
A waitlist-controlled randomised trial is a different kind of evidence from the surveys and forum analyses common elsewhere in this record set. Randomising who received Therabot immediately, versus who waited, means a measured difference between groups is far less likely to be explained by who chose to seek help, or when, than in a self-selected sample. That is what lets a trial claim a treatment effect rather than a correlation. Therabot is also a distinct product category from a companion app: fine-tuned specifically as a therapeutic tool and evaluated against clinical scales, not a persona-driven companion evaluated on engagement.
What it leaves open
Dartmouth's account quotes the senior author stating that AI-powered therapy remains in critical need of clinician oversight and that no generative AI system is ready to operate fully autonomously in mental health, calling for further work to quantify the risks involved. Neither the press release nor the bibliographic record for this article states a funding source, so that cannot be confirmed here. This is an editorial reading: an eight-week trial of a purpose-built clinical tool, run by an academic team, says little on its own about consumer companion apps not built or tested to the same clinical standard.
- Was the comparison a randomised waitlist control, or a before-and-after measure with no comparison group?
- Was the chatbot purpose-built and clinically evaluated, or a general companion app used for an unintended purpose?
- Did the authors call for more oversight or further study, and does that caution travel with the result when it is cited elsewhere?
Read as intended, the trial is early but real evidence that a purpose-built, clinically supervised chatbot produced measurable symptom change against a randomised control over eight weeks. It is not evidence that a general-purpose companion app, built for a different purpose and without this trial design, would show the same result.
Sources & reading trail
The trial's own abstract: randomised, waitlist-controlled design, N=210 (106 vs 104), three screened conditions, 4-week and 8-week outcome points, and the analysis method.
Source published: 27 March 2025 · Retrieved: 16 September 2026
Reports the topline percentage symptom reductions and quotes the senior author's caution that clinician oversight remains necessary.
Source published: 27 March 2025 · Retrieved: 16 September 2026
Independently confirms the paper's exact title, ten-author byline, journal and 27 March 2025 publication date.
Source published: 27 March 2025 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- A small trial gave Woebot's chatbot claims their evidence template
- A 2025 study pair found heavy chatbot use linked to worse outcomes
- Crisis pop-ups differ across chatbots and one new law
- Browse the complete the archive
Sources & reading trail
- Randomized Trial of a Generative AI Chatbot for Mental Health Treatment
Source published: March 27, 2025 · Retrieved: September 16, 2026 - First therapy chatbot trial yields mental health benefits, Dartmouth-led study finds
Source published: March 27, 2025 · Retrieved: September 16, 2026 - Crossref metadata record for DOI 10.1056/AIoa2400802
Source published: March 27, 2025 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.