“People felt heard,” “loneliness fell,” and “treats depression” are not stronger versions of the same statement. They are different claims requiring different evidence and safeguards. A claim ladder helps readers slow down without dismissing useful experiences or converting them into medical conclusions.
Place the exact sentence
Rung one is a product property: the app offers journaling or reflective prompts. Rung two is reported experience: some users say it feels supportive. Rung three is an association: heavier or lighter use correlates with an outcome. Rung four is a controlled intervention result in a defined sample. Rung five is a medical purpose, such as diagnosis or treatment, that raises regulatory and clinical questions. Never upgrade a claim merely because the prose sounds confident.
Inspect the study frame
Record population, sample size, recruitment, comparator, duration, attrition, outcome measure, preregistration, funding and adverse-event handling. Ask whether the studied product is the one being marketed now. NIST’s TEVV-Athlon Initial Public Draft emphasizes evaluation around a specified objective and context; it is a 2026 public-comment draft, not a final standard, and a result from a short student study cannot silently become a promise for every age, culture or level of distress.
Separate wellness from medical purpose
The FDA’s 2026 clinical decision support guidance explains that software intended for patients or caregivers may fall outside the non-device CDS exclusion and that existing digital-health policies still apply. That does not let a reader classify a product from a slogan. It does mean “not medical advice” is not a magic phrase if the actual function makes disease-specific recommendations or treatment claims.
Work a hypothetical headline
Suppose a four-week, company-funded survey of 120 self-selected users reports lower loneliness scores with no control group. A fair summary is that participants reported change during use. It does not establish causation, durability, comparative benefit or safety in crisis. Before publication, seek protocol details, full results and conflicts; if unavailable, say so.
Use an action rule
For ordinary companionship claims, uncertainty may simply guide a purchase. For diagnosis, medication, self-harm or treatment decisions, do not rely on a companion or this guide; seek qualified human and emergency support appropriate to the situation. Report dangerous product behavior through available safety channels.
Let the rung change the decision
A rung-two self-report may justify trying a low-cost journaling feature, but not replacing human care. A well-run controlled study may justify greater confidence in its measured population, yet still leave long-term harms and product-version changes unresolved. A medical-purpose claim demands regulatory and clinical scrutiny beyond a blog summary. Write the allowed decision beside the rung. Evidence appraisal fails when every level ends in the same enthusiastic recommendation.
The ladder is a reading discipline, not a verdict on companionship. Use the study checklist for deeper appraisal and the wellbeing study record for an example where methods and limitations matter more than a single number.
Sources & reading trail
- Clinical Decision Support SoftwareUS Food and Drug Administration · retrieved 2026-09-19
- The TEVV-Athlon Framework for Evaluating AI Systems (Initial Public Draft)National Institute of Standards and Technology · published 2026-08-07 · retrieved 2026-09-19