RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
Lovebot Journaljournal
← The archive

XiaoIce's own paper names conversation length as its target metric

The 2018 XiaoIce paper optimises for conversation-turns per session, not correctness, and reports over 660 million users.

Historical event
December 21, 2018
First source published
December 1, 2018
Site publication
September 18, 2026
Visual for this record: XiaoIce's own paper names conversation length as its target metric
Visual published by ai2-s2-public.s3.amazonaws.com, shown for identification of the record. Credit: ai2-s2-public.s3.amazonaws.com · source page ↗ Rights: owner-review-pending. Source

What happened

In December 2018, Microsoft researchers Li Zhou, Jianfeng Gao, Di Li and Heung-Yeung Shum posted The Design and Implementation of XiaoIce, an Empathetic Social Chatbot, also released as Microsoft Research technical report MSR-TR-2018-42. XiaoIce itself had launched in China in 2014; the paper is the company's own account, four years on, of what the system was built to do and how it measures success.

What the documents show

Both versions of the paper state that XiaoIce is designed around 'IQ' and 'EQ' in combination, and report that, by the paper's writing, XiaoIce 'has communicated with over 660 million users' and established long-term relationships with many of them. The paper names its central success metric directly: Conversation-turns Per Session, or CPS, and reports XiaoIce achieving an average CPS of 23, which the authors describe as significantly higher than competing chatbots and, they state, than typical human conversations. Neither version reports accuracy, task-completion or factual-correctness rates as a primary metric.

The mechanism

Optimising for CPS optimises for the conversation continuing, not for any particular claim within it being true or useful. The paper's own emphasis on an 'empathetic computing' component, designed to detect a user's emotional state and respond in ways suited to it, is explicitly in service of extending sessions - positioning emotional responsiveness as an engagement mechanism much as a recommendation algorithm treats watch time. A system built to maximise turns per session and a system built to answer a question correctly are not the same system, and a paper that discloses CPS as its headline metric is telling a reader which one it is.

What it leaves open

The paper does not report user wellbeing outcomes, retention over months or years, or how emotional-state detection was validated against ground truth rather than user-reported or inferred signals; those questions sit outside what the authors chose to measure and disclose. Whether a high CPS represents a good outcome for the person on the other end of the conversation is a judgment this document does not make, and this record will not make on its behalf.

A paper that names its own optimisation target plainly is, in that respect, more transparent than most companion products' marketing; the target it names is still worth reading carefully before treating engagement as evidence of benefit.

Sources & reading trail

States the IQ/EQ design goal, the CPS metric and average score of 23, and the 660 million user figure.

Source published: 21 December 2018 · Retrieved: 16 September 2026

Microsoft's own publication record confirming authorship, venue, and the same CPS and user figures.

Source published: 1 December 2018 · Retrieved: 16 September 2026

Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.

Continue reading

Sources & reading trail

The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.

Published September 18, 2026, not on the date of the event described.