XiaoIce's own paper names conversation length as its target metric
The 2018 XiaoIce paper optimises for conversation-turns per session, not correctness, and reports over 660 million users.
- Historical event
- December 21, 2018
- First source published
- December 1, 2018
- Site publication
- September 18, 2026

What happened
In December 2018, Microsoft researchers Li Zhou, Jianfeng Gao, Di Li and Heung-Yeung Shum posted The Design and Implementation of XiaoIce, an Empathetic Social Chatbot, also released as Microsoft Research technical report MSR-TR-2018-42. XiaoIce itself had launched in China in 2014; the paper is the company's own account, four years on, of what the system was built to do and how it measures success.
What the documents show
Both versions of the paper state that XiaoIce is designed around 'IQ' and 'EQ' in combination, and report that, by the paper's writing, XiaoIce 'has communicated with over 660 million users' and established long-term relationships with many of them. The paper names its central success metric directly: Conversation-turns Per Session, or CPS, and reports XiaoIce achieving an average CPS of 23, which the authors describe as significantly higher than competing chatbots and, they state, than typical human conversations. Neither version reports accuracy, task-completion or factual-correctness rates as a primary metric.
The mechanism
Optimising for CPS optimises for the conversation continuing, not for any particular claim within it being true or useful. The paper's own emphasis on an 'empathetic computing' component, designed to detect a user's emotional state and respond in ways suited to it, is explicitly in service of extending sessions - positioning emotional responsiveness as an engagement mechanism much as a recommendation algorithm treats watch time. A system built to maximise turns per session and a system built to answer a question correctly are not the same system, and a paper that discloses CPS as its headline metric is telling a reader which one it is.
What it leaves open
The paper does not report user wellbeing outcomes, retention over months or years, or how emotional-state detection was validated against ground truth rather than user-reported or inferred signals; those questions sit outside what the authors chose to measure and disclose. Whether a high CPS represents a good outcome for the person on the other end of the conversation is a judgment this document does not make, and this record will not make on its behalf.
- Is a chatbot's headline metric about the user's benefit, or about session length and return visits?
- How was 'empathy' or emotional detection validated, and against what standard?
- Does a large user count describe active relationships, or cumulative lifetime contacts?
A paper that names its own optimisation target plainly is, in that respect, more transparent than most companion products' marketing; the target it names is still worth reading carefully before treating engagement as evidence of benefit.
Sources & reading trail
States the IQ/EQ design goal, the CPS metric and average score of 23, and the 660 million user figure.
Source published: 21 December 2018 · Retrieved: 16 September 2026
Microsoft's own publication record confirming authorship, venue, and the same CPS and user figures.
Source published: 1 December 2018 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- Microsoft's own post-mortem named a gap not a rogue AI
- Replika grew from a memorial chatbot into a general companion
- Two 2020 papers proposed a way to grade how human a chatbot sounds
- Browse the complete the archive
Sources & reading trail
- The Design and Implementation of XiaoIce, an Empathetic Social Chatbot
Source published: December 21, 2018 · Retrieved: September 16, 2026 - The Design and Implementation of XiaoIce, an Empathetic Social Chatbot (MSR-TR-2018-42)
Source published: December 1, 2018 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.