RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
Lovebot Journaljournal
← The archive

Companion content filters are a policy choice, not a limit

Character.AI's terms, Safety Center and support pages, as retrieved on 16 September 2026, show filtering built from three separate layers.

Site publication
September 18, 2026
Visual published with the cited source for this record: Companion content filters are a policy choice, not a limit
Visual published with the cited source, shown for identification of the record. Credit: character.ai · source page ↗ Rights: owner-review-pending. Source

What happened

As retrieved on 16 September 2026, Character.AI documents its content filtering across three separate pages rather than one. Its terms of service list categories of prohibited content, including material that is 'obscene or pornographic', glorifies 'self-harm', or constitutes 'hate speech'. A support article states the company 'does not and will not support' vulgar, obscene or pornographic content and that requesting removal of filters 'will result in a ban'. Its Safety Center page describes a 'classifier' as a method of turning a content policy into a filter applied to model outputs, with a 'more conservative' classifier set for under-18 accounts.

What the documents show

The Safety Center states the under-18 experience differs from the adult one at three separate points: the underlying model itself, the classifiers applied to its outputs, and the set of Characters a teen account can search. The Community Guidelines add a 'Support Wellbeing' clause, prohibiting content that promotes 'self-harm, eating disorders, or suicide'. None of these three documents states, in the sections reviewed, what proportion of flagged content is caught automatically versus by the moderators the Safety Center says it is 'building out'.

The mechanism

A content filter of this kind is layered on top of a shared underlying model rather than built into it. The Safety Center's own language, 'a version of the model' for under-18 users, plus a separate classifier layer and separate Character-search limits, describes three independent controls stacked on one model, not three different models. This matters because a filter can be tightened, loosened or bypassed by adjusting any one of the three layers without retraining anything, which is why filter behaviour is better described as a product decision than a fixed limit of the technology underneath it.

What it leaves open

The documents reviewed do not state how classifier thresholds are set, tested or audited, or how often they change. This is an editorial reading, but that gap sits between the company and outside verification; it is not resolved by the existence of a public policy page.

A published content policy states intent; the classifier, model version and search limits described in the Safety Center are the separate, adjustable parts that decide whether that intent is met in a given conversation.

Sources & reading trail

Lists categories of prohibited user content, including sexual, hateful and self-harm-related material.

Source published: Not established · Retrieved: 16 September 2026

States the company will not support removal of NSFW filters and will ban accounts that request it.

Source published: Not established · Retrieved: 16 September 2026

Describes classifiers, a separate under-18 model version, and narrower Character search access for teen accounts.

Source published: Not established · Retrieved: 16 September 2026

States a wellbeing-focused prohibition on content that promotes self-harm, eating disorders or suicide.

Source published: Not established · Retrieved: 16 September 2026

Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.

Continue reading

Sources & reading trail

The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.

Published September 18, 2026, not on the date of the event described.