RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
Lovebot Journaljournal
← The archive

A written constitution trained a model, not just labels

Anthropic's Constitutional AI paper trains behaviour from stated principles, distinct from a persona prompt layered on top.

Historical event
December 15, 2022
First source published
December 15, 2022
Site publication
September 18, 2026
Visual for this record: A written constitution trained a model, not just labels
Visual published by cdn.sanity.io, shown for identification of the record. Credit: cdn.sanity.io · source page ↗ Rights: owner-review-pending. Source

What happened

On 15 December 2022, Anthropic researchers posted a paper, Constitutional AI: Harmlessness from AI Feedback, describing a method for shaping a model's behaviour using a written list of principles rather than relying only on human labels of which outputs are harmful. Anthropic's own research page, published the same day, summarises the method in similar terms and states that "the only human oversight is provided through a list of rules or principles."

What the documents show

Both documents describe a two-stage process. In the first, supervised stage, the model is prompted to critique and revise its own earlier responses against the stated principles, and it is then fine-tuned on the improved, self-revised answers. In the second, reinforcement stage, an AI model, rather than a human, judges which of two candidate responses better fits the principles, and those AI-generated preference judgments train a reward model, a step the paper names reinforcement learning from AI feedback, or RLAIF. The stated outcome, in the paper's own words, is "a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them," achieved with far fewer human-written labels than a purely human-feedback pipeline would need.

The mechanism

The mechanism is training the model's weights against a fixed set of written principles, a different thing from writing a persona description into a prompt at the time of use. A persona prompt, of the kind Persona-Chat popularised, tells an already-trained model what to pretend to be for this conversation, and can be overridden, ignored or forgotten as the context changes. Constitutional training instead uses the principles during the training process itself, to generate the self-critiques and the AI preference judgments that adjust the model's weights, so the resulting tendencies persist across conversations rather than depending on a prompt being present. A "character" produced this way is closer to a trained disposition; a character produced by a system prompt is closer to an instruction the model happens to be following for now.

What it leaves open

The paper is Anthropic's own account of its own method and evaluates it primarily on the specific target of harmlessness combined with helpfulness, not on companion-style traits such as warmth, memory or long-term consistency, and it does not claim the technique guarantees a fixed personality across all future interactions. Whether any named companion product uses this or a related AI-feedback method is not addressed here and is not asserted by this record.

The distinction is not academic: a trained disposition and a prompted persona fail differently, and knowing which mechanism is in play changes what a reasonable person should expect a character to resist.

Sources & reading trail

Describes the two-stage supervised and RLAIF process trained against written principles.

Source published: 15 December 2022 · Retrieved: 16 September 2026

Anthropic's own summary stating the only human oversight is a list of rules or principles.

Source published: 15 December 2022 · Retrieved: 16 September 2026

Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.

Continue reading

Sources & reading trail

The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.

Published September 18, 2026, not on the date of the event described.