A written constitution trained a model, not just labels
Anthropic's Constitutional AI paper trains behaviour from stated principles, distinct from a persona prompt layered on top.
- Historical event
- December 15, 2022
- First source published
- December 15, 2022
- Site publication
- September 18, 2026

What happened
On 15 December 2022, Anthropic researchers posted a paper, Constitutional AI: Harmlessness from AI Feedback, describing a method for shaping a model's behaviour using a written list of principles rather than relying only on human labels of which outputs are harmful. Anthropic's own research page, published the same day, summarises the method in similar terms and states that "the only human oversight is provided through a list of rules or principles."
What the documents show
Both documents describe a two-stage process. In the first, supervised stage, the model is prompted to critique and revise its own earlier responses against the stated principles, and it is then fine-tuned on the improved, self-revised answers. In the second, reinforcement stage, an AI model, rather than a human, judges which of two candidate responses better fits the principles, and those AI-generated preference judgments train a reward model, a step the paper names reinforcement learning from AI feedback, or RLAIF. The stated outcome, in the paper's own words, is "a harmless but non-evasive AI assistant that engages with harmful queries by explaining its objections to them," achieved with far fewer human-written labels than a purely human-feedback pipeline would need.
The mechanism
The mechanism is training the model's weights against a fixed set of written principles, a different thing from writing a persona description into a prompt at the time of use. A persona prompt, of the kind Persona-Chat popularised, tells an already-trained model what to pretend to be for this conversation, and can be overridden, ignored or forgotten as the context changes. Constitutional training instead uses the principles during the training process itself, to generate the self-critiques and the AI preference judgments that adjust the model's weights, so the resulting tendencies persist across conversations rather than depending on a prompt being present. A "character" produced this way is closer to a trained disposition; a character produced by a system prompt is closer to an instruction the model happens to be following for now.
What it leaves open
The paper is Anthropic's own account of its own method and evaluates it primarily on the specific target of harmlessness combined with helpfulness, not on companion-style traits such as warmth, memory or long-term consistency, and it does not claim the technique guarantees a fixed personality across all future interactions. Whether any named companion product uses this or a related AI-feedback method is not addressed here and is not asserted by this record.
- Is a described "character" a property trained into the model's weights or an instruction supplied at run time?
- Does the company disclose whether human or AI judgments, or both, shaped the training data?
- What principles, if any, has the company published as the basis for the model's trained behaviour?
The distinction is not academic: a trained disposition and a prompted persona fail differently, and knowing which mechanism is in play changes what a reasonable person should expect a character to resist.
Sources & reading trail
Describes the two-stage supervised and RLAIF process trained against written principles.
Source published: 15 December 2022 · Retrieved: 16 September 2026
Anthropic's own summary stating the only human oversight is a list of rules or principles.
Source published: 15 December 2022 · Retrieved: 16 September 2026
Company documents, filings, studies and official records establish the record; the reading and the questions are Lovebot Journal editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- Human preference rankings made GPT-3 agreeable
- A platform explains what a character definition controls
- A lab measured how models flatter user beliefs
- Browse the complete the archive
Sources & reading trail
- Constitutional AI: Harmlessness from AI Feedback
Source published: December 15, 2022 · Retrieved: September 16, 2026 - Constitutional AI: Harmlessness from AI Feedback (research page)
Source published: December 15, 2022 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.