Building responsibly

Source/event record · Guide · prepared 19 September 2026

Write the companion incident postmortem around user impact, not embarrassment

A blameless record should reconstruct detection, exposure, containment, communication, recovery and prevention without publishing private conversations.

Prepared for local review · Site publication: not set · 496 words

When a companion exposes private content, gives a dangerous response or changes behavior unexpectedly, the first public impulse is often reassurance. The better first internal impulse is preservation: what happened, who may be affected, what changed, and what can be stopped without destroying the evidence needed to learn.

Open an incident record early

Assign an owner, severity, detection time and decision log. Preserve relevant model, prompt, policy, retrieval and deployment versions. Save the minimum conversation evidence necessary, redact identities, restrict access and document every copy. CISA’s logging guidance recommends deciding what to log, protecting logs, monitoring high-risk events and assigning incident roles; companion teams must also avoid turning intimate chats into an unlimited forensic archive.

Build the timeline from facts

Separate first occurrence, first report, detection, escalation, containment, user notice, recovery and closure. Label estimates. Include what the team knew at each decision point rather than rewriting history with later knowledge. If a vendor outage or model update contributed, cite evidence and keep responsibility for the product decisions your team controlled.

Measure exposure carefully

Count confirmed affected users, potentially affected accounts, duration, surfaces and regions. Do not equate messages scanned with people harmed. Describe qualitative impact categories—privacy, financial, emotional, child safety, identity or availability—without diagnosing users. NIST’s Generative AI Profile encourages incident disclosure, monitoring and feedback as part of lifecycle risk management.

Explain containment and recovery

Record why the team disabled a feature, rolled back a model, delisted a character or changed a filter. State what remains uncertain and what users can do now. A hypothetical memory leak might require suspending cross-account retrieval, notifying potentially affected users, rotating identifiers, validating isolation and offering deletion support. “Fixed” should require a test and monitoring window.

Turn causes into owned actions

Avoid ending at “human error.” Ask why review, tooling or incentives allowed the action. Each follow-up needs an owner, due date, verification method and public/private status. Publish a concise account when useful, but never include prompts or transcripts that expose people or make abuse easier.

Make communication a tested control

Prepare messages for affected users, all users, creators, support staff and regulators where applicable. Each audience needs confirmed facts, immediate actions, uncertainty and the next update time. Do not let the companion persona deliver a serious breach notice in character. In a drill, verify that support can identify the incident and that translations preserve meaning. Slow, vague communication can become a second failure even when technical containment succeeds.

Close the record only after each corrective action has evidence: a merged control, a completed test, a trained on-call owner or a monitored threshold. “Discussed with the team” is not verification. Schedule a later review to learn whether the same failure pattern returned under another name. Recurrence is evidence that the original cause or action was too narrow.

A postmortem is complete when prevention work can be checked, not when attention fades. Feed its cases into the regression suite and compare release decisions against the safety-case and rollback drill.

Sources & reading trail