Building responsibly

Source/event record · Guide · prepared 19 September 2026

Do not ship a companion update until someone can argue for rollback

A lightweight safety case links the change, evidence, uncertainties, monitoring and stop conditions to a rehearsed recovery path.

Prepared for local review · Site publication: not set · 464 words

Companion teams often test whether a new model is faster or more engaging. The release decision also needs a compact argument that the update is acceptably safe for its intended use—and a credible way back if production contradicts the test. This is a lightweight safety case, not a certificate.

State the claim narrowly

Write one sentence: “Version B is acceptable for this audience and feature because…” Name the changed components, users, languages and excluded uses. List evidence beneath the claim: regression results, privacy review, accessibility checks, moderation sampling and vendor change notes. NIST’s Generative AI Profile organizes risk work around governance, mapping, measurement and management; a safety case connects those activities to one release.

Record uncertainty and dissent

Include known gaps, small samples, untested languages, dependencies and any reviewer objection. Assign an owner to each open risk. A model that passes common prompts but has not been tested for long conversations should not inherit a blanket “safe” label. Decision-makers must see what the evidence does not cover.

Define stop conditions before launch

Choose observable triggers: cross-account memory exposure, a critical child-safety regression, an abnormal spike in crisis-response failures, unauthorized voice output or a sustained latency threshold. Set who can halt rollout, what data they need and whether partial containment is possible. A rollback trigger should not depend on the team first agreeing why the failure occurred.

Rehearse the path

In staging, switch back to the prior model or configuration, verify memory-schema compatibility, confirm old safety settings, and time the operation. Test user communication and support scripts. If rollback would lose new memories or strand conversations, state that cost in the release decision. NIST’s AI RMF playbook resources stress documentation and risk treatment across the lifecycle; recovery belongs in that lifecycle.

Use a hypothetical decision

Version B improves response speed but changes retrieval. Tests pass except two of fifty correction cases revive stale facts. The team limits rollout to five percent, blocks the affected memory path, monitors correction complaints and sets any cross-user recall as an immediate stop. After a staged rollback drill succeeds, the named owner approves the bounded release. That is a decision under uncertainty, not proof of harmlessness.

Assign an independent challenger

Someone other than the feature owner should test the argument, ask what evidence would reverse the decision, and verify that business pressure has not weakened the stop conditions. The challenger need not veto every release; they need authority to record unresolved disagreement and escalate critical gaps. If nobody can state what would make the team roll back, the safety case is advocacy rather than a decision instrument.

After launch, preserve the evidence and update the case with real outcomes. Publish the user-facing portion through the change-log discipline; if a stop condition fires, document it with the incident postmortem.

Sources & reading trail