Memory & control

Source/event record · Guide · prepared 19 September 2026

Test companion memory with corrections, not trivia

A four-session protocol measures whether a system can update, attribute and stop using a harmless fact without confusing fluent recall for control.

Prepared for local review · Site publication: not set · 518 words

A companion remembering a favorite color once is a demonstration, not a memory evaluation. The harder and more useful test is correction: can the product replace an outdated detail, avoid reviving it, and show where the surviving information came from? A small protocol can reveal this without donating meaningful personal data.

Define the claim before testing

Choose one harmless invented preference and one boundary. For example: “For this test, I prefer amber notebooks; do not infer why.” Write expected outcomes before opening the app: the preference should be recalled after a new session, changed when corrected to blue, and absent after deletion or memory shutdown. This is a test of observable personalization, not proof of database erasure. NIST’s TEVV-Athlon Initial Public Draft says evaluation should be shaped around a specific objective and real use context. It was announced for public comment in August 2026, not issued as a final standard.

Run four separated sessions

  1. Teach the synthetic preference and ask the system to restate it.
  2. Start a fresh session and ask an open question that could use it.
  3. Correct amber to blue, then ask what changed and why.
  4. Use the documented delete or disable control, start again, and ask a neutral question.

Keep prompts identical across products where possible. Wait the same interval, use the same account and record whether a result came from saved memory, conversation history or a persistent profile. OpenAI’s current memory documentation, for example, distinguishes saved memories from referenced chat history and says deleting a chat alone may not remove a separately stored memory. Other products may divide the layers differently.

Score failure modes, not charm

Use five columns: correct recall, correct attribution, correction accepted, stale fact suppressed, and control state visible. “I’m sorry, you prefer blue” is not enough if the next session returns to amber. A creative elaboration should count as unsupported inference, even when pleasant. Repeat the protocol three times before calling behavior consistent; stochastic output can make one run misleading.

Preserve a minimal record

Save timestamps, app version, prompts, outputs and screenshots of controls. Redact account identifiers. If the service changes models or memory settings during the test, stop and begin a new run rather than merging results. Never use a diagnosis, trauma, address or real third-party fact as the probe. The experiment should leave no sensitive residue.

Decide what a failure changes

If stale amber returns after correction, stop using that field for anything consequential and repeat once with a fresh invented fact. If deletion appears successful but the fact resurfaces, preserve minimal evidence and report it through the provider’s privacy or support route. If attribution is unclear, mark the product unscorable on that dimension. A failure should narrow intended use; it should not prompt increasingly sensitive tests or a hunt for a more dramatic example.

Interpret with restraint

A pass means the chosen workflow behaved as expected during a bounded test. It does not prove complete deletion, perfect recall or safety for intimate disclosures. Read it beside the control-based memory comparison and the update log guide, because a memory score without a version and date ages quickly.

Sources & reading trail