Voice & avatars

Source/event record · Guide · prepared 19 September 2026

A voice companion needs a text door, a repeat button and time to answer

Accessibility testing should cover modality switching, transcripts, correction and pacing—not just whether speech recognition works in a quiet room.

Prepared for local review · Site publication: not set · 463 words

Natural voice can make a companion feel effortless, but voice-only design transfers effort to anyone who is deaf, hard of hearing, speech-disabled, fatigued, processing language slowly or simply standing in a noisy station. An accessibility audit asks whether the same relationship remains usable when speech is not the easiest channel.

Test input and output independently

Build a four-cell matrix: speech in/speech out, speech in/text out, text in/speech out, and text in/text out. The W3C’s draft Natural Language Interface Accessibility User Requirements calls for multiple methods, synchronized text with spoken output, text input alongside speech output, and the ability to switch input methods during a dialogue. Treat that document as design guidance, not a conformance badge.

Probe recovery, not perfect recognition

Use a harmless sentence containing a name, a number and an uncommon word. Try it at normal pace, slow pace and with ordinary background noise. When recognition is wrong, can the user edit the transcript before the companion acts? Can they say or tap “repeat”? Does the interface expose what it heard? A confident but invisible mishearing is more consequential than a visible low-confidence request for confirmation.

Inspect pacing and control

Check whether speech rate, volume and captions can be adjusted; whether the companion interrupts; whether the user can pause output; and whether silence is mistaken for abandonment. The US Access Board’s ICT requirements are written for covered federal and telecommunications contexts, not every consumer app, but their requirements that speech output be coordinated with the display and be repeatable and pausable offer a strong test lens.

Run a task-based audit

Hypothetically, ask the companion to summarize a three-step plan, correct the second step, then confirm before saving it. Complete the task using each matrix cell. Record completion, errors, forced modality changes and time pressure. Invite disabled testers when possible and compensate them; a non-disabled reviewer using earplugs does not reproduce lived experience.

Publish barriers precisely

“Has captions” is weaker than “live transcript remains scrollable, can be corrected before send, and was checked on version X.” Also record language, device and assistive technology. Do not turn one tester’s success into a universal accessibility claim.

Turn barriers into release decisions

If text input disappears during a voice call, the product is not fully usable for someone who must switch modalities mid-conversation. If captions lag but remain accurate, document the lag and prioritize it according to the task. If a destructive action cannot be reviewed in text before confirmation, block that flow from voice-only release. These are specific consequences tied to observed barriers. “Needs accessibility work” is too vague to guide a builder or warn a reader.

Voice presence is meaningful only when control survives the performance. Continue with the latency field test for timing and the consent ledger for the identity behind a voice.

Sources & reading trail