Building responsibly

Source/event record · Guide · prepared 19 September 2026

A report button is the front door of moderation, not the system

Character platforms need intake, severity, evidence preservation, response targets, appeals and feedback loops that creators can actually operate.

Prepared for local review · Site publication: not set · 465 words

User-made characters combine profile text, model output, images and private conversations. A generic abuse inbox cannot tell an operator what is urgent, what evidence is needed or whether a fix worked. Moderation becomes operational only when a report moves through a defined queue with accountable decisions.

Design intake around the object

Let a reporter identify the character, public listing, generated message, creator behavior or direct user interaction at issue. Capture a permalink or immutable reference, product version, time, locale and whether a minor may be involved. Offer free text without requiring a user to relive harm. Google Play’s UGC policy expects in-app reporting and blocking plus ongoing moderation; Apple’s review guideline similarly calls for filtering, reporting, timely responses, blocking and contact information.

Triage by harm and reach

Use a small severity model. Priority zero covers imminent threats or legally mandated emergency pathways. Priority one includes suspected child sexual exploitation, non-consensual intimate imagery or credible targeted threats. Lower levels cover harassment, impersonation, mislabeled mature content and ordinary policy disputes. Add reach and recurrence: a discovery-feed character used by thousands deserves faster containment than an unseen draft, even when the underlying category matches.

Separate containment from judgment

Temporary delisting, generation limits or preservation holds can reduce exposure while review continues. They are not final findings. Give moderators a policy citation, evidence checklist and permitted actions. Protect the reporter’s identity from the creator unless disclosure is necessary and lawful. For difficult cases, require a second reviewer and log disagreements.

Close the loop

Tell reporters that a review occurred and give creators a reason category plus an appeal route, subject to safety and privacy limits. Measure median response time, oldest open case, reversal rate, repeat-offender rate and prevalence sampling—not just reports closed. Ofcom’s moderation research emphasizes identification, action and tracking; its later sector summary noted that effectiveness data remained scarce even where moderation systems existed.

Test the queue before crisis

Run a tabletop with a hypothetical popular character whose public greeting is safe but whose generated replies solicit private contact from teenagers. Trace report, evidence capture, delisting, model-side mitigation, creator notice, appeal and re-test. Record where ownership is ambiguous.

Plan staffing from the queue

Sample arrival volume by category, language and hour, then estimate reviewer time and escalation demand. A 24-hour target is meaningless if one trained reviewer covers several languages and weekends. Decide what can be automated for routing, never assume automation makes the final sensitive judgment, and create an overflow rule for sudden events. If backlog exceeds the safe limit, reduce distribution or creation features rather than quietly allowing the oldest serious reports to age.

A queue is credible when the team can explain a decision and verify remediation. Connect it to the incident postmortem for systemic failures and the marketplace governance guide for rules before upload.

Sources & reading trail