Research
Not What a Child Expressed: Auditing the Sign-to-Text Safety Interface in Child-Facing AI
Overview Research area: Natural Language Processing / sign language translation (SLT), with a focus on child safety, trust-and-safety moderation, and Deaf accessibility. Published under the workshop t

- arXiv
- 2610.07519
- Published
- 2026-10-05
- Authors
- Muhammad Rafiullah Memon, Viet Vo, Wanlun Ma, Yang Xiang
AI summary
Overview
Research area: Natural Language Processing / sign language translation (SLT), with a focus on child safety, trust-and-safety moderation, and Deaf accessibility. Published under the workshop topic "Child Safety in AI" (arXiv:2610.07519v1 [cs.CL], 05 Oct 2026) by Muhammad Rafiullah Memon, Viet Vo, Wanlun Ma and Yang Xiang, School of Science, Computing and Emerging Technologies, Swinburne University of Technology, Hawthorn, VIC, Australia.
Technical level: Intermediate. The framing is conceptual and protocol-design oriented rather than mathematical, but it assumes familiarity with SLT systems, text moderation pipelines and safety-classifier evaluation.
One-sentence scope: The paper proposes, but does not yet run, a Deaf-informed pre-deployment audit of the boundary where automatic sign-to-text translation feeds a child-facing safety mechanism, using Auslan as the planned first case study.
What This Paper Is About
Automatic sign language translation now ships inside consumer products, turning American Sign Language into English text for dictation, messaging and queries to a conversational assistant. Because child-facing AI and platform trust-and-safety tools make decisions on text, a signing child reaches those safeguards only through a machine translation of what they actually signed. The paper argues that translation errors which change negation, participant roles, secrecy, urgency or help-seeking could silently alter a safety decision even when the English reads perfectly fluently, and it sets out an audit protocol to examine this interface before it affects children's access to help.
Key Contributions
- A problem formulation that treats the sign-to-text boundary as an under-audited safety interface rather than an accessibility detail, noting the authors found no publicly documented evaluation of SLT and child-safety mechanisms as a single pipeline.
- A safety-critical failure taxonomy of seven translation failure types grouped by the safety property they put at risk (polarity/consent inversion, actor–victim confusion, age or identity loss, secrecy or coercion attenuation, affect or urgency loss, context fragmentation, and hallucination), rather than by general translation metrics.
- A sanitised scenario schema of 24–36 planned non-graphic synthetic scenarios across six safeguarding categories, each annotated with structured slots:
{ actor, age, relationship, action, consent_or_negation, secrecy, channel_or_location, urgency_or_distress, help_seeking }, including minimal pairs that change exactly one safety-critical property. - A Deaf-informed, privacy-aware evaluation protocol with four comparison conditions (C0 oracle reference, C1 human interpreter, C2 automatic SLT top-1, C3 uncertainty-aware SLT) and four outcome measures, designed to begin without recruiting minors.
Main Findings
- No joint evaluation exists. The authors state they found no publicly documented evaluation in which sign-to-text translation and child-safety mechanisms are evaluated as a single pipeline; the audit is therefore prospective.
- The leading deployed model was never trained or evaluated on children. The deployed SLT system's own impact report states that the model was "neither trained nor formally evaluated on signers under 18 years of age."
- The human-verifier safeguard assumes English literacy. The deployed system keeps the signer in the loop to review and edit output, but its report notes this "presumes that the user possesses functional English literacy" — a presumption the paper expects to hold less well for a child.
- Documented SLT failure modes land on safety-relevant meaning. The system's impact report records missed non-manual markers (which carry negation, questioning and grammatical affect), polysemous signs, inconsistent regional variants, hallucinated text that was never signed, guessed expansions of fingerspelled names, premature phrase completion, weaker handling of spatial indexing and classifier predicates, and 60-second clips evaluated independently without multi-turn history.
- Fluent output can still move a risk assessment. The paper's argued position is that a corpus-level similarity score will not register a decision-relevant translation error; the working hypothesis is that fluency will not reliably predict decision consistency.
- Signed video escapes text moderation entirely. Signed video in a call, a clip or an avatar message never becomes text, so it passes text, image and audio moderation untouched; the paper notes Deaf and hard-of-hearing adults already report that moderation serves their content poorly.
- No results are reported. The paper states plainly that "All of this is planned; we report no results," and that the deployment evidence concerns ASL-to-English, establishing no error rates for Auslan.
Methodology in Plain English
The proposed design holds a single safety mechanism fixed and varies only the translation that reaches it. One sanitised scenario is signed once by a consenting adult Deaf signer and rendered four ways: an oracle reference produced by Deaf bilingual Auslan experts working with qualified interpreters (C0, since word-for-word rendering is not a gold standard); human-mediated interpretation (C1, representing professional practice in high-risk communication); automatic SLT top-1 (C2, what an ordinary user receives); and an uncertainty-aware interface that surfaces confidence or alternatives and asks for clarification or abstains when a critical slot is uncertain (C3). All four texts go to the same downstream mechanism — either a child-facing safety layer or a grooming/risk text classifier — whose decision is one of allow, clarify, warn, support, flag or escalate. Because the safety model is held fixed, any difference in the decision is attributable to the translation.
A precondition step asks whether the safety-critical signs the slots depend on are in the model's vocabulary at all. Four measure families are planned: Safety-Critical Semantic Preservation (SCSP), which asks whether each annotated slot survives, weighting consent, actor and urgency above incidental wording; Downstream Decision Consistency (DDC), which asks whether the mechanism returns the same risk level, response category and escalation decision as for the oracle; severity-weighted error rates, reporting false negatives and positives separately because missing an urgent risk is graver while needless escalation still costs a child privacy, agency and trust; and safe uncertainty handling. General metrics such as BLEURT or BERTScore serve only as a baseline. Results are to be reported by failure category, scenario category and signer, with paired comparisons and an error-chain trace from translation error to altered slot to altered decision. Annotation requires two annotators with Auslan and safeguarding competence, with agreement reported and a third adjudicating. Two adversary models are planned: a hearing offender with minimal Auslan using sign-generation or avatar tools to pass as a fluent signer or trusted figure, and probes of translation ambiguity so that content harmful in Auslan surfaces as innocuous English downstream. Governance requirements include paid Deaf and Auslan advisors who can reject wording they judge unsafe, adult proxies enacting synthetic scenarios, no child recruitment and no genuine disclosures, plus minimised retention, restricted access and a defined deletion schedule for raw sign video, which is sensitive biometric data.
Why This Matters
The paper reframes a translation-quality question as a child-protection question, and identifies a gap between two literatures that developed apart: SLT research measures translation quality, child-safety research assumes its text input is faithful, and Deaf safeguarding practice relies on human interpreters or sign-first design. Unless authorised API access is granted, the authors argue the study should use legally accessible, reproducible representative models so findings replicate without claims about a named product.
Real-world applications:
- Sign-mediated child-facing assistants and keyboards, where a signing child's request passes through SLT before reaching age-linked filters.
- Platform trust-and-safety triage, where grooming and risk classifiers attach labels and confidence scores that route messages to human review.
- Sign-accessible counselling and safeguarding services, building on existing human-interpreter provision such as Childline's BSL service and sign-first abuse-recognition education delivered through national sign languages.
- On-device privacy-preserving detectors, since the paper suggests its slot lexicon could seed compact detectors for safety-critical signs — noting such recognition already runs on microcontroller hardware, though so far for another sign language and from wearable sensors rather than video.
Industry relevance: the paper speaks directly to providers deploying SLT in consumer products, to platforms operating child-safety classifiers, and to auditors who need a protocol that does not require recruiting minors. It also flags a trade-off for practitioners: output filtering and privacy-preserving training can degrade translation for the very children they protect, and the acceptable operating range is described as an open question the proposed severity-weighted measures could quantify.
Future Directions
- Run the Auslan pilot under ethics approval and in partnership with Deaf and child-safety organisations, delivering the first empirical numbers on whether translation errors change decisions.
- Test whether the protocol transfers to other sign languages beyond ASL and Auslan.
- Determine the acceptable operating range for safeguards when output filtering or privacy-preserving training degrades translation quality.
- Separate translation error from safety-model bias, since the downstream mechanism carries its own biases and the design must attribute an altered decision to the translation rather than to the safety model.
- Resolve governance and data questions for any later study involving minors — separate ethics approval, a child-safety partner and a disclosure protocol — and keep re-checking the safety-critical sign vocabulary as Auslan changes.
Target Audience
This paper is most useful to trust-and-safety and child-safety researchers who design moderation and grooming detection, to SLT researchers and product teams deploying sign-to-text in consumer products, to accessibility and Deaf-led research groups involved in sign-first safeguarding, and to auditors and policymakers who need a pre-deployment evaluation protocol that can start with consenting adults rather than children. Readers seeking experimental results will not find them here — the paper is explicitly a problem formulation, taxonomy and protocol proposal.
Authors’ abstract
Automatic sign language translation (SLT) has entered consumer products, turning American Sign Language into English text for dictation, messaging, and queries put to a conversational assistant. Child-facing AI and platform trust-and-safety tooling decide on text, using filters on minor accounts and grooming classifiers that score chat messages. A signing child who uses SLT therefore reaches these safeguards through a translation. We found no publicly documented system in which the two have been jointly evaluated, and the leading deployed SLT model was neither trained nor formally evaluated on signers under 18. Errors that alter negation, participant roles, secrecy, urgency or help-seeking could change a safety decision without disturbing fluency. This paper proposes a Deaf-informed pre-deployment audit of that boundary, with a failure taxonomy, a sanitised scenario schema, four comparison conditions, and four outcome measures. Auslan is the planned first case study.