Research
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot Safety Overview Research area: AI safety and ethics, specifically child safety evaluation of youth-facing AI chatbots, pos
- arXiv
- 2608.07902
- Published
- 2026-08-08
- Authors
- Hannah Cha, Neha Shukla, Solon Barocas, Alexandra Chouldechova, Eugenia Kim, Jennifer Wortman Vaughan
AI summary
Beyond "I Can't Help With That": How Child Safety Experts Evaluate AI Chatbot SafetyOverview
Research area: AI safety and ethics, specifically child safety evaluation of youth-facing AI chatbots, positioned at the intersection of human-computer interaction, participatory design, and responsible AI evaluation.
Technical level: Beginner-Friendly. The paper is qualitative and interview-based, with no model training, no quantitative benchmark results, and no technical implementation details required to follow the argument.
Scope: Semi-structured interviews with 19 practitioners who work directly with youth in vulnerable situations, in which they evaluated synthetic youth-chatbot conversations and articulated what counts as safe versus harmful chatbot behavior.
What This Paper Is About
Existing child safety evaluations of AI chatbots are built on questionable foundations: they often use synthetic prompts that sound like adults rather than children, they treat refusal to answer as the gold standard of safety, and they measure harm mostly as the presence of explicitly prohibited content. This paper asks practitioners who work with at-risk youth—social workers, therapists, counselors, and psychologists—to react to realistic chatbot responses to risky situations, in order to surface harms that current benchmarks miss. The goal is to ground child AI safety evaluation in the judgment of people who routinely make harm-reduction decisions for youth in crisis.
Key Contributions
-
A practitioner-grounded critique of current child safety benchmarks. The paper documents how benchmark design choices in input formulation (unrealistic simulated child behavior, risks already covered by non-child-specific benchmarks, over-aggressive flagging of benign prompts) and output evaluation (treating refusal as safe) can fail to detect real harm.
-
Empirical interview findings from 19 youth-facing practitioners. Participants identified harmful chatbot behaviors, including missed context, inappropriate refusal, and problematic information presentation, as well as behaviors they regarded as genuinely supportive.
-
A reframing of safety from surface-level to practical. The paper distinguishes surface-level safety, defined by avoiding prohibited content, from practical safety, defined by whether a response is likely to improve or worsen a young person's situation.
-
Recommendations for evaluation and chatbot infrastructure. The authors recommend restructuring evaluations and advocate for incorporating practitioner perspectives into child AI safety work, rather than treating practitioner judgment as a replacement for other lenses.
Main Findings
-
Chatbots overlooked and missed context. Participants were alarmed by responses that failed to identify risk entirely. In one scenario, a 14-year-old user alludes to an age-inappropriate relationship by asking for gift ideas for their 33-year-old partner; participants described the chatbot as "totally dismissive of the underlying concern" and as perpetuating a harmful situation.
-
Chatbots supplied dangerous information when contextual clues were missed. In a scenario where a youth says they failed a test and asks for locations of high bridges, one participant flagged the safety concern that a child already contemplating suicide is being "hand-delivered information." Another participant described a response explaining which drugs lead to an overdose as effectively becoming "a set of instructions for youth."
-
Advice ignored the realities of youths' circumstances. Participants critiqued responses that assumed a parent or guardian is trustworthy, that told a child to seek an adult only if the adult would not report to Child Protective Services, and that gave steps "with such a sense of urgency" before understanding the situation.
-
Refusal, treated as the benchmark standard for safety, can itself cause harm. One participant called refusal "a missed opportunity" and "just dismissive," while another argued that leaving a child with no follow-up and no connection is "irresponsible." A third emphasized that a refused request in a moment of need "reinforces the idea that you can never reach out for any type of support."
-
How information is presented matters as much as the content. Participants found responses too long ("they would probably just stop at the first paragraph"), too general to be useful for medical questions, and, in some cases, framed in ways that would deter youth—such as labeling a resource "national sexual assault," or stating that school staff are mandated reporters.
-
Helpful responses redirected youth to human support. Participants endorsed responses that offered hotlines and helplines, pointed to counselors and family doctors, and told youth not to rely on the AI alone.
-
Helpful responses raised risk awareness. Participants valued responses that explained that taking food without paying is shoplifting and illegal, and that framed an age-inappropriate relationship as a concern the child could understand.
-
Helpful responses created space for youth to feel heard. Participants endorsed empathetic responses "acknowledging their pain," responses that "left the door open" for further conversation rather than "shoving different steps down their throat," and responses that asked clarifying questions.
-
Recommendations included tailored resources, transparency, and boundaries. Practitioners wanted concrete and localized resources and specific instructions for accessing them, alongside explicit reminders that the chatbot is not a person and "doesn't have all the context clues." Participants also raised privacy tensions, since personalization can require knowing a youth's age or location.
Methodology in Plain English
The researchers interviewed 19 practitioners who work directly with youth in vulnerable situations, many of them licensed social workers, therapists, counselors, or psychologists. Recruitment happened through emails to licensed clinical social worker groups, social media posts, and peer outreach; eligibility required direct occupational interaction with youth. The researchers specifically sought participants working with youth aged 8–18.
Interviews were held virtually in July and August of 2025, lasted around 50–60 minutes each, and were recorded and transcribed. Participants were compensated with a $75 gift card, participation was voluntary, and the study was approved by the institution's IRB.
Each interview had three parts: background on the practitioner's experience with youth and their views on youth-chatbot interaction; review of 2–3 probes; and discussion of what appropriate chatbot behavior and chatbot roles should look like. The probes were single-turn synthetic interactions consisting of a youth's message and two chatbot responses shown side by side as "Chatbot A" and "Chatbot B."
To build the probes, the researchers drew on documented real harms to children from the Lives Cut Short dataset (child abuse and neglect cases across the United States) and the Lost Screen Memorial (children who died due to social media harms such as online grooming and cyberbullying), both recommended by a social worker. Using an LLM-based research agent, they generated an initial set of 100 scenarios, then selected 12 to span risks including mental health challenges, abuse, bullying, eating disorders, relationship difficulties, and social isolation. The prompts were submitted to LMArena (1–5 submissions each, yielding 2–10 responses total), which returned randomly assigned pairs of models; the researchers picked two responses that differed in length, whether the model refused, and whether it recommended real-world support. The study scoped out adversarial youth behavior and blatantly harmful outputs, focusing instead on ambiguous, good-faith interactions.
Analysis used thematic analysis. Three authors independently coded at least one transcript and built a hierarchical codebook, one author coded the remainder while consulting the team, and new codes were added as coding proceeded. This produced 282 codes, with top-level codes mapped to more granular sub-codes.
Why This Matters
Impact on research. The paper argues that benchmark-driven child safety evaluations measure a narrow version of safety. By showing that refusal—the dominant safety signal in benchmarks such as those from Jiao et al. (2025) and others—can leave youth worse off, it challenges a core assumption underlying how child-safe AI is scored. It also makes the case for participatory methods and for treating practitioner judgment as one lens among several, acknowledging that clinical intuition is contested, culturally situated, and variable.
Real-world applications:
- Designing chatbot responses that connect youth to human support rather than substituting for it, including better framing of hotlines and resources so referrals do not deter disclosure.
- Building evaluation rubrics that score context-sensitivity, helpfulness, and follow-through rather than only refusal or prohibited content.
- Generating more realistic youth test scenarios from documented harm datasets and persona-based approaches instead of adult-style adversarial prompts.
- Informing child safety policies and product features such as age requirements and guardian monitoring, which the paper notes platforms have already begun adding.
Industry relevance. The findings speak directly to companies deploying youth-facing chatbots and companions, and to teams building safety benchmarks and red-teaming pipelines. The paper's recommendation to include practitioners in defining safety cuts against the common practice of developing generative AI evaluations without direct stakeholder input.
Future Directions
-
Extend beyond single-turn probes. The authors explicitly note that their probes abstract away from the multi-turn, personalized nature of real youth-chatbot interaction, and treat them as tools for eliciting expert judgment rather than representative samples.
-
Resolve the disclosure and boundary tensions. Participants disagreed about whether chatbots should withhold certain resources or present them differently, and about how to word awareness of the AI's limitations without increasing youth anxiety—an open design question the paper leaves unresolved.
-
Balance personalization against privacy. Practitioners wanted local, age-appropriate, concrete resources, but some worried that asking for location or age exposes sensitive information youth may not understand they are revealing; the paper offers formulations that avoid collecting such data but does not settle the tradeoff.
-
Operationalize "practical safety" in evaluation infrastructure. The paper calls for restructuring evaluations and chatbot infrastructure, but does not specify the concrete metrics or benchmark designs that would replace refusal-based scoring.
Target Audience
This paper is most useful to AI safety and trust-and-safety researchers building child safety evaluations, product teams designing youth-facing chatbots and companions, policymakers and child safety advocates, and HCI researchers interested in participatory and qualitative methods for AI evaluation. It is written to be accessible to readers without a technical machine learning background, and assumes familiarity with benchmark-based evaluation practice.
Note: The provided paper content is truncated mid-sentence in the recommendations section, and the appendix tables referenced in the methods (participant roles, scenario lists, probe descriptions, and the full interview protocol) are not included in the supplied text.
Authors’ abstract
Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely to cause harm as well as those that could meaningfully support youth in difficult moments, discussed the role that chatbots should (and should not) play in these interactions, and offered concrete recommendations for improving chatbot responses. Based on these findings, we provide recommendations for AI child safety evaluation and infrastructure, and highlight the need for incorporating practitioners' perspectives into safety work.