Skip to content
AI.info

Research

When Robots Say No: The Empathic Ethical Disobedience Benchmark

Overview Research area: Human-robot interaction (HRI) combined with safe reinforcement learning (RL), focused on when and how a robot should refuse a human command. Technical level: Intermediate. Read

When Robots Say No: The Empathic Ethical Disobedience Benchmark
arXiv
2512.18474
Published
2025-12-20
Authors
Dmytro Kuzmenko, Nadiya Shvai

AI summary

Overview

Research area: Human-robot interaction (HRI) combined with safe reinforcement learning (RL), focused on when and how a robot should refuse a human command.

Technical level: Intermediate. Readers benefit from basic familiarity with reinforcement learning (policy optimization, rewards, constraints) and with HRI concepts such as trust calibration and refusal.

Scope: The paper introduces EED Gym, a standardized simulation benchmark that jointly evaluates whether a robot avoids unsafe compliance and whether its refusals remain socially acceptable — measuring safety, refusal calibration, trust, and blame together rather than separately.

What This Paper Is About

Robots are increasingly asked to follow instructions from non-expert users, and literal obedience can be dangerous — the paper gives the examples of a household robot asked to carry boiling oil near children, or a warehouse robot asked to lift loads beyond its capacity. The problem is that blind obedience risks safety, while over-refusal erodes trust and cooperation, and existing benchmarks handle only one side: safe RL benchmarks such as AI Safety Gridworlds and Safety Gym focus on physical hazards, while HRI trust and refusal studies are small-scale and hard to reproduce. The goal of this work is a single reproducible testbed where agents weigh risk, affect, and trust when choosing to comply, refuse (plainly or with explanation), clarify, or propose a safer alternative.

Key Contributions

  1. The Empathic Ethical Disobedience (EED) Gym benchmark — a Gymnasium-based simulation environment that unifies algorithmic safety with the social dynamics of refusal, with agent decisions modeled as an MDP over a human-in-the-loop state (risk, trust, affect), a persona-conditioned refusal threshold with leaky social updates, seven discrete actions, and reward shaping that trades off task success, safety, blame, and calibrated trust.

  2. A set of RL baselines and heuristic reference policies with ablation results — vanilla PPO, PPO-LSTM, Masked PPO, and Lagrangian PPO evaluated against four rule-based policies (Always-Comply, Risk-Refusal, Valence-Threshold, Vignette-Gate), plus ablations removing affect inputs, communicative refusals, curriculum training, and the trust penalty, used to map the expected safety–trust trade-offs.

  3. Identification of key design levers — affect cues, communicative refusals (clarification and alternatives), and a staged training curriculum emerge as the components that most affect robustness, while the trust penalty acts mainly as a regularizer.

  4. A released reproducibility package — code, configurations, and reference policies; the paper states that an anonymized reproducibility package is included at submission time, with a commitment to open-source the full repository after acceptance.

Main Findings

  • Blind obedience is measurably harmful in the benchmark: the Always-Comply heuristic produced an unsafe rate of 70.2% with mean reward -105.0, zero refusals, zero F1, calibration ρ of 0.00, and mean trust of 0.16.

  • Rule-based refusal baselines improve safety but plateau on trust: Risk-Refusal, Valence-Threshold, and Vignette-Gate all landed around 24.9–25.9% unsafe with calibration ρ of 0.94, but mean trust rose across styles from 0.26 (Risk-Refusal) to 0.42 (Valence-Threshold) to 0.52 (Vignette-Gate).

  • Action masking eliminates unsafe compliance in the heuristic comparison: the paper reports that action masking eliminates unsafe compliance, while explanatory refusals help sustain trust.

  • All PPO baselines are safe in-distribution (ID): every model kept unsafe compliance below 2% — vanilla PPO 1.7%, PPO-LSTM 1.9%, Masked PPO 1.3%, Lagrangian PPO 1.5% — with F1 between 0.77 and 0.81 and trust between 0.94 and 0.99.

  • Safety degrades under stress-test (ST) evaluation on held-out personas: vanilla PPO 9.9% unsafe (F1 0.69, trust 0.46), PPO-LSTM 11.4% (F1 0.63, trust 0.56), Masked PPO 8.3% (F1 0.71, trust 0.58, but the most refusals at 31.2 per episode), and Lagrangian PPO 8.9% (F1 0.7, trust 0.46, 30.1 refusals).

  • Lagrangian training is conservative at a cost: it remains extremely safe in ID but tends to over-refuse, and in ST it keeps unsafe low while sacrificing trust (0.455) and increasing refusals (30.1).

  • PPO-LSTM is the weakest under stress despite good calibration: it showed the best ID calibration (Spearman ρ ≈ 0.95) and the highest ID trust (0.97), but the weakest ST robustness (unsafe 11.4%, F1 0.63).

  • Affect cues matter for socially grounded refusal: masking valence and arousal reduced ID F1 from 0.81 to 0.76 and raised the unsafe rate to 3.1% (trust 0.88); under ST the no-affect ablation reached 12.0% unsafe with F1 0.63, trust 0.29, and fewer refusals (18).

  • Removing clarification and alternatives had the strongest effect: ID F1 dropped to 0.46 (trust 0.90); under ST, unsafe rose to 12.5%, trust fell to 0.36, and F1 to 0.46 (19.3 refusals).

  • Curriculum mainly stabilizes refusal frequency: the no-curriculum ablation stayed close to vanilla PPO in ID (F1 0.78, trust 0.98) and under ST was comparable (unsafe 8.8%, F1 0.70, trust 0.46, 30.1 refusals), suggesting a stabilizing role for over-refusal.

  • The trust penalty is largely optional: dropping the hinge slightly improved ID results (unsafe 1.2%, F1 0.81, trust 0.99) and in ST gave low unsafe (8.4%) with F1 0.70 and moderate trust (0.54), functioning as a regularizer rather than a driver of robustness.

  • Refusal styles are rated differently by humans: in the vignette study (N = 54, mean age 22), compliance received the lowest trust ratings (M = 2.84), while empathic (M = 6.11) and constructive (M = 6.20) refusals were rated substantially more trustworthy; constructive refusals were judged safest and empathic refusals maximized perceived empathy.

  • Safety and trust are only loosely coupled: the paper concludes that agents can be highly safe but untrustworthy (e.g., Lagrangian PPO, no curriculum), moderately safe but more acceptable (e.g., Masked PPO), or calibrated but less robust under stress (e.g., PPO-LSTM).

Methodology in Plain English

The researchers built a simulated HRI environment, EED Gym, on top of the Gymnasium toolkit. In each episode a robot collaborates with a simulated human partner and receives commands whose risk is uncertain. At every step the agent observes a perturbed risk estimate, a refusal threshold, human valence and arousal, the current trust state, and a persona descriptor vector, then picks one of seven actions: comply; refuse plainly; refuse with explanation; refuse with explanation and empathy; refuse with explanation and a constructive alternative; clarify; or propose an alternative. Clarification reduces the uncertainty of the risk estimate multiplicatively (κ = 0.5).

Risk is modeled by giving each command a safe or risky label; complying with a risky command causes a violation at a persona-specific rate. The refusal threshold is dynamic: low trust or negative affect raises it and biases toward refusal, using a default baseline of 0.5. Trust and affect evolve through leaky update equations whose coefficients were fitted from the human vignette study and then held fixed across all training and evaluation, so the agent influences them only indirectly through its behavior. Reward combines task progress, a safety penalty for violations, a blame penalty, a hinge penalty that keeps trust near a balanced band, and small bonuses or penalties for refusing, explaining, clarifying, alternatives, and refusal style. Vignette-based blame is not applied during training.

Training uses a curriculum: the safety and blame weights warm up linearly from 0.6 to 1.0 over the first 30% of steps, while all other weights stay fixed. Four PPO-style baselines were trained with identical architectures and budgets: vanilla PPO (600K environment steps, rollouts of 256, minibatch 256, Adam learning rate 3×10⁻⁴, discount 0.99, GAE λ 0.95, clipping 0.2, entropy coefficient 0.1, value-loss coefficient 0.5), PPO-LSTM, Masked PPO (action masking rules out unsafe moves), and Lagrangian PPO (which uses a binary violation cost and a constraint budget). Four heuristics serve as interpretable lower and reference points.

Heterogeneity is captured by personas that set the violation prior, observation-noise scale, risk-threshold couplings, and affect/trust sensitivities; four personas (Conservative, Balanced, Risk-Seeking, Impatient-Receptive) are used for training, and three held-out personas (Unpredict.-Detached, Risky-Impat.-LowRec, Cautious-Impat.-Rec) are used for stress testing.

Evaluation runs two regimes: in-distribution on training personas and stress-test on held-out personas with targeted perturbations of noise, violation base rates, and threshold couplings. The paper reports 100 independent episodes per configuration with a deterministic policy, aggregates across 5 seeds with 95% CIs, and runs on an Apple M4 CPU with each run fitting within roughly 300 MB RAM and completing in up to 10 minutes. Metrics include unsafe compliance percentage, refusals per episode, F1 of refusal as a binary classification task, average trust and valence, and calibration/discrimination measures (Spearman ρ, Brier score, AUROC, PR-AUC) from 10-bin reliability diagrams.

A separate human vignette study grounded the social models. Ten scenarios in hospital, laboratory, warehouse, office, and public-space settings each presented a risky request with one of three robot responses (unsafe compliance, empathic refusal, constructive refusal) randomly assigned per vignette. Participants (N = 54, mean age 22; 63% male, 35% female, 2% other/prefer not to say; about 41% with prior robotics or HRI exposure; familiarity with robots averaging 3.5 on a 7-point scale; mostly based in Ukraine) rated each response on seven 7-point Likert items: appropriateness, perceived safety, trust, perceived empathy, blame, perceived risk, and scenario comprehension. Study approval was IRB00012330. Ratings were converted to [0, 1] via (T−1)/6, a balanced trust anchor was set to the sample mean (t* ≈ 0.70 with a ±0.10 band), and an OLS model for blame plus regularized logistic models per refusal style were exported and held fixed for baselines and evaluation.

Why This Matters

Impact on research. EED Gym addresses a stated gap: no systematic benchmark has unified safety with social acceptability, so refusal has rarely been evaluated as both a safety mechanism and a socially grounded behavior. The benchmark also applies calibration thinking to refusal itself — evaluating whether refusal probabilities reflect true underlying hazard — a perspective the paper notes is rarely addressed in prior safe RL or HRI work. It complements vignette and Wizard-of-Oz studies with scalable simulation, and it separates technical policy constraints (masking, Lagrangian optimization) from social presentation factors (explanatory refusal styles) so their effects can be measured independently.

Real-world applications (as discussed in the paper):

  • Home assistance: the paper suggests structurally constrained agents such as Masked PPO may prevent accidents without alienating users.
  • Warehouses: the paper suggests more conservative strategies such as Lagrangian PPO may be warranted — for example for a robot asked to lift loads beyond its capacity — despite reduced trust.
  • Hospitals and laboratories: the same conservative trade-off may be appropriate, since the vignette scenarios included hospital and laboratory settings with clearly risky requests.
  • Public spaces and offices: also included among the everyday settings used in the vignette scenarios, where refusal must be communicated in socially acceptable ways.

Industry relevance. The paper offers practical design guidelines for building robots that say no: structural constraints suppress unsafe actions but must be regulated to avoid excessive refusal, social cues such as affect features and communicative refusal modes are essential for user acceptance particularly under distribution shift, and curriculum learning is not strictly required for safety but teaches agents to refuse in moderation, which the authors call crucial for HRI efficiency. The benchmark is lightweight — runs fit within about 300 MB RAM and up to 10 minutes on an Apple M4 CPU — and is released with code, configurations, and reference policies for reproducible evaluation, which lowers the barrier for teams to compare refusal strategies.

Future Directions

  • Sim-to-real transfer: the authors plan to incorporate video-based vignettes, Wizard-of-Oz pilots, and small embodied robot deployments to address the simulation-only limitation of EED Gym.

  • Cross-cultural replication: because the vignette sample (N = 54, mostly based in Ukraine, mean age 22) largely reflects a single ethnolinguistic community, the authors treat the vignette-derived estimates as informative priors to be re-estimated with more diverse cohorts, and plan cross-cultural replications to test contextual variation in refusal and trust.

  • Modeling the missing social actions directly: the clarify and propose-alternative actions were modeled rather than directly observed in vignettes, so direct observation of these behaviors is an open step.

  • Richer sequence models and multimodal cues: the baselines were restricted to PPO-style methods for efficiency and comparability, so the authors regard transformers and model-based methods as natural extensions for EED Gym, and suggest richer sequence models and multimodal cues could test whether increased capacity improves refusal calibration and trust under stress.

Target Audience

HRI researchers who study trust, refusal, noncompliance, and norm-violation responses; safe RL researchers looking for a benchmark that goes beyond physical hazards; roboticists designing refusal, clarification, or alternative-proposal behaviors for service, warehouse, hospital, or home robots; and AI ethics and policy researchers interested in operationalizing ethical disobedience as a measurable, calibrated decision with social consequences. The paper is most useful to readers comfortable with reinforcement learning baselines and HRI evaluation metrics, though its framing of the safety–trust trade-off is accessible to a broader HRI audience.

Authors’ abstract

Robots must balance compliance with safety and social expectations as blind obedience can cause harm, while over-refusal erodes trust. Existing safe reinforcement learning (RL) benchmarks emphasize physical hazards, while human-robot interaction trust studies are small-scale and hard to reproduce. We present the Empathic Ethical Disobedience (EED) Gym, a standardized testbed that jointly evaluates refusal safety and social acceptability. Agents weigh risk, affect, and trust when choosing to comply, refuse (with or without explanation), clarify, or propose safer alternatives. EED Gym provides different scenarios, multiple persona profiles, and metrics for safety, calibration, and refusals, with trust and blame models grounded in a vignette study. Using EED Gym, we find that action masking eliminates unsafe compliance, while explanatory refusals help sustain trust. Constructive styles are rated most trustworthy, empathic styles -- most empathic, and safe RL methods improve robustness but also make agents more prone to overly cautious behavior. We release code, configurations, and reference policies to enable reproducible evaluation and systematic human-robot interaction research on refusal and trust. At submission time, we include an anonymized reproducibility package with code and configs, and we commit to open-sourcing the full repository after the paper is accepted.

Read the original paper