Skip to content
AI.info

Research

Reward Learning through Ranking Mean Squared Error

Overview Research area: Reinforcement learning (RL) and reward learning from human feedback, specifically rating-based reward learning (learning reward functions from discrete, multi-class human ratin

Reward Learning through Ranking Mean Squared Error
arXiv
2601.09236
Published
2026-01-14
Authors
Chaitanya Kharyal, Calarina Muslimani, Matthew E. Taylor

AI summary

Overview

Research area: Reinforcement learning (RL) and reward learning from human feedback, specifically rating-based reward learning (learning reward functions from discrete, multi-class human ratings rather than binary preferences).

Technical level: Advanced. The paper combines a formal theoretical analysis (solution-set minimality and completeness proofs) with deep RL benchmarks and a human-user study.

Scope: This paper introduces R4 (Ranked Return Regression for RL), a rating-based RL algorithm built on a new ranking mean squared error (rMSE) loss, and evaluates it against rating- and preference-based baselines using simulated feedback and a human-user study.

What This Paper Is About

Designing reward functions by hand is a well-known bottleneck in applying RL to real-world problems, so researchers instead try to learn reward functions from human feedback. Most such work uses binary preferences (which behavior is better?), but preferences convey only one bit per comparison and say nothing about absolute quality. Recent work introduced rating-based RL (RbRL), where a human assigns a single trajectory a discrete rating such as "bad," "neutral," or "good." This paper's goal is to make better use of that ordinal rating structure by training reward functions with a ranking-based squared error loss instead of the cross-entropy-style loss used in prior rating-based work.

Key Contributions

  1. A new algorithm, R4. The paper proposes Ranked Return Regression for RL, a rating-based RL method that samples one trajectory per rating class, computes each trajectory's predicted return under the learned reward model, ranks those returns with a differentiable sorting operator (soft ranks), and minimizes a mean squared error between the resulting soft ranks and the human's ratings.

  2. The first rating-based RL objective with formal guarantees. The paper shows that the rMSE solution set is minimal and complete under mild assumptions, meaning it contains exactly the reward functions consistent with the ratings and no objective can shrink that set further without additional assumptions. A relaxed result extends this when the ranking operator is only approximately correct.

  3. Empirical validation in offline and online feedback settings. Using simulated ratings, R4 is compared against rating-based and preference-based RL algorithms across robotic benchmarks from OpenAI Gym and the DeepMind Control Suite.

  4. A human-user study plus an open-sourced dataset. An ethics-approved study with 8 participants providing ratings on two robotic tasks shows R4 outperforms RbRL despite large inter-participant variability, and participants report low cognitive workload. The collected rating dataset is released, along with code at https://github.com/IRLL/R4.

Main Findings

  • Offline gains over RbRL: With identical training conditions on OpenAI Gym Reacher, Inverted Double Pendulum, and Half Cheetah, R4-trained reward functions consistently produced better downstream Soft Actor-Critic (SAC) performance than RbRL-trained ones, yielding either statistically faster learning or higher final returns across all tested domains (p < 0.03).

  • Online gains across six control tasks: In DeepMind Control Suite environments (Walker-walk, Walker-stand, Cheetah-run, Quadruped-walk, Quadruped-run, Humanoid-stand), R4 matched or outperformed RbRL, PEBBLE, SURF, and QPA. The SAC agent learned significantly faster than all baselines in three of the six environments and achieved higher final returns in four of the six (p < 0.0125, Bonferroni-corrected threshold of α/4 = 0.0125).

  • Human-participant results: 8 participants (4 male, 4 female) each provided 100 ratings, choosing their own number of rating classes — which ranged from 4 to 12 across individuals. With five reward models trained per participant, R4 achieved significantly faster learning and higher returns in both Reacher and Hopper (p < 0.006). Comparing R4 trained on simulated versus human ratings showed only a small performance gap.

  • Human ratings are noisy and imbalanced: Ratings showed significant class imbalance (some participants assigned only a single trajectory to certain classes), behavior varied substantially between participants, and humans frequently assigned different ratings to trajectories with similar environment returns — unlike simulated ratings.

  • Low perceived workload: Mean NASA Task Load Index (TLX) scores were low on a 1 (low) to 7 (high) scale. Overall workload averaged 2.05 ± 0.53; other subscales were Mental Demand 2.29 ± 1.38, Physical Demand 1.57 ± 0.98, Hurried 1.71 ± 0.76, Effort 2.29 ± 0.76, Insecure 1.29 ± 0.49, and Success 4.86 ± 0.69.

  • rbRL's reward can fail to be a minimizer: Proposition 1 shows the human's true reward r* is always in the rMSE solution set, but is not guaranteed to be in the RbRL solution set, because the RbRL loss pulls predicted returns within a class toward class boundaries or their midpoint.

  • Theoretical guarantee: Theorem 1 establishes that the rMSE solution set equals the set of feasible reward functions consistent with the rating order, and Theorem 2 shows this still holds under a relaxed assumption where soft ranks deviate from hard ranks by at most ε, with 0 ≤ ε < (√(2n) − 2)/(n − 2) and n > 2.

  • Ablations: In two of three tested domains, R4 performed comparably without the dynamic feedback schedule and stratified sampling additions; using either technique alone was sufficient, and combining both gave small but consistent improvements. On the Trajectory Alignment Coefficient (TAC), R4's learned rewards were generally better aligned with the environment reward than the baselines. As the number of rating bins varied, R4 remained stable while RbRL degraded.

Methodology in Plain English

R4 treats a dataset of trajectory–rating pairs as an ordering problem. Ratings like "bad," "neutral," and "good" are ordinal — a "good" trajectory should have a higher return than a "bad" one under the true (unknown) human reward. To exploit this, R4 samples one trajectory from each rating class at each training step, feeds each trajectory's state-action pairs into the reward model, and computes a predicted discounted return. It then runs those predicted returns through a differentiable sorting routine (from Blondel et al., 2020), which produces continuous "soft ranks" instead of hard integer positions. The loss is simply the mean squared error between these soft ranks and the human's rating labels — R4's rMSE objective. Because soft rankings are differentiable, gradients flow back into the reward model.

The paper argues three advantages over RbRL's cross-entropy-style loss: no need to specify return-class decision boundaries; no forced collapse of all trajectories in a class toward a single midpoint value, which preserves within-class diversity; and support for dynamic rating schemes where classes can be added or merged during training.

For the online setting, the authors add three practical mechanisms: a dynamic feedback schedule that collects feedback more frequently early and less often later; stratified sampling over a maintained dataset of the 50 most recent trajectories, blending high-predicted-return and lower-predicted-return trajectories and extracting fixed-length sub-trajectories either randomly or by highest predicted return with equal probability; and dynamic rating classes that start fine-grained and merge as behavior improves, reflecting response shift and recalibration concepts from the measurement literature.

Evaluations follow standard practice: simulated ratings derived from environment-return thresholds for the offline Gym experiments, an SAC agent trained on the learned reward, five random seeds per condition, learning curves showing individual runs and means, and t-tests (α = 0.05; Bonferroni-corrected to 0.0125 for the four online comparisons per domain). The human study follows the offline protocol, with participants rating a fixed dataset of trajectory segments.

Why This Matters

Impact on research. The paper contributes the first rating-based RL objective with provable minimality and completeness, connecting reward learning from ordinal human feedback to the differentiable-ranking literature. It also adds to a documented gap: the authors cite a review of preference-based RL literature from 2012–2024 finding that less than 50% of proposed algorithms were validated with non-author human participants, and this study adds real human data and an open dataset.

Real-world applications (note: the paper's own evaluation is on robotic locomotion benchmarks, not deployed systems; these are areas the framing implies):

  • Robotic locomotion and control, where the paper's Gym and DeepMind Control Suite tasks are standard proxies.
  • Assistive and interactive robotics, where end users rather than RL experts may supply feedback.
  • Reward modeling for large language models, where recent work cited in the paper has moved beyond binary preferences to ordinal feedback.
  • Any RL deployment where hand-specifying a reward risks misspecification and unintended agent behavior.

Industry relevance. Rating interfaces may reduce the cognitive burden and human time needed for feedback compared to pairwise comparisons — the paper notes a fixed feedback budget requires a human to assess twice as many trajectories in preference-based methods, since each comparison involves two trajectories. Lower-effort, richer feedback is directly relevant to teams collecting human labels for RLHF-style pipelines, and the released dataset and code lower the barrier to adoption.

Future Directions

  • Extending the guarantees beyond the stated assumptions. The theory assumes a noise-free rating process, an unknown deterministic human reward function, and that the human's reward lies in the hypothesis class. Handling noisy or inconsistent human labels theoretically is left open.
  • Dynamic rating schemes. RbRL degrades when the number of bins deviates from an optimal range, while R4 tolerated varying bin counts in the paper's bin-number experiment. Understanding how best to add, merge, or split rating classes online remains open.
  • Reducing the gap between simulated and human feedback. The paper reports only a small performance gap between R4 trained on simulated versus human ratings; characterizing where that gap comes from and whether it grows in other domains is a natural next question.
  • Broader human evaluation. The study covers 8 participants across two OpenAI Gym tasks (Reacher with n = 7 and Hopper with n = 6), with considerable inter-participant variability. Larger, more diverse studies, especially with non-author participants, are an evident next step, as is testing whether the approach transfers to non-locomotion domains.

Target Audience

RL researchers working on reward learning and human-in-the-loop learning; practitioners building systems that learn objectives from human feedback, including RLHF-style LLM reward modeling; HCI and human-AI interaction researchers interested in the usability of ratings versus preferences as a feedback modality; and graduate students or advanced undergraduates comfortable with MDPs, policy optimization, and basic supervised learning theory.

Authors’ abstract

Reward design remains a significant bottleneck in applying reinforcement learning (RL) to real-world problems. A popular alternative is reward learning, where reward functions are inferred from human feedback rather than manually specified. Recent work has proposed learning reward functions from human ratings rather than traditional binary preferences, enabling richer and potentially less cognitively demanding supervision. Building on this paradigm, we introduce a new rating-based RL method, Ranked Return Regression for RL (R4). At its core, R4 uses a novel ranking mean squared error loss that learns from a dataset of trajectory-rating pairs, treating the human-provided discrete ratings (e.g., bad, neutral, good) as ordinal targets. Unlike prior rating-based approaches, R4 offers formal guarantees: its solution set is provably minimal and complete under mild assumptions. Empirically, using both human-provided and simulated ratings, we demonstrate that R4 consistently matches or outperforms existing rating and preference-based RL methods on robotic benchmarks from OpenAI Gym and the DeepMind Control Suite. Code released at https://github.com/IRLL/R4.

Read the original paper