Research
Rubric Rewards from Item Response Theory
Overview Research area: Reinforcement learning for large language models, specifically reward design for rubric-based (non-verifiable) tasks, drawing on item response theory (IRT) from psychometrics.

- arXiv
- 2609.35646
- Published
- 2026-09-28
- Authors
- Milad Yazdani, Yaser Souri, Xiren Zhou, Pranit Chawla, Dena Shahriari, Subhojit Som, Xia Song
AI summary
Overview
- Research area: Reinforcement learning for large language models, specifically reward design for rubric-based (non-verifiable) tasks, drawing on item response theory (IRT) from psychometrics.
- Technical level: Advanced. The paper assumes familiarity with GRPO-style policy optimization, Bayesian inference, expectation maximization, and Fisher information.
- Scope: The paper introduces Rubric Response Theory (RRT), a method that replaces the additive rubric-point reward in GRPO with a maximum a posteriori (MAP) quality estimate from a two-parameter item response model, and uses the same model to select which rubric criteria to judge under a budget.
What This Paper Is About
Many language tasks have no single automatically checkable answer, so they are graded by rubrics — natural-language criteria plus assigned points — with a large language model as judge. The standard way to turn criteria verdicts into a training reward is to sum or average the points of satisfied criteria, which means two rollouts with different verdict patterns can get the same reward, and it still requires judging every criterion for every response. RRT instead treats each criterion as a test question with its own difficulty and discrimination, infers each rollout's latent quality from its full verdict pattern, and uses that quality as the reward — while also ranking unjudged criteria by how much information they would add.
Key Contributions
-
Rubric aggregation as Bayesian inference. RRT uses a two-parameter item response model where each criterion j has difficulty b_j and discrimination a_j, and a Response Parameter Network (RPN) ψ predicts (a_j, b_j) = ψ(q, c_j) from the prompt q and criterion text c_j, with a_j > 0. The scalar reward is the posterior mode of quality, R_i = ẑ_i = argmax_z ℓ_i(z), combined with a fixed Gaussian prior of mean zero and variance σ_z². This distinguishes rollouts that share the same rubric-point total (Equation 2 and Figure 2).
-
An online EM procedure to keep the RPN calibrated as the policy changes. One online EM sweep is run per policy step: the E-step recomputes each rollout's posterior mode, and the M-step minimizes the criterion log-likelihood at those fixed modes, with a regularizer that pulls a_j toward one when λ_a > 0 (Equation 3).
-
A theoretical analysis tying the reward to local signal-to-noise ratio. Under F = Φ, the rubric likelihood score is, up to a nonzero affine map, the unique scalar function of the verdicts that attains the local SNR bound I(z_i) = Σ_j I_j(z_i) for small changes in quality (Theorems 1 and 2). For the rubric-points reward, the bound is attained only if w_j ∝ a_j φ(u_ij)/[P_ij(1 − P_ij)] for every j, which no fixed vector of points can satisfy once criteria differ in their parameters.
-
Adaptive Fisher selection of criteria under a budget. Criteria are ranked by Σ_i I_j(ẑ_i) at the currently inferred rollout qualities, and those qualities are updated after each selected criterion is judged (Figure 1).
Main Findings
-
Criterion and normalized points scores. With Qwen3.5-4B as the base policy, RRT + online RPN is highest or tied in every column of the main policy table and exceeds Vanilla GRPO by 1.7 points and POW3R by 0.8 points on both macro metrics (RRT + online RPN macro criterion score 75.5, normalized points score 75.1; Vanilla GRPO 73.8 and 73.4). The frozen RPN is 1.0 to 1.2 macro criterion points above the empirical pass-rate conditions, and the online RPN scores 0.5 points above the frozen RPN.
-
Gains concentrate on hard criteria. RRT + online RPN exceeds Vanilla GRPO by 2.8 to 5.6 points in seven of eight difficulty bands on Medical and Science, including every Medium, Hard, and Very hard band. On rollouts from trained policies, RRT gains 2.8 to 5.6 points over GRPO on hard and very hard criteria in Medical and Science.
-
Selection reduces judge requests. Adaptive Fisher selection reaches 95.0% mean Pearson correlation with GRPO advantages from full rubric judging while leaving 21.0% of criteria unjudged in the macro mean — 11.0 points more than random selection and 1.7 points more than static Fisher. Discrimination alone matches random selection on RaR Science and RubricBench, where static Fisher leaves 13.2 and 9.6 points more criteria unjudged respectively.
-
Half-budget training stays competitive. At criterion budget 0.50, adaptive Fisher selection keeps both RRT macro scores within 0.1 points of Vanilla GRPO with full judging, and reduces RRT's judge requests by 49.0% on Medical and Science. Under random selection at the same budget, macro criterion scores fall by 2.3 points for Vanilla GRPO and 1.6 for RRT relative to each method's full-judging score.
-
Better verdict prediction. In leave-one-criterion-out prediction, RRT gains 10.1 points in pooled macro ROC-AUC over the mean of the other verdicts and exceeds the RPN without rollout evidence on every dataset. All three methods using rollout evidence reach 65.8% macro ROC-AUC within each criterion.
-
Fewer reward ties and more variation. RRT has 1.2 to 2.2 times the variance share within prompts and 2% to 58% fewer tied pairs in all 12 dataset and policy cells. With groups drawn from 48 rollouts per prompt, the tied-pair share falls by 2.2 to 2.3 points on Medical and 1.5 to 1.6 on Science across tested group sizes from 2 to 32, including the training size n = 8.
-
Response-function choice matters for ordering. With a_j = 1 and matched difficulties, the Gaussian CDF separates 83.3% of rollout pairs with equal pass counts on Medical and 39.3% on Science, while the logistic CDF leaves all such pairs tied. The adopted RPN configuration improves macro ROC-AUC by 0.7 points over the model with a_j = 1.
-
RPN text predictions correlate with empirical difficulty. Predicted and empirical criterion difficulty have Spearman correlations of 38.0% to 47.9%. Online RPN updates increase macro Pearson correlation with empirical pass rates by 1.7 points on the next policy step.
-
Stability under judge noise. Mean disagreement with the consensus verdict is 1.27% across all 12 dataset and policy cells. On repeated verdicts, RRT with marginal calibration increases its macro mean stable nonzero ordering share by 5.4 points over rubric points, with mean gains of 0.99 to 2.00 points under corrupted consensus verdicts.
-
Transfer across policies and benchmarks. RRT's macro criterion score is 0.2 points below Vanilla GRPO on Qwen3.5-2B and 0.1 points above it on Llama-3.1-8B-Instruct, and scores higher on RubricBench with both. Policies trained on Medical and Science and evaluated on HealthBench and ResearchQA gain 0.1 to 0.7 macro criterion points over Vanilla GRPO.
-
Shorter responses and lower cost. RRT has lower median response length in all 12 dataset and policy combinations, with the mean of dataset medians 10.6% to 47.9% below Vanilla GRPO across policies. Adaptive Fisher selection at budget 0.50 reduces median judging time by 49.1% to 49.6% and total step time by 23.5% to 28.7%, while generation time stays within 1.0% of full judging. RRT + online RPN adds 0.123% of a Vanilla GRPO Medical step, and the trained policy adds zero parameters or inference components at deployment.
Methodology in Plain English
The researchers start from the baseline reward: sum the points of satisfied criteria and divide by the total available points, giving a number in [0, 1]. That score ignores how a rollout passed or failed, so different verdict patterns can collapse into the same reward.
RRT borrows a model from educational testing. Each criterion is treated like a test question with two properties: a difficulty (the quality level at which a rollout is about to start passing it) and a discrimination (how sharply the pass probability rises around that level). A small neural network, the Response Parameter Network, reads the prompt and criterion text and predicts both properties, so it can handle rubrics it has never seen. Together with a Gaussian prior on quality, the full set of verdicts defines a posterior over the rollout's quality; the reward is the peak of that posterior, found by bisection.
Because the policy keeps changing during training, the network's predictions would drift out of date. RRT therefore runs one pass of online expectation maximization per policy step, using the verdicts just produced by the current rollouts to nudge the predicted parameters — an E-step that re-estimates each rollout's quality and an M-step that fits the network at those fixed qualities.
The same parameters give a second capability: Fisher information, a standard measure of how much a question is expected to reveal about an unknown quantity. RRT ranks unjudged criteria by their total Fisher information at the currently estimated qualities, judges the most informative one, updates the estimates, and repeats until the budget is spent. The paper's theoretical result shows the likelihood-based reward is the best possible scalar summary of the verdicts at small quality differences, which is the formal justification for aggregating this way rather than summing points.
Experiments use Qwen3.5-4B as the default policy, with Qwen3.5-2B and Llama-3.1-8B-Instruct for transfer, on Medical, Science, Rubrics as Rewards (RaR) Science, and RubricBench. Each policy step uses 32 prompts, eight rollouts per prompt, and one PPO epoch; GPT-5.5 with reasoning disabled supplies one binary verdict per rollout and criterion. Scores are percentages, differences are percentage points, and reported intervals are 95% bootstrap intervals over prompt groups.
Why This Matters
Rubric-based judging is one of the main ways to train language models on open-ended tasks, but every criterion judged is an expensive API call, and the additive reward it produces discards information about which criteria a rollout actually passed. RRT shows that a psychometric model can recover that information and simultaneously decide which criteria are worth judging at all.
- Education and assessment technology. Rubric criteria are already the standard for grading essays, lab reports, and open-ended exam answers; a model that infers a quality estimate from a partial set of criteria could grade consistently while asking far fewer questions of the response.
- Medical and scientific question answering. The paper evaluates on Medical and Science datasets, domains where the criteria mix required content, reasoning steps, and errors to avoid, and where RRT's largest gains fall on the criteria base policies most often miss.
- Cheap reward modeling for agent and assistant training. Any pipeline that trains a model against a checklist of requirements — customer support quality, instruction-following, safety constraints — can use criterion selection to cut judge volume while keeping the training signal close to full judging.
- Benchmark construction and evaluation. The leave-one-criterion-out analysis provides a way to estimate how much each criterion in an evaluation rubric contributes, which is useful for pruning redundant criteria from benchmark suites.
For industry, the concrete relevance is cost: at criterion budget 0.50, adaptive Fisher selection cuts judging time roughly in half and total policy-step time by roughly a quarter on Medical and Science, while the added online RPN computation is a small fraction of a step (0.123%), the trained policy adds no parameters or inference components, and one RPN warm start per dataset serves all RPN variants.
Future Directions
- Modeling local dependence between criteria. Fitted quality explains 72.8% to 85.3% of pairwise mutual information, but residual redundancy is concentrated among criteria with similar text, exceeding the bootstrap null rate by 3.4 to 11.9 points there. The paper notes that unidimensional IRT models are known to have redundancy and bias, and that explicitly modeling dependence could address this.
- Rubrics with multiple quality targets. The paper suggests that rubrics containing explicit tradeoffs between goals may need more than one latent quality dimension rather than the single shared target RRT assumes.
- Applying RRT only where its monotonicity assumption holds. RRT is developed for rubrics whose criteria are monotone indicators of one shared target; characterizing when real rubrics violate this, and what to do at that boundary, remains open.
- Extending the adaptive-versus-static selection result. Mean savings are greatest with static Fisher selection on rollouts from base policies and adaptive Fisher selection on rollouts from trained policies, which the paper reads as support for adapting judge allocation as the policy changes — a direction it does not fully resolve.
Target Audience
This paper is most useful to reinforcement learning researchers and practitioners who train language models against rubric or LLM-judge rewards and want to reduce judging cost without giving up training signal. It also speaks to applied psychometricians and educational measurement researchers interested in transferring item response theory from fixed-examinee assessment into an on-policy training loop. Readers will get the most from it with working knowledge of policy gradient methods (GRPO, PPO), Bayesian posterior inference, and the two-parameter IRT model; the theoretical sections assume comfort with Fisher information and expectation maximization.
Authors’ abstract
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this aggregation problem, judging the full rubric needs more judge requests as the criterion count grows. To address these limitations, Rubric Response Theory (RRT) measures quality and selects criteria when rubric criteria are monotone indicators of a shared target. Rather than adding assigned points, RRT uses a two parameter item response model that treats the verdict pattern as evidence about scalar quality specific to the rubric. Under this model, its likelihood score maximizes the local signal-to-noise ratio for quality. Its Response Parameter Network (RPN) reads the prompt and criterion text to predict criterion difficulty and discrimination. As the policy distribution changes during training, RRT uses online expectation maximization to update the RPN from current rollout verdicts. With Qwen3.5-4B as the policy, RRT's macro criterion score across Medical, Science, Rubrics as Rewards Science, and RubricBench is 1.7 points above that of group relative policy optimization (GRPO). On hard and very hard criteria in Medical and Science, RRT gains 2.8 to 5.6 points over GRPO. At half the criterion budget, adaptive Fisher selection with a frozen RPN keeps the macro criterion score across four datasets within 0.1 points of GRPO with full judging. These results show RRT can reduce judge requests while remaining competitive with GRPO.