Skip to content
AI.info

Research

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Overview Research area: Natural Language Processing, specifically reward modeling and reinforcement learning for legal-domain la

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models
arXiv
2609.39071
Published
2026-09-30
Authors
Yida Cai, Xin Dai, Bingxiang He, Huiyuan Xie, Yuxiao Ye, Zhenghao Liu, Yang Bai, Zhiyuan Liu

AI summary

LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models

Overview

Research area: Natural Language Processing, specifically reward modeling and reinforcement learning for legal-domain large language models.

Technical level: Intermediate. The framework itself is described conceptually, but the paper assumes familiarity with reward models, Direct Preference Optimization (DPO), Group Relative Policy Optimization (GRPO), and Bradley–Terry preference training.

Scope (one sentence): The paper proposes a three-dimension taxonomy of legal response quality (Style, Element, Chain), turns it into rubric-based reward functions, and shows that the resulting rewards and trained reward models improve legal text generation through DPO, test-time scaling, and GRPO.

What This Paper Is About

Most reward signals used to train legal language models judge a response holistically or check only the final answer, such as a predicted charge or statute. That makes it hard to tell whether a model improved because its legal reasoning got better or because its output simply got longer, fluent, or stylistically convenient. LexReward addresses this by first defining what a high-quality legal response actually contains, then building interpretable rewards for each part of that definition. The authors work in Chinese legal text, using datasets such as CLASE, JEC-QA, and LexChain.

Key Contributions

  1. A structured taxonomy of legal response quality. The authors decompose quality into three complementary dimensions — Style (lexical and syntactic quality), Element (legal subjects, facts, statutes, and decisions), and Chain (order, completeness, correctness, and non-redundancy of reasoning) — with finer-grained criteria under each. The taxonomy was built by combining top-down expert specification with bottom-up error analysis of model outputs compared against reference answers.

  2. Interpretable rubric-based rewards for every criterion. Context-independent criteria (Style) use rule-based scorers calibrated against authentic Chinese judicial documents; context-dependent criteria (Element, and Chain when reasoning steps must be inferred) use an LLM-as-a-judge. The paper reports that these rewards reliably distinguish responses of different quality along their respective dimensions.

  3. Rubric-derived preference data, DPO training, and LexRM. Rubric scores are used to build pairwise preference data, which both trains a family of dimension-specific reward models (LexRM — described as the first collection of reward models developed for the Chinese legal context) and drives DPO. A multi-dimensional model, LexRM-Merge, is formed by combining the three dimension-specific task vectors.

  4. Reinforcement learning with learned reward models. Each dimension-specific LexRM is integrated into a GRPO pipeline and improves policy performance in its corresponding dimension without requiring reference answers at reward time.

Main Findings

  • Rubrics beat random and multi-model baselines. On the CLASE style dataset, rubric-based scoring reached 79.00 accuracy versus 50.00 for random selection. On LegalΔ, the rubric reached an average of 60.27 versus 49.63 for the five-model average. On LexChain, the rubric reached an overall 47.80 versus 45.41 for the five-model average.

  • LexRM is the strongest selector on Style and Chain. LexRM-Style reached 81.75 on CLASE (the best result in that column, above the rubric at 79.00 and Skywork-Qwen at 78.50), and LexRM-Chain reached 50.46 overall on LexChain (best, above Skywork-Qwen at 50.19). LexRM-Element ranked second on LegalΔ with an average of 66.32, behind Skywork-Qwen at 68.04.

  • The merged model is mixed. LexRM-Merge reached 65.77 on LegalΔ, matching the Element expert

Authors’ abstract

Legal language models require reward signals that capture not only answer correctness but also the multidimensional quality of legal responses. Existing reward methods, however, often rely on coarse-grained holistic judgments, providing limited domain specificity and interpretability. We introduce LexReward, a taxonomy-driven framework for legal reward modeling. LexReward characterizes legal response quality along three complementary dimensions: Style, covering lexical and syntactic quality; Element, assessing legal subjects, facts, statutes, and decisions; and Chain, evaluating the order, completeness, correctness, and non-redundancy of legal reasoning. For each dimension, we develop rubrics that specify evaluation criteria and quality levels. The resulting rewards are used to construct pairwise preference data for Direct Preference Optimization (DPO) and reward-model training. Experiments show that the rubric-based rewards reliably distinguish legal responses of different quality and that DPO training on the preference data improves performance across all three dimensions. The learned reward models, LexRM, also support effective downstream optimization: each dimension-specific reward model improves policy performance in its corresponding dimension through reinforcement learning, without requiring reference answers at reward time. Dimension-wise analyses further support the effectiveness of the proposed taxonomy and reward construction.

Read the original paper