Skip to content
AI.info

Research

Rubric-to-Code Credit Assignment for Reinforcement Learning

Overview Research area: Reinforcement learning (RL) for code generation, specifically the generation of interactive web applications (HTML/CSS/JavaScript) from natural language requests. The work sits

Rubric-to-Code Credit Assignment for Reinforcement Learning
arXiv
2608.27906
Published
2026-08-28
Authors
Rui Jin, Jikai Chen, Yihan Chen, Hao Zhou, Demin Zhu, Kaichen Yang, Dong Wang, Linjian Mo, Chenyi Zhuang

AI summary

Overview

Research area: Reinforcement learning (RL) for code generation, specifically the generation of interactive web applications (HTML/CSS/JavaScript) from natural language requests. The work sits at the intersection of post-training for large language models, reward design for GRPO-style RL, and benchmarks for user-facing software artifacts.

Technical level: Advanced. The paper assumes familiarity with GRPO, policy optimization, group-relative advantages, clipping, token-level weighting, and evaluator-based reward models.

Scope (one sentence): The paper introduces Rubric-to-Code Credit Assignment (RCCA), an RL framework that converts rubric-level functional feedback about a generated web application into localized, token-level optimization signals, and uses it to train Ling-RCCA-Flash from Ling-3.0-Flash.

What This Paper Is About

Interactive web application generation is judged by multiple user-facing behaviors — clicking a button opens a panel, typing in a field updates displayed content, an interaction triggers the right state change — and each of these behaviors typically depends on only a small, localized part of the generated code. Standard GRPO collapses all of these separate outcomes into one sequence-level reward and then applies the same advantage to every token in the response, so the model learns whether an application worked but not which lines of code caused it to work or fail. RCCA addresses this mismatch by keeping the rubric structure intact: it builds training tasks around explicit functional rubrics, scores responses through a staged hierarchy of failure types, and then uses evaluator-generated textual attributions to weight the GRPO objective toward the specific code spans responsible for each outcome.

Key Contributions

  1. A rubric-driven synthesis pipeline for building RL training data for interactive web application generation. Each training example is a pair (x, R), where x is a natural-language request and R = {r₁, …, r_M} is a set of independently checkable user-facing rubrics split into initial-state requirements (verified after page load) and dynamic requirements (verified after executing an interaction path). The same rubric set is used for task construction, reward computation, and code localization.

  2. A hierarchical reward that evaluates each response through four progressively more semantic stages — output-format validity, source-code validity, runtime validity, and rubric-level requirement satisfaction — implemented as a gated reward that separates fundamentally invalid artifacts from partially correct ones. Rubric-level scoring weights violations by rubric importance (core functionality, behavioral correctness, rendering quality) and by violation severity, with score ceilings for severe violations of essential requirements.

  3. Rubric-to-Code Credit Assignment (RCCA), which takes evaluator diagnostics D(y) = {d₁, …, d_K} arising from source-code validation, runtime validation, or rubric checking, expands the initial evidence locations using source-level relations (enclosing functions, event bindings, state reads/writes, function calls), and aligns the resulting source spans to generated tokens via tokenizer offset mapping. These spans become token-level diagnostic weights inside the GRPO loss without changing its group-relative advantage or clipping mechanism.

  4. Ling-RCCA-Flash, the resulting model, trained by applying SFT to Ling-3.0-Flash and then running RCCA on top of the SFT model, evaluated on MiniAppBench and ArtifactsBench.

Main Findings

  • MiniAppBench average pass rate: Ling-RCCA-Flash scores 41.25%, which the abstract describes as improving Ling-3.0-Flash by 32.20 points (Ling-3.0-Flash scores 9.05%) and slightly surpassing Claude-Opus-4.5 at 41.14%.
  • Full training trajectory: Ling-3.0-Flash 9.05% → Ling-3.0-Flash + SFT 26.85% → Ling-RCCA-Flash 41.25%, meaning RCCA adds 14.40 points over the SFT model on MiniAppBench.
  • Category strengths on MiniAppBench: hard tasks 37.60%, Humanities 54.29%, Visualization 63.46%, and Lifestyle 70.00%.
  • Comparison context on MiniAppBench: other evaluated models include GLM-5.1 at 33.50%, GPT-5.1 at 32.00%, Gemini-3-Pro-Preview at 27.52%, Claude-Sonnet-4.5 at 26.36%, Gemini-3-Flash at 17.62%, MiniMax-M2.1 at 17.12%, Grok-4.1-Fast-Reasoning at 13.77%, Mimo-V2-Flash at 12.48%, Kimi-K2-Instruct at 6.19%, Qwen3-235B-A22B at 2.88%, Qwen3-Coder-480B-A35B-Instruct at 1.83%, Qwen3-32B at 0.66%, and Hunyuan-Turbos-Latest at 2.32%. Token consumption and inference time are listed as "–" (not reported) for the three Ling models in Table 1.
  • ArtifactsBench result: Ling-RCCA-Flash ranks 1st at 76.19, improving the pre-RL (SFT) model by 4.48 points (SFT model: 71.71) and surpassing the top official leaderboard entry, GPT-5 at 72.55, by 3.64 points.
  • Cross-benchmark transfer: The gain on ArtifactsBench suggests the benefit is not benchmark-specific; the authors attribute it to learning at the level of implementation behavior rather than benchmark-specific output patterns.
  • Reward structure specifics: the gated reward returns 0 for format or source-code failures, 0.1 for runtime failures, and R_rubric ∈ [0.2, 1.0] otherwise, where R_rubric = max(0.2, 1 − Σ p(ℓ_j, q_j)) over violated rubrics.
  • Token weighting specifics: tokens overlapping directly identified spans receive diagnostic weight 3.0, tokens in contextually related spans receive 1.5, all others keep the default 1.0, and the diagnostic weight is capped at 4.0. Weighting is asymmetric by advantage sign: for negative-advantage responses the diagnostic weights apply; for non-negative advantages, diagnosed regions keep weight 1 and unaffected tokens receive weight c = 1.2.
  • Limitations noted by the authors: evaluation does not fully cover large multi-page applications, backend services, persistent storage, authentication flows, or production deployment constraints; and incorrect evaluator diagnostics can assign credit to the wrong implementation spans, especially when failures arise from interactions among distant code regions.

Methodology in Plain English

The researchers start from an existing model, Ling-3.0-Flash, described as a 124B-parameter hybrid-linear MoE with 5.1B activated parameters per token. They first run supervised fine-tuning (SFT) on it to establish basic artifact-generation ability, then run RCCA reinforcement learning on top of that SFT model.

The RL training data is not just prompts with answers. Each prompt is paired with a rubric set: a list of concrete, independently checkable requirements, split into ones about the page's initial state and ones about behavior after interaction. An evaluator then checks generated applications against these rubrics by actually running them.

Scoring is staged rather than holistic. A response that violates the output format gets 0. One that fails source-code checks gets 0. One that fails at runtime gets 0.1. Only responses that survive all of these reach rubric-level scoring, where each violated rubric subtracts a penalty based on how important that requirement is and how badly it was broken, with a floor of 0.2 and score ceilings for severe failures of essential requirements. This produces better separation between samples inside a GRPO group, which matters because GRPO's advantage is relative to other samples in the group.

The distinctive step is what happens after scoring. The evaluator also produces textual diagnostics describing what went wrong and where — pointing at code excerpts, functions, event handlers, state variables, DOM regions, or interaction paths. Because a single interactive behavior can depend on several related pieces of code, the system expands those initial pointers using source-level relationships before mapping them onto the exact generated tokens. Those tokens then get multiplied by a larger weight in the loss. The weight is applied asymmetrically: when a response did worse than its group peers, the update is concentrated on the diagnosed error regions; when a response did better, those regions are not further reinforced, and the rest of the code gets a slightly stronger positive update (weight 1.2). The group-relative advantage and the clipping mechanism of GRPO are left unchanged.

Why This Matters

Impact on research. The paper targets a general problem in RL post-training: reward signals are often sequence-level while the causes of success or failure are localized. RCCA offers one concrete recipe for bridging that gap in a domain where failures are executable and diagnosable — turning an environment harness into a source of training signal, not just a source of evaluation scores. It also connects rubric-based evaluation, which is widely used as a measurement tool, to training itself.

Real-world applications:

  • Rapid UI prototyping from natural language, where a user describes an interface and receives working HTML/CSS/JavaScript with functioning interactions.
  • Design-to-code and mockup-to-app workflows, producing interactive artifacts rather than static pages.
  • Internal tool and dashboard generation, where requirements can be expressed as checkable behaviors such as filters updating tables or panels toggling.
  • Agentic development pipelines, where an agent generates an application, executes it, diagnoses failures, and revises the code.

Industry relevance. Web application generation is a high-demand capability for coding assistants and agent platforms. A model that exceeds Claude-Opus-4.5 on MiniAppBench (41.25 vs. 41.14) and edges past GPT-5 on the ArtifactsBench leaderboard (76.19 vs. 72.55) is directly relevant to product decisions about which models to deploy for UI generation. The paper's reliance on an execution harness also reflects the industry trend of building agent environments first and then mining them for training data.

Future Directions

  • Broaden the evaluation scope. The authors explicitly note that large multi-page applications, backend services, persistent storage, authentication flows, and production deployment constraints are not fully covered.
  • Make credit assignment robust to bad diagnostics. When failures arise from interactions among distant code regions, evaluator attributions can be wrong and credit can land on the wrong spans. Improving diagnostic accuracy or making the weight assignment robust to it is an open problem.
  • Test whether the token-weighting scheme generalizes to other domains where structured, requirement-level feedback exists — the paper only demonstrates it on interactive and visual artifact generation.
  • Investigate the interaction between the reward hierarchy thresholds and the token weights. The specific values used (0.1 for runtime failure, R_rubric floor of 0.2, weights 3.0/1.5/1.0 with cap 4.0, c = 1.2) are design choices the paper does not ablate in the provided content.

Target Audience

Researchers and engineers working on reinforcement learning for large language models, particularly those interested in reward design, credit assignment, and fine-grained token-level supervision. It is also relevant to practitioners building interactive web application generation systems or agentic coding tools, to teams designing benchmarks for user-facing software artifacts, and to readers already comfortable with GRPO who want to see how structured evaluator feedback can be converted into policy optimization signals.

Authors’ abstract

Interactive web application generation requires models to produce usable HTML, CSS, and JavaScript applications from natural language requests. Unlike conventional code generation, application quality depends on multiple user-facing functional requirements, each often tied to localized code regions such as event handlers, state updates, DOM fragments, or CSS selectors. Standard GRPO collapses these structured outcomes into a single sequence-level reward and applies the resulting advantage uniformly to all tokens, weakening credit assignment. We propose \textbf{Rubric-to-Code Credit Assignment} (RCCA), a reinforcement learning framework that converts rubric-level functional feedback into localized optimization signals over generated code. RCCA builds training tasks around explicit functional rubrics, uses a hierarchical reward to separate format, source-code, runtime, and functional failures, and aligns evaluator-generated textual attributions with responsible code spans and generated tokens. The resulting model, \textbf{Ling-RCCA-Flash}, scores 41.25 on MiniAppBench, improving Ling-3.0-Flash by 32.20 points and slightly surpassing Claude Opus 4.5. It also reaches 76.19 on ArtifactsBench, improving the SFT model by 4.48 points and establishing a new top score under the official ArtifactsBench leaderboard setting by surpassing the GPT-5 score by 3.64 points, suggesting transferable implementation-level gains.

Read the original paper