Research
Beyond Timestamps: Decision-Aligned On-Policy Distillation for Long-Horizon Agents
Overview Research area: Reinforcement learning for language agents (RLVR / on-policy self-distillation), with emphasis on credit assignment in long-horizon, multi-turn agentic tasks. Technical level:

- arXiv
- 2609.33391
- Published
- 2026-09-27
- Authors
- Mingju Chen, Can Lv, Jinrong Liu, Huan Zhang, Heng Chang, Shiji Zhou
AI summary
Overview
Research area: Reinforcement learning for language agents (RLVR / on-policy self-distillation), with emphasis on credit assignment in long-horizon, multi-turn agentic tasks.
Technical level: Advanced. The paper assumes familiarity with policy-gradient methods (GRPO-style group advantages), knowledge distillation, KL regularization, and semi-Markov abstractions.
Scope: The paper diagnoses a specific failure mode in privileged-feedback training for long-horizon agents ("Decision–Timestamp Mismatch"), proposes a method called AlignOPSD to fix it, and validates it on three agentic benchmarks with two Qwen2.5 backbone scales.
What This Paper Is About
Long-horizon language agents are usually trained with reinforcement learning using verifiable rewards, but the reward typically only arrives at the end of a task, so every token in a trajectory inherits the same coarse, trajectory-level signal. On-policy self-distillation (OPSD) tries to fix this by giving the student dense feedback from a "teacher" view of the same model that has access to privileged task information during training. The authors argue that this dense feedback is often pointed at the wrong place: the teacher's guidance at a given timestamp may not correspond to the decision the student is actually making, because the same functional decision can happen at different turns in different rollouts and can also stretch across several turns. AlignOPSD is their answer — align supervision to functional decisions first, then assign credit.
Key Contributions
-
Problem identification — Decision–Timestamp Mismatch. The authors formalize why timestamp-local privileged evidence can be misaligned with the student's functional decision, both across sibling rollouts (corresponding decisions occur at different turns) and within a rollout (one decision can span multiple turns, so per-turn credit fragments or mixes decisions).
-
Decision-Aligned Supervision Rectification. A component that estimates functional decision correspondence across sibling rollouts using thinking traces as semantic surrogates of the decision state, then re-scores the same student-sampled response under functionally matched privileged contexts, mixing cross-context and local evidence according to correspondence confidence.
-
Semi-Markov Hierarchical Credit Assignment. A component that derives variable-duration decision spans from changes in the correspondence profile (measured with Jensen–Shannon divergence between adjacent turns' correspondence distributions), then allocates an outcome-grounded credit budget across spans and, within each span, across constituent turns.
-
Empirical validation. Evaluation on ALFWorld, WebShop, and Search-QA with Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct backbones against prompting, RLVR, and self-distillation baselines, plus ablations of the two alignment stages, hyperparameter sensitivity analysis, and mechanistic diagnostics of correspondence, span lengths, and gap changes.
Main Findings
-
Aggregate performance. AlignOPSD reaches 80.5% / 89.1% average success on ALFWorld, 45.1% / 49.1% accuracy on Search-QA, and 86.8% / 87.9% Score on WebShop for the 3B / 7B backbones respectively.
-
Ranking across the eight comparisons. Across the eight backbone–aggregate-metric comparisons, AlignOPSD ranks first in six and second in two, beaten only by GRPO+OPSD on 3B ALFWorld (81.2%) and Skill-GRPO* on 7B WebShop Acc (81.2%).
-
Gains over GRPO and StepOPSD. Compared with GRPO, AlignOPSD improves all eight aggregate comparisons by 5.5% / 7.9% on ALFWorld, 8.7% / 7.1% on Search-QA, 7.0% / 7.0% on WebShop Score, and 5.5% / 6.3% on WebShop Acc (3B / 7B). It also outperforms StepOPSD on all eight comparisons, indicating the gains go beyond step-level supervision alone.
-
Privileged information alone is not enough. Because the skill-conditioned baselines share the same skill source, the authors read their results as evidence that behavioral alignment and outcome-conditioned credit assignment also matter.
-
Ablation — rectification matters. On WebShop with the 7B backbone (Table 2), removing Decision-Aligned Supervision Rectification while keeping adaptive allocation drops Score from 87.9% to 84.2% and Acc from 78.9% to 71.9% — drops of 3.7 and 7.0 points.
-
Ablation — span adaptation matters. With rectification fixed, replacing adaptive spans with token-, turn-, or random allocation gives AlignOPSD improvements of 6.3 / 7.8 (Score/Acc) over token-level allocation, 1.6 / 1.6 over turn-level allocation, and 9.0 / 9.4 over random allocation. Turn-level allocation beats token-level, which beats random.
-
Learned span lengths are task-dependent. The average learned span contains 4.41 turns on ALFWorld, 3.92 on WebShop, and 2.24 on Search-QA — ALFWorld and WebShop favor longer decision units, while Search-QA concentrates mass on short spans.
-
Rectification materially changes supervision. Diagnostic plots show the rectified teacher–student gaps systematically depart from the identity gap, with correction magnitudes spread across many turns rather than driven by isolated cases.
-
Hyperparameter sensitivity is task- and scale-dependent. Retrieval top-K ∈ {2,3,4} favors K=2 on ALFWorld 3B (endpoints 78.1%, 73.4%, 76.6%), K=3 on WebShop 3B (60.2%, 69.5%, 57.8%), and K=4 on WebShop 7B, with endpoint spreads of 4.7% (ALFWorld 3B), 4.6% (WebShop 7B), and 11.7% (WebShop 3B). The threshold γ_H ∈ {0.60, 0.70, 0.80} lifts WebShop 3B from 51.6% to 69.5%, peaks ALFWorld 3B at γ_H=0.70 with 78.1%, and reaches 76.6% on WebShop 7B at γ_H=0.80, with spreads of 3.9%, 17.9%, and 9.4%. Credit temperature T ∈ {0.2, 0.5, 1.0} favors T=0.5 consistently (ALFWorld 3B: 59.4%–78.1%–62.5%; WebShop 3B: 51.6%–69.5%–49.2%; WebShop 7B peak 75.0% with only 3.1% variation).
-
Correspondence is selective and not positional. High-scoring cross-rollout matches need not occur at the same turn index; matches that preserve decision semantics are kept while superficially nearby but functionally mismatched candidates are rejected.
Methodology in Plain English
The setup: an agent interacts with an environment over multiple turns. At each turn it produces a response (possibly a thinking trace plus an action), gets an observation, and continues. At the end it gets a binary-ish verifiable reward. Standard GRPO normalizes that reward within a group of sibling rollouts for the same task and broadcasts the resulting advantage to every generated token — coarse, but at least directionally correct.
OPSD adds a second "view" of the same underlying model: the same policy conditioned on privileged task information (retrieved task-relevant knowledge available only during training). Comparing the student's log-probability of a token under the ordinary view versus the privileged view gives a dense local signal. The paper's complaint is that this comparison happens at the same timestamp, and timestamps do not line up with decisions.
AlignOPSD works in two stages.
Stage one — rectification. The method takes the thinking trace at a target turn and the privileged thinking from other sibling rollouts' turns, embeds both with a frozen semantic encoder, and computes cosine similarities to build a correspondence matrix. A selection operator keeps structurally valid sources, filters out low-similarity candidates below a threshold γ_H, keeps at most one source per sibling rollout, and retains the top-K matches. There is also an action-consistency filter: parsed actions must be valid and have an action-embedding cosine of at least 0.8, and on WebShop (where actions have exact symbolic form) operation types must match and click targets must be identical. The privileged teacher then teacher-forces the same target response under each matched source history, instead of requiring imitation of a different response. These matched log-probabilities are combined into a mixture, and the mixture is blended with the ordinary identity-view gap using a coefficient α_u that scales with match confidence (set to zero if no valid match exists, reducing to the conventional gap).
Stage two — hierarchical credit. Adjacent turns' normalized correspondence distributions are compared with Jensen–Shannon divergence; large shifts are treated as boundaries between decision spans, producing a contiguous partition of each trajectory into variable-duration segments. For each turn, an evidence score measures how strongly the rectified gaps agree with the sign of the trajectory-level advantage. A two-level allocator then hands out the conserved trajectory advantage: an outer level distributes a budget across spans based on span-aggregate evidence, and an inner level distributes within each span based on turn-level evidence. After normalization, turn weights reweight the trajectory advantage (with stop-gradient), and the resulting advantage is broadcast to that turn's tokens. KL regularization keeps both allocation levels from drifting far from the neutral (uniform) credit measure. No separate auxiliary distillation loss is added — the rectified evidence enters only through the advantage weights.
Training used Qwen2.5-3B-Instruct and Qwen2.5-7B-Instruct on a single node with up to eight NVIDIA A800 GPUs. Baselines span three families: training-free prompting (Vanilla, Skill-Prompt*), RLVR (GRPO, Skill-GRPO, Skill-GRPO*), and self-distillation RL (OPSD, GRPO+OPSD, Skill-SD, RLSD, SDAR, StepOPSD). Notably, the OPSD baseline alone collapses badly on Search-QA (0.0 average for both backbones), which the authors' framing helps explain: naive timestamp-local privileged supervision can actively mislead on some tasks.
Why This Matters
Impact on research. The paper reframes a dense-supervision problem as an alignment problem rather than a density problem. Most work in agentic distillation asks how to get more or better teacher signal; this work argues the signal must first be pointed at the right functional decision before it becomes credit. That distinction — align supervision before assigning credit — is a transferable principle that could apply to any method producing local per-token or per-turn discrepancies, including rubric-based and hindsight-based approaches. The variable-duration span construction is also a concrete alternative to the fixed credit units (single token, single turn, fixed window) that prior step- and segment-level methods assume.
Real-world applications.
- Customer-service and web-automation agents that must click through multi-step workflows (like the WebShop setting), where one "decision" such as comparing options legitimately spans several turns.
- Tool-using and retrieval-augmented assistants that alternate between searching, reading, and answering, where the credit horizon is short and the paper's Search-QA results are most relevant.
- Embodied and household robotics-style instruction following (the ALFWorld setting), where the paper reports the longest average decision spans (4.41 turns) and therefore the largest mismatch between per-turn supervision and actual decisions.
- Any deployment where a stronger, privileged-signal teacher is available at training time but must not be available at inference — the method adds no privileged inputs at deployment.
Industry relevance. Teams training agentic models with RLVR commonly rely on GRPO-style group advantages because they are simple and critic-free. This paper offers a drop-in refinement layer on top of that objective — it changes only how the trajectory advantage is reweighted, leaves other base-objective regularization unchanged, and adds no auxiliary loss term. The practical caveats are also stated plainly: the preferred retrieval breadth K varies by task and model scale, and the method requires training-time privileged information plus a frozen semantic encoder.
Future Directions
-
Choosing top-K and γ_H without per-task tuning. The paper shows no single value dominates: K=2 wins on ALFWorld 3B, K=3 on WebShop 3B, K=4 on WebShop 7B, and γ_H preferences differ across tasks. An adaptive or self-calibrating selection rule is an obvious next step.
-
Validating the thinking-trace surrogate. The method uses thinking traces as a proxy for the decision state, a choice the authors say is audited in Appendix C.2. Whether this proxy holds for models or tasks with weaker or absent reasoning traces is open.
-
Better small-model sensitivity. Credit temperature spreads reach 18.7% on ALFWorld 3B and 20.3% on WebShop 3B, markedly larger than the 3.1% spread on WebShop 7B, suggesting the smaller backbone needs more careful hyperparameter control.
-
Testing whether the principle generalizes beyond self-distillation. The claim "align supervision before assigning credit" could plausibly be combined with hindsight-value methods, process-reward models, or rubric-based evaluation, none of which the paper tests.
Target Audience
Researchers and engineers working on reinforcement learning for language agents, especially those using GRPO-style group-relative objectives, on-policy distillation, or privileged-information training. It will be most useful to readers already comfortable with policy-gradient objectives, importance ratios, and KL regularization. Practitioners building multi-turn tool-use, web-automation, or embodied agents who need finer credit assignment than a trajectory-level advantage provides will find the two-stage recipe and the reported baselines directly applicable. Readers looking for an introductory treatment of agentic RL, or for large-scale dataset statistics — the truncated content does not report dataset sizes, only that composition and splits appear in Appendix D — should look elsewhere.
Authors’ abstract
Reinforcement learning with verifiable rewards (RLVR) often relies on sparse outcome rewards, providing coarse supervision for long-horizon agents. On-policy self-distillation (OPSD) complements this signal with dense privileged feedback. However, we identify \emph{Decision--Timestamp Mismatch}: privileged guidance may be misaligned with the student's functional decision because the corresponding decision can occur at a different timestep, while the student's decision itself may span multiple timesteps rather than being tied to a single timestamp. Thus, timestamp-local supervision can misalign both the context and the temporal scope of credit. To address this mismatch, we introduce \textsc{AlignOPSD}, following the principle of aligning supervision before assigning credit. Decision-Aligned Supervision Rectification re-scores the same student-sampled response in functionally matched contexts across sibling rollouts to calibrate local teacher evidence. Semi-Markov Hierarchical Credit Assignment then derives variable-duration decision spans from correspondence changes and uses rectified evidence to allocate outcome-grounded credit across spans and their constituent turns. We evaluate \textsc{AlignOPSD} with Qwen2.5-3B and Qwen2.5-7B on ALFWorld, WebShop, and Search-QA against representative baselines. \textsc{AlignOPSD} outperforms both GRPO and StepOPSD across all eight backbone--aggregate-metric comparisons, improving on GRPO by 5.5--8.7 \% and ranking first in six. Additional analyzes examine the two alignment stages and hyperparameter sensitivity between tasks. Our code is avaliable at https://github.com/mingju-c/Align-OPSD