Research
Learning to Steer, Steering to See: Unveiling the Geometry of RLVR in Large Language Models via Trainable Vectors
Overview Research area: Reinforcement learning post-training of large language models (RLVR), interpretability of activation-space geometry, and training stability. Technical level: Advanced — assumes

- arXiv
- 2609.34344
- Published
- 2026-09-28
- Authors
- Yuchen Cai, Ding Cao, Qixiang Yin, Xin Xu, Kai Yang, Siye Wu, Pengyuan Wang, Jiaxuan Wang, Weijie Liu, Saiyong Yang, Guangzhong Sun, Guiquan Liu, Junfeng Fang
AI summary
Overview
Research area: Reinforcement learning post-training of large language models (RLVR), interpretability of activation-space geometry, and training stability.
Technical level: Advanced — assumes familiarity with transformer hidden states, LoRA, policy-gradient RL for LLMs (GRPO/DAPO), knowledge distillation, principal component analysis, and gradient projection.
Scope: The paper uses trainable "steering vectors" injected into frozen base models as a probe to characterize a low-dimensional effective manifold underlying RLVR-induced gains, then builds a training-stabilization wrapper (Alpha-Stabler) from the resulting geometric findings.
What This Paper Is About
RL post-training of LLMs produces large reasoning gains, but because parameter updates live in very high dimensions, it is hard to tell which internal changes actually matter. The authors sidestep parameter space and instead freeze the base model and learn input-invariant shared vectors at selected layers, asking whether RL-induced gains can be recovered by low-dimensional activation interventions. From this probe they derive two geometric properties of the RL effect — Effective Manifold Capacity and Control Manifold Separation — and turn the second into a practical training stabilizer.
Key Contributions
-
A steering-vector methodology for studying RLVR. The paper reframes RL post-training as an activation-space intervention problem: a frozen base model is distilled toward its RL-trained counterpart using only input-invariant, layer-specific vectors shared across all inputs and token positions (plus a gated low-rank variant), isolating whether RL gains admit a low-dimensional description.
-
Property 1: Effective Manifold Capacity. A single vector per controlled layer recovers over 85% of the full-fine-tuning gain on most evaluated tasks, but this compressibility has limits governed by the teacher–student distribution gap and by injection depth, and is relieved by token-dependent modulation.
-
Property 2: Control Manifold Separation. Effective control directions concentrate in the low-variance complement of the activation principal subspace, are largely reproducible across teacher parameterizations and distillation regimes within a task, and their cross-topic alignment correlates with capability transfer (R² = 0.91 across six ordered source–target pairs).
-
Alpha-Stabler. A plug-and-play framework with a Predictor that monitors principal-subspace intrusion (PSI) as an early warning of collapse and a Controller that removes the principal-subspace component of activation gradients during backpropagation. It adds zero trainable parameters, requires no modification to the optimizer, reward function, or architecture, and stabilizes training for 2,000 steps.
Main Findings
-
Extreme but bounded compressibility. Under on-policy distillation, LoRA and full-parameter training perform comparably, while a single steering vector per controlled layer still recovers over 85% of the RL gain. Off-policy distillation with teacher-generated trajectories shows a comparably high recovery rate on evaluation datasets.
-
Teacher identity matters. Vector steering effectively transfers capabilities from RL-trained teachers, but degrades substantially with the tested non-RL teachers, with some settings collapsing as the teacher–student gap widens. The authors attribute this to corrections becoming more input-dependent as the gap grows.
-
Expressive capacity ordering under large gaps. Under off-policy distillation with a small gap, full-parameter training, LoRA, and vector steering perform comparably; with a larger gap a clear ordering emerges: full-parameter training best, then LoRA, then vector steering worst. This matches their expressive capacities (model-wide adjustments, low-rank input-dependent corrections, fixed offsets).
-
Depth-dependent degradation of fixed vectors. Selected-layer full-parameter updating and LoRA remain effective across depths, whereas fixed-vector steering works well at lower and intermediate layers but deteriorates near the output. High-layer steering plateaus early and extended training does not noticeably close the gap.
-
Gradient magnitude is not the explanation. Shallow and high-layer injections receive gradients of comparable magnitude yet yield substantially different final performance, so weak gradient signals do not explain high-layer degradation.
-
Token-dependent modulation fixes high-layer degradation. Replacing the fixed offset with a Rank-1 gated correction (h' = h + B σ(A h), with B zero-initialized) reduces teacher KL divergence and improves converged performance at all four tested high-layer injection sites. A parameter-matched static (input-independent) control does not reproduce the improvement, supporting the role of token-dependent coefficients; increasing rank to r = 16 yields further gains toward the full-parameter reference.
-
Gradient alignment increases with depth but does not predict success. Mean pairwise cosine similarity between per-token gradient contributions to the steering vector rises from about 0.019 at lower and intermediate layers to about 0.14 near the output — a trend that would favor shared-direction interventions at higher layers, where fixed-vector steering in fact degrades.
-
Gating activations become more selective at high layers. At lower and intermediate layers the Rank-1 gating coefficients spread broadly over [0, 1]; at higher layers they concentrate near 0 and 1.
-
The leading direction is effective but not unique. On SciKnowEval with DeepSeek-R1-Distill-Qwen-1.5B, the first direction achieves 67.5% standalone accuracy and the eighth direction v(7) achieves 62.0% — only 5.5 percentage points lower — despite being orthogonal to the preceding seven directions at every controlled layer. Standalone accuracy generally declines with extraction order, and the decline is sharper in smaller models, whereas larger models retain higher-order directions closer to the leading direction.
-
Directions are reproducible within a task. Same-index alignment exceeds cross-index alignment across the tested teacher parameterizations (full-parameter fine-tuning vs. LoRA with different ranks), and on- vs. off-policy directions show matching-index alignment.
-
Alignment tracks transfer across topics. Partitioning Science into physics, chemistry, materials, and biology and evaluating six ordered source–target pairs (excluding self-pairs), stronger directional alignment is associated with larger normalized transfer gains, with a linear fit of R² = 0.91.
-
Effective directions avoid the principal subspace. Measured against the leading orthonormal eigenvectors of the frozen base model's token-level activation covariance at each layer, with r corresponding to 10% of the hidden dimension, unconstrained steering places only about 1%–2% of its energy in the Top-10% principal subspace — well below the 10% isotropic reference — across four scientific topics and the displayed directions v(0) through v(8).
-
Forcing alignment with the principal subspace does not help. Restricting steering vectors to the span of the Top-30% activation-PC components after every optimizer step raises the Top-10% energy fraction to about 15%, but accuracy does not improve meaningfully.
-
Principal-subspace intrusion precedes collapse. In collapse-prone RL runs, PSI stays low (well below the 10% isotropic reference) during stable training, then departs from that range and becomes volatile before any visible score degradation; as collapse progresses, PSI keeps rising while the score falls more sharply.
-
Alpha-Stabler stabilizes without costing performance. Uncontrolled runs show a pronounced PSI increase coinciding with sharp score deterioration; under control, PSI remains near its early-training range while task scores continue to improve.
-
Ablations favor triggered, principal-subspace-targeted control. Controlled training keeps peak PSI at a low single-digit percentage across the tested learning rates while uncontrolled training eventually collapses at each. Triggered control matches always-on control in both score and peak PSI, whereas delayed intervention progressively degrades the score and raises peak PSI. Removal of Top-PC components achieves the highest score among compared gradient-control methods; random-subspace, Middle-PC, and Bottom-PC control fail to match it and remain close to the uncontrolled baseline, and Top-PC also compares favorably with global gradient clipping. Score trajectories with and without Alpha-Stabler are nearly identical under the same measured wall-clock time, indicating almost no additional time overhead.
-
Mechanism summary. Alpha-Stabler warms up for T_warm = 50 updates, then fixes principal bases with r = ceil(0.10d), sets warning and release thresholds from layer-wise medians and scaled median absolute deviations, smooths PSI with an exponential moving average, activates control after p consecutive valid checks above the warning threshold, and releases it below the lower threshold. The Controller applies G̃ = G − a(GUUᵀ), requiring two low-rank matrix multiplications per active hook.
Methodology in Plain English
The authors treat RL post-training as an intervention on activations rather than as a change to weights. They take an RL-trained model as a teacher, freeze a base model as the student, and learn one shared vector per selected layer that is added to that layer's output. Because the vector is the same for every input and every token position, the number of trainable parameters is tiny — yet if the RL gain is recoverable this way, the gain must live in a low-dimensional structure. They compare this against full-parameter training and LoRA, under both on-policy and off-policy distillation, and vary teacher type to change the distribution gap.
To probe depth, they slide the injection layer across the network and compare against placing LoRA or full-parameter updates at the same layer. To separate "weak gradients" from "weak expressiveness," they inspect per-token gradient norms and the mean pairwise cosine similarity of per-token gradient contributions, then replace the fixed offset with a rank-1 nonlinear gate that keeps a single shared direction per layer but lets its strength vary per token, contrasting it against a parameter-matched constant-coefficient control.
For geometry, they learn multiple orthogonal vectors per layer in sequence: each new vector is initialized in the remaining orthogonal complement, optimized to convergence, and reprojected after every optimizer step, while earlier vectors are frozen and used only to define orthogonality constraints. Directions learned under different teacher parameterizations and distillation regimes are compared by layer-averaged cosine similarity, and topic-to-topic alignment is compared against measured transfer gains.
Finally, they characterize where the directions sit relative to the base model's activation principal subspace — the leading eigenvectors of the token-level activation covariance covering 10% of the hidden dimension — and define principal-subspace intrusion (PSI) as the fraction of activation-shift energy that falls in that subspace. Two operational tools follow: the Predictor monitors PSI against fixed thresholds as a collapse early warning, and the Controller projects the principal-subspace component out of backward gradients while preserving the orthogonal complement, leaving forward activations, rewards, likelihood ratios, and the GRPO/DAPO loss untouched.
Why This Matters
Impact on research. The paper argues that RLVR's apparently complex training dynamics are governed by a surprisingly simple structure — a low-dimensional manifold confined to the low-variance complement of the activation principal subspace — and recasts RLVR from black-box parameter optimization into an interpretable geometric process. It offers a diagnostic (PSI) and a control mechanism that connect interpretability findings to training-time practice, and reports results across 5 LLMs and 6 tasks with verifiable rewards.
Real-world applications:
- More reliable post-training pipelines for reasoning models, where training collapse wastes large compute budgets.
- Early-warning monitoring for long RL runs, since PSI rises and becomes volatile before any visible score degradation.
- Cheaper capability transfer: distillation into a frozen base model using a small number of shared per-layer vectors instead of full fine-tuning.
- Selecting where to inject lightweight adaptation — the results indicate fixed offsets suit lower and intermediate layers, while high layers need token-dependent modulation.
Industry relevance. Alpha-Stabler is described as plug-and-play, introduces no additional trainable parameters, requires no change to the optimizer, reward function, or model architecture, needs only two low-rank matrix multiplications per active hook, and adds almost no measured wall-clock overhead — properties that matter for large-scale RL pipelines where any per-step cost is multiplied across long runs. The paper's author list spans USTC, BUPT, and Tencent Hunyuan, and the code is released.
Future Directions
- Ground the theory. The theoretical framework rests on empirically motivated assumptions that are not yet derived from first principles; the authors suggest relaxing them via neuron attribution and causal tracing.
- Extend beyond RLVR. The analysis is restricted to RLVR on tasks such as mathematical reasoning and code generation; whether the same geometry holds for RLHF, multi-turn agentic decision-making, and open-ended generation remains open.
- Turn off-principal localization into a design principle. The authors propose adding complement-space regularizers or using PSI as an adaptive signal for learning rate and intervention strength.
- Remove the monitoring overhead. A further goal is low-rank control without layer monitoring.
Target Audience
Researchers and engineers working on LLM post-training and RL fine-tuning, mechanistic interpretability of activation space, and parameter-efficient adaptation; practitioners who run long GRPO/DAPO training jobs and need collapse early-warning or stabilization tooling; and readers interested in connecting geometric analysis of hidden states to practical training interventions. The paper presumes comfort with transformer internals, PCA-style subspace language, and RL fine-tuning objectives, so it is best suited to readers at an advanced level in this area.
Authors’ abstract
Reinforcement learning (RL) has become a key paradigm for enhancing the reasoning of large language models, yet the high dimensionality of parameter updates makes its training dynamics hard to analyze. We study reinforcement learning with verifiable rewards (RLVR) and use vector steering to identify a low-dimensional effective manifold in activation space associated with RL-induced gains. We uncover two geometric properties. (1) Effective Manifold Capacity: the capacity needed to reproduce RL gains can be very small but is not infinitely compressible; at extremely low capacity, intervention dimensionality and input-dependent expressiveness become key constraints, and this requirement varies with injection depth. (2) Control Manifold Separation: effective control directions lie mainly in the low-variance complement of the activation principal subspace. Within a task and base model, the learned geometry stays largely consistent across training configurations, and across tasks geometric alignment correlates with capability transfer. Experiments on 5 LLMs and 6 verifiable-reward tasks support these findings. We then propose Alpha-Stabler, a plug-and-play framework with a Predictor that monitors principal-subspace intrusion for early collapse warnings, and a Controller that removes the principal-subspace component of activation gradients during backpropagation while preserving the orthogonal complement. Alpha-Stabler stabilizes training for 2,000 steps and consistently improves RL gains, offering practical insights for robust post-training. Code: https://github.com/caiyuchen-ustc/On_Policy_Vector_Training