Research
Reasoning Stabilization Point: A Training-Time Signal for Stable Evidence and Shortcut Reliance
Overview Research area: Interpretability and training dynamics for fine-tuned NLP classifiers; specifically, tracking token-level feature attributions across fine-tuning epochs. Technical level: Inter

- arXiv
- 2601.11625
- Published
- 2026-01-12
- Authors
- Sahil Rajesh Dhayalkar
AI summary
Overview
Research area: Interpretability and training dynamics for fine-tuned NLP classifiers; specifically, tracking token-level feature attributions across fine-tuning epochs.
Technical level: Intermediate. The method relies on gradient-based attributions and rank correlation, which are standard tools, but the contribution is a conceptual framing (explanations as a time series) rather than a new explainer.
Scope: The paper introduces a training-time diagnostic called explanation drift and a summary event called the Reasoning Stabilization Point (RSP), then tests it on two lightweight transformer classifiers across two classification benchmarks plus one controlled shortcut setting.
What This Paper Is About
Fine-tuning usually gets judged by accuracy alone, but accuracy does not reveal which tokens a model is actually using to make its decisions — a model can score well while quietly leaning on spurious cues. Explanations are typically reported as a single static snapshot at the end of training, so nobody sees how the evidence evolves in between. This paper treats token attributions as a time series over training epochs, measures how much they change from epoch to epoch, and asks when the model's "reasons" stop changing.
Key Contributions
- Explanation drift: a model-agnostic scalar curve measuring epoch-to-epoch change in normalized token attributions on a fixed probe set.
- Reasoning Stabilization Point (RSP): a compact training-time signal marking the earliest epoch after which drift stays consistently low, computed from within-run drift dynamics with no tuning on out-of-distribution data.
- Empirical evidence on timing: drift stabilizes by (or before) accuracy saturation across multiple models and tasks, with post-RSP accuracy gains remaining marginal.
- Shortcut detection: under a controlled label-correlated trigger-token setting, attribution dynamics expose growing reliance on the shortcut even while validation accuracy stays competitive.
Main Findings
-
RSP lands at epoch 3.0 in every run. Across all eight reported model/task/experiment settings, RSP occurred at epoch 3.0 with a window of w=2 and the data-driven threshold τ = median(D₂, …, D_T).
-
Accuracy barely moves after RSP. In Experiment 1, DistilBERT + SST-2 reached 90.21 ± 0.48 accuracy at RSP versus a peak of 91.21 ± 0.59; DistilBERT + QNLI was 88.08 ± 0.29 at RSP versus 88.91 ± 0.22 peak; MiniLM + SST-2 was 91.40 ± 0.41 versus 92.09 ± 0.36; MiniLM + QNLI was 91.16 ± 0.23 versus 91.39 ± 0.24.
-
Drift collapses early and then stays low. Validation accuracy and drift trajectories show a consistent drop in drift early in training followed by a low-drift regime beginning at RSP, where the model's token-importance profile is stable on the probe set.
-
Shortcut reliance is visible in attributions, not in accuracy. With a label-correlated token prepended to training examples at probability p = 0.8 (validation kept clean), spur attribution mass M_t — the fraction of total attribution assigned to the injected token — peaks near RSP while validation accuracy remains high.
-
Shortcut-conditioned accuracy at RSP is lower than the clean setting. Experiment 2 accuracy at RSP was 88.03 ± 0.81 (DistilBERT + SST-2), 83.51 ± 0.25 (DistilBERT + QNLI), 89.68 ± 0.64 (MiniLM + SST-2), and 88.55 ± 0.16 (MiniLM + QNLI), with peaks of 89.87 ± 0.29, 84.56 ± 0.17, 90.63 ± 0.48, and 88.82 ± 0.29 respectively.
-
Not reported: the paper states an early-stopping corollary (stopping at or near RSP may preserve or improve robustness while matching in-domain accuracy) but explicitly flags it with a TODO in the text, and no OOD or challenge-set robustness experiments are reported.
Methodology in Plain English
The authors fine-tune two small pretrained transformers — DistilBERT (distilbert-base-uncased) and MiniLM (microsoft/MiniLM-L12-H384-uncased) — on SST-2 and QNLI for T = 5 epochs, using batch size 128, tokenizer length 128, AdamW with learning rate 2 × 10⁻⁵ and weight decay 0.01, implemented in PyTorch on an NVIDIA GeForce RTX 4060 GPU, repeated over 3 random seeds with mean ± std reported.
At every epoch they take a fixed probe set of n = 500 validation examples and compute token attributions with gradient × input at the embedding layer for the gold label. Attributions are converted to a distribution over tokens by taking absolute values and normalizing. Drift between two consecutive epochs is defined as 1 minus the Spearman rank correlation between the two normalized attribution vectors, averaged over the probe set to give one scalar per epoch. An optional label-conditional version averages within a label class to show whether stabilization happens uniformly across classes.
RSP is the first epoch where a short window of drift values stays at or below a threshold: with window w (they use w = 2 or 3), RSP = min{ t : mean of D_t … D_{t+w−1} ≤ τ }, where τ is set to the median of D₂ through D_T. That rule anchors stabilization to the run's own typical drift level, so no OOD or challenge set is needed to pick the threshold. In the shortcut experiment, they also measure spur attribution mass on a spur-probe set of 200 spurious training examples.
Why This Matters
The paper reframes interpretability as something you monitor during training rather than inspect after the fact. If a model's evidence stabilizes well before its accuracy saturates, then the extra epochs are buying marginal accuracy while potentially reallocating importance toward brittle cues — a signal that pure accuracy curves cannot show. Because the method is post-hoc over saved checkpoints and uses off-the-shelf gradient attributions, it adds minimal overhead beyond standard fine-tuning.
Real-world applications:
- Checkpoint selection: choosing a checkpoint in a stable-evidence regime rather than the final epoch.
- Spurious-correlation auditing: flagging when a classifier is leaning on dataset artifacts or trigger-like tokens in production training pipelines.
- Model monitoring during continued fine-tuning or domain adaptation: detecting when decision evidence is shifting even if metrics look flat.
- Debugging NLP classifiers in high-stakes domains: providing a cheap diagnostic alongside accuracy for text classifiers where "why" matters as much as "how accurate."
Industry relevance: teams that fine-tune transformer classifiers on limited compute can run this diagnostic over existing checkpoints without new explainer infrastructure, making it a low-cost addition to standard training monitoring.
Future Directions
- Scale and task type: the paper notes that behavior of drift and RSP in larger models, longer training regimes, and generative tasks remains an open question.
- Multiple stabilization phases: models trained for many epochs or with curriculum-style schedules may exhibit several stabilization phases that a single RSP would not capture.
- Explainer dependence: results characterize stabilization with respect to a given explainer and similarity metric; different attribution methods may yield quantitatively different drift curves, and no explainer-independent guarantee is offered.
- Real-world shortcuts: real spurious correlations may be subtler and harder to isolate than injected trigger tokens, so drift should be treated as a complementary signal rather than a standalone robustness criterion.
Target Audience
Researchers and practitioners working on interpretability, fine-tuning, and robustness of NLP classifiers; machine learning engineers who select checkpoints or monitor fine-tuning runs; and anyone studying shortcut learning and spurious correlations, who would benefit from a cheap training-time signal that surfaces evidence shifts hidden by aggregate accuracy.
Authors’ abstract
Fine-tuning pretrained language models can improve task performance while subtly altering the evidence a model relies on. We propose a training-time interpretability view that tracks token-level attributions across finetuning epochs. We define explanation driftas the epoch-to-epoch change in normalized token attributions on a fixed probe set, and introduce the Reasoning Stabilization Point(RSP), the earliest epoch after which drift remains consistently low. RSP is computed from within-run drift dynamics and requires no tuning on out-of-distribution data. Across multiple lightweight transformer classifiers and benchmark classification tasks, drift typically collapses into a low, stable regime early in training, while validation accuracy continues to change only marginally. In a controlled shortcut setting with label-correlated trigger tokens, attribution dynamics expose increasing reliance on the shortcut even when validation accuracy remains competitive. Overall, explanation drift provides a simple, low-cost diagnostic for monitoring how decision evidence evolves during fine-tuning and for selecting checkpoints in a stable-evidence regime.