Skip to content
AI.info

Research

FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff

Overview Research area: Continual learning / deep learning optimization — specifically the stability–plasticity tradeoff in models trained on nonstationary data. Technical level: Intermediate to Advan

FIRE: Frobenius-Isometry Reinitialization for Balancing the Stability-Plasticity Tradeoff
arXiv
2602.08040
Published
2026-02-08
Authors
Isaac Han, Sangyeon Park, Seungwon Oh, Donghu Kim, Hojoon Lee, Kyung-Joong Kim

AI summary

Overview

Research area: Continual learning / deep learning optimization — specifically the stability–plasticity tradeoff in models trained on nonstationary data.

Technical level: Intermediate to Advanced. The paper combines a constrained-optimization formulation (Orthogonal Procrustes) with four supporting theorems and empirical validation across vision, language, and reinforcement learning. Some familiarity with weight matrices, Frobenius/spectral norms, and continual learning baselines will help.

Scope (one sentence): The paper proposes FIRE, a reinitialization method that projects neural network weights onto an isotropy manifold to simultaneously preserve prior knowledge (stability) and restore adaptability (plasticity), evaluated on continual visual learning, LLM continual pretraining, and reinforcement learning.

What This Paper Is About

Deep networks trained on data that changes over time tend to lose plasticity — the ability to keep learning new things — while also risking catastrophic forgetting of what they already know. Existing fixes either reinitialize weights too conservatively (plasticity is not restored) or too aggressively (useful knowledge is destroyed), and choosing between these extremes requires careful tuning. FIRE reframes reinitialization as a constrained optimization problem: find weights close to the old ones (high stability) that are also close to orthonormal (high plasticity), and solve it efficiently with a Newton–Schulz iteration.

Key Contributions

  1. Two differentiable proxies for the tradeoff. The paper introduces Squared Frobenius Error (SFE) as a stability measure and Deviation from Isometry (DfI) — defined as ‖WᵀW − I‖²_F — as an optimizable plasticity measure, arguing that the previous plasticity indicators (loss curvature, dormant neurons, feature rank) are data-dependent and non-differentiable.

  2. A constrained-optimization framing of reinitialization. FIRE minimizes SFE subject to DfI = 0, which the authors show is mathematically equivalent to the classical Orthogonal Procrustes Problem with the closed-form solution W̃⋆ = W(WᵀW)^(−1/2) via polar decomposition.

  3. An efficient Newton–Schulz approximation. The closed form is approximated with the iteration X_{k+1} = aX_k + bX_k(X_kᵀX_k) with a = 1.5, b = −0.5, and X₀ = W/‖W‖_F, adding less than 1% to training time and applying kernel-wise in convolutional layers.

  4. Theory linking DfI to known plasticity symptoms. Four theorems connect the chosen quantities to prior diagnostics: Theorem 1 bounds output-feature-covariance difference by SFE; Theorem 2 bounds Hessian spectral norm by layerwise DfI; Theorem 3 gives a lower bound on effective rank controlled by DfI; Theorem 4 bounds per-neuron activity scores by DfI, addressing dormant neurons.

Main Findings

  • Consistent superiority across three domains. The abstract and experiments report that FIRE outperforms both naive training without intervention and standard reinitialization methods on continual visual learning, language modeling, and reinforcement learning. Exact accuracy numbers are presented in figures rather than in the text provided.

  • Vision results vary by architecture. On CIFAR-10 with ResNet-18 and Tiny ImageNet with VGG-16, FIRE outperforms all baselines. On CIFAR-100 with ViT-Tiny, the improvement is less pronounced: FIRE still beats all baselines except DASH and is competitive with S&P, which the authors attribute to DASH's data-dependent shrinking strategy being helpful for transformer reinitialization.

  • Reinitialization methods degrade sharply after each reset. In the continual visual setting, full reset and DASH show sharp drops immediately after each reset, S&P avoids such drops but remains suboptimal; FIRE incurs only a slight or negligible drop.

  • Node-resetting and gradient-orthogonalization baselines underperform. CBP, SNR, and ReDo show poor overall performance, consistent with prior findings that trainability-focused methods yield limited generalization gains. The Muon optimizer also performs poorly, which the authors read as evidence that periodically orthogonalizing weights is more effective than applying the same iteration to gradients.

  • Continuous constraints slow convergence. Parseval regularization and L2init converge more slowly than FIRE, particularly on CIFAR-10/ResNet-18 and Tiny ImageNet/VGG-16, suggesting that enforcing the orthogonality constraint throughout training hinders convergence.

  • Full reset is a poor strategy for LLM continual pretraining. In the GPT-0.1B setting (pretrained on WikiText-103, then trained on OpenWebText + WikiText-103), full reset cannot outperform the base model, because the instability introduced by erasing all prior information outweighs the plasticity benefit. The gap between the base model and full reset narrows as pretraining progresses, consistent with prior observations that plasticity loss worsens with longer pretraining.

  • FIRE works untuned in the LLM setting. FIRE was applied with no tuning and a fixed 5 iterations, while S&P was carefully tuned over varying reinitialization degrees — yet FIRE maintained strong performance even when initialized from the 60k pretraining checkpoint.

  • Reinforcement learning gains. With DQN on Asterix, BeamRider, and DemonAttack, FIRE consistently outperforms S&P, surpasses full reset in Asterix, and remains competitive elsewhere. With SAC/SimBa on HumanoidBench (balance, walk, run), FIRE achieves superior or competitive results. Plasticity Injection performs poorly across discrete and continuous control tasks.

  • Ablation confirms the mechanism. FIRE achieves the lowest DfI while maintaining the lowest SFE, and produces a smoother loss landscape than S&P while preserving lower SFE. DASH smooths the loss landscape effectively but has the highest SFE, which the authors link to erasure of learned knowledge and post-reset instability.

  • Robust to its single hyperparameter. The number of Newton–Schulz iterations is the only hyperparameter; FIRE already provides strong performance gains with as few as five iterations.

Methodology in Plain English

The authors start from the observation that after training on one dataset, a network's weight matrices drift away from the well-conditioned geometry that made them easy to train in the first place. Rather than picking an arbitrary reset point, they ask: what is the closest set of weights to the current ones that is also perfectly orthonormal? "Closest" is measured by SFE (sum of squared entrywise differences), and "perfectly orthonormal" is measured by DfI being zero. This is a textbook problem — the Orthogonal Procrustes Problem — and it has an exact answer: multiply the weight matrix by the inverse square root of its Gram matrix. Because computing that matrix inverse square root is expensive for large networks, they approximate it with the Newton–Schulz iteration, a cheap fixed-point method that pushes singular values toward 1. In practice, FIRE is applied after training on the current dataset but before learning on new data, and in the reinforcement learning experiments it is applied once at the midpoint of training. For convolutional layers, the orthogonalization is applied per filter along the spatial dimensions.

Why This Matters

Research impact. The paper offers a unified recipe that is intended to work across vision, language, and reinforcement learning, rather than being tailored to one regime. It also provides a theoretical bridge: it shows that a single differentiable quantity (DfI) is linked to three previously separate diagnostics of plasticity loss — loss curvature, feature rank, and dormant neurons — which may make future plasticity research easier to formalize and optimize.

Real-world applications.

  • Autonomous driving systems that must recognize new traffic signs, road layouts, or weather conditions absent from training data.
  • Large language models that need continual updates to overcome a fixed knowledge cutoff date.
  • Robots operating in dynamic physical environments that must adjust perception and control policies as conditions change.
  • General foundation-model or agent training pipelines on expanding datasets where past data remain accessible.

Industry relevance. The method adds less than 1% to training time and is applied as a discrete intervention rather than a continuous constraint, which the paper argues avoids the slower convergence seen with regularizers like Parseval regularization and L2init. The minimal hyperparameter surface (one: iteration count, robust down to five) reduces tuning cost — a practical consideration for large-scale training pipelines.

Future Directions

  • Remove the past-data assumption. The authors state their main limitation is assuming access to past data during continual learning, and call for evaluating FIRE under restricted data access.
  • Scale to larger language models. LLM experiments used only GPT-0.1B, which the authors acknowledge as small; they propose larger models as a promising direction.
  • Extend beyond pretraining. The paper suggests applying FIRE to continual fine-tuning of LLMs, not only continual pretraining.
  • Understand data-dependent resets. The authors note that DASH's data-dependent shrinking was especially effective on ViT-Tiny, hinting that combining data guidance with the principled projection could be worth exploring for transformer architectures.

Target Audience

Researchers and practitioners working on continual learning, plasticity loss, and nonstationary-data training — particularly those who apply reinitialization or regularization to mitigate forgetting. It is also relevant to reinforcement learning engineers dealing with high replay ratio settings, and to LLM pretraining teams considering continual updates. Readers without a background in matrix decompositions or spectral norms will find the theory section demanding, but the algorithmic core (Algorithm 1) is only a few lines of PyTorch-like pseudocode.

Authors’ abstract

Deep neural networks trained on nonstationary data must balance stability (i.e., retaining prior knowledge) and plasticity (i.e., adapting to new tasks). Standard reinitialization methods, which reinitialize weights toward their original values, are widely used but difficult to tune: conservative reinitializations fail to restore plasticity, while aggressive ones erase useful knowledge. We propose FIRE, a principled reinitialization method that explicitly balances the stability-plasticity tradeoff. FIRE quantifies stability through Squared Frobenius Error (SFE), measuring proximity to past weights, and plasticity through Deviation from Isometry (DfI), reflecting weight isotropy. The reinitialization point is obtained by solving a constrained optimization problem, minimizing SFE subject to DfI being zero, which is efficiently approximated by Newton-Schulz iteration. FIRE is evaluated on continual visual learning (CIFAR-10 with ResNet-18), language modeling (OpenWebText with GPT-0.1B), and reinforcement learning (HumanoidBench with SAC and Atari games with DQN). Across all domains, FIRE consistently outperforms both naive training without intervention and standard reinitialization methods, demonstrating effective balancing of the stability-plasticity tradeoff.

Read the original paper