Research
Local Support Learning
Overview Research area: Continual learning and catastrophic forgetting in large pre-trained language models, specifically during parameter-efficient post-training/finetuning. Technical level: Intermed

- arXiv
- 2610.02126
- Published
- 2026-10-01
- Authors
- Assaf Ben-Kish, Akarsh Kumar, James Glass, Raja Giryes
AI summary
Overview
Research area: Continual learning and catastrophic forgetting in large pre-trained language models, specifically during parameter-efficient post-training/finetuning.
Technical level: Intermediate. The paper is readable with a basic grasp of gradient descent, weight matrices, and LoRA-style finetuning, but it assumes familiarity with continual learning terminology.
Scope: The paper diagnoses why gradient-based finetuning overwrites prior capabilities and proposes Local Support Learning (LSL), a method that pairs a standard weight adapter with a Gaussian-Mixture-Model gate so that updates only affect the region of input space that produced them.
What This Paper Is About
Modern LLMs are pretrained once and then repeatedly finetuned, but each new finetuning phase tends to erase capabilities learned earlier — a phenomenon called catastrophic forgetting. Existing fixes either require storing old data (impractical at pretraining scale), restrict updates to low-rank subspaces (which still forgets and caps learning capacity), or regularize the model's outputs. This paper argues that all of these apply their updates to all inputs, and instead proposes restricting each weight update to the input region that actually produced it, using a gate trained only on the current phase's data.
Key Contributions
-
A geometric analysis of forgetting. Viewing the problem at the level of a single weight matrix, the authors show that any prior input not orthogonal to a gradient update has its mapping altered. They argue this makes updates from gradient-based optimizers — including Adam (which rescales gradients by their second moment) and Muon (which orthogonalizes them) — suboptimal for retention, since both still produce a matrix update that acts on all inputs.
-
Local Support Learning (LSL). A general-purpose framework that augments ordinary gradient-based training with two components: a standard weight adapter trained as usual, and a gating function that enables the adapter only on input activations from its own training distribution. No access to prior data is required.
-
A theoretical link between optimal gates and density estimators. The authors decompose a gate's worst-case error into deficit (optimized directly by the training objective) and excess (the region outside the training distribution where the gate wrongly stays open — this is the forgetting term). They show that, unlike conventional MLP classifiers, density estimators can recover optimal gates under explicit conditions.
-
Evaluation at scale. Experiments on LLMs up to 7 billion parameters, across multiple sequential finetuning phases, demonstrating retention, hyperparameter robustness, scaling behavior, and efficiency in memory and compute.
Main Findings
-
LSL retains pretrained capability across three diverse post-training tasks. Finetuning Qwen2.5-7B-Instruct on cybersecurity instruction tuning, English→Igbo translation, and chemistry instruction tuning, LSL learns each new task while staying close to optimal retention on GSM8K (math), HumanEval (code), and IFEval (instruction following). LoRA adapted strongly but forgot in all settings; OP-LoRA performed very similarly to LoRA (the authors attribute this to forgetting being only weakly tied to the top-k singular-value subspace); LwF beat the weight-based baselines but still trailed LSL on all tasks.
-
Retention holds across multiple sequential phases. With the order (1) Igbo translation, (2) Chemistry, (3) Cybersecurity, LSL retained pretrained capabilities throughout, while the LoRA baseline (a fresh adapter per phase, merged before the next phase) dropped by 44% after just one phase. Within the finetuning benchmarks, the LoRA baseline lost 8% on translation after two phases and 5% on chemistry after one phase; zero-shot performance on cybersecurity fell 11% after the translation phase. LSL avoided all of these drops.
-
Learning hyperparameters decouple from forgetting. Sweeps over learning rate, adapter rank, and batch size show LSL reaching full learning capacity without sacrificing pretrained capability, whereas LoRA faces a direct learning-versus-forgetting trade-off.
-
Retention improves with model scale. Tested from 1.5B to 7B parameters, LSL worked at all sizes and improved with scale. LoRA also improved with scale but far more slowly, remaining well below LSL at every size.
-
The GMM gate is the dominant mechanism, not the smoothing. The GMM improves retention from a raw score of 55.3 to 69.8 (76.6% → 96.6% of the base model's performance); temporal smoothing adds a smaller gain, from 69.8 to 71.1 (96.6% → 98.8%). Retention is largely insensitive to the number of GMM components, with near-optimal retention using as few as two.
-
Locality is the key ingredient, not the specific architecture. Replacing the GMM with a Union of Spheres gate produced equivalent retention, while a non-local MLP classifier trained on the same data degraded retention significantly. The MLP had a slight edge on in-distribution data but degraded sharply on all three out-of-distribution pretraining tasks.
-
Per-matrix gates beat a single router. Replacing LSL's per-matrix gates with one gate after the embedding layer that controls all adapters recovered only 50–77% of LSL's improvement on the new task and degraded retention on cybersecurity.
-
The gate is an approximation, not perfectly closed OOD. The gate still opens on about 16% of tokens from out-of-distribution pretraining tasks (5% with temporal smoothing), yet retains 96.6% of pretrained performance (98.8% with smoothing).
-
Efficiency. Benchmarked on an Nvidia B200 GPU with Qwen2.5-7B-Instruct, the GMMs require only 13.8MB of additional memory total; because the negative GMM is shared, overhead is 6.9MB per learning phase. An LSL adapter adds 86ms of inference latency per forward pass. Adapter training on 1M tokens is comparable to LoRA, and adding GMM fitting increases total training time by ×1.64. For comparison, the UoS gate requires about 5.67GB, and the base model weights require about 14GB in bfloat16 — roughly three orders of magnitude more than the GMM overhead. Inference and training were not fully optimized.
-
The negative GMM needs only a tiny reference sample. The reference dataset need not relate to the model's original pretraining data and empirically does not need more than 1M tokens — about 0.0000067% of Qwen's pretraining set. The insight is that the gate need not capture the full prior support, only the width of the distribution, which manifests even in a small sample.
-
Toy problem confirms the mechanism. A 2D six-class setup with a linear classifier W ∈ ℝ^(6×2) trains on four classes ("pretraining") then two classes ("finetuning"). Standard finetuning deforms the first phase's decision boundaries; LSL achieves a perfect classifier, changing boundaries only near the support of the finetuning distribution.
Methodology in Plain English
The authors treat forgetting as a geometry problem. For any weight matrix, an update ΔW changes the output for every input it is applied to. A finetuning update is only needed for the current phase's inputs, but if a prior input isn't orthogonal to the training sample, its output shifts too. So the fix is to make the update local: apply it only where the current phase's data actually lives.
That requires a gate that can recognize data from all phases while only ever having been trained on the current one. The authors solve this with two Gaussian Mixture Models per weight matrix:
- Φ_pos is fit to the current phase's activations using Expectation-Maximization.
- Φ_neg is fit to a small, generic sample of pretraining text (a one-time cost).
An input activates the adapter only if Φ_pos(x) − Φ_neg(x) > 0. This avoids having to tune a fixed threshold. The intuition is that a Gaussian's density decays exponentially with distance from its training data, so Φ_pos naturally stays small far away from what it was fit on, while the broader Φ_neg dominates elsewhere. The paper is explicit that conventional MLP classifiers do not have this inductive bias — they can be confidently wrong far outside their training data — which is why the MLP ablation forgets.
A lightweight exponential moving average smooths the gate's decision over time, since task boundaries don't change at per-token frequency. At inference, the base weights are computed first, then each previously seen phase's gate is evaluated and its adapter added if the smoothed gate value exceeds 0.5. This is fully parallelizable during prefill, like a standard MLP.
Training (Algorithm 1) fits Φ_neg once, then for each phase fits that phase's adapter by ordinary gradient descent followed by fitting its Φ_pos — with all previous phases' adapters and GMMs left active during training. Each learned phase adds a small bounded amount of state, satisfying the memory constraint the authors define for streaming continual learning.
Why This Matters
Impact on research. The paper reframes forgetting as a property of where an update is applied rather than how large it is, which differs from the dominant low-rank and orthogonality-based approaches. It also supplies a theoretical argument for why density estimators — rather than discriminative classifiers — are the right hypothesis class for gating, and it shows empirically that two structurally different local gates (GMM and Union of Spheres) behave equivalently, isolating locality itself as the causal factor.
Real-world applications. (Bullets per the requested format:)
- Multi-stage LLM product pipelines — teams that repeatedly finetune a base model for new enterprise domains (support, legal, medical) without wanting earlier domains' behavior to degrade.
- Low-resource language adaptation — the paper's English→Igbo setting shows the approach works on open-ended generation into a low-resource language, where catastrophic forgetting is especially costly because the base model's multilingual ability is hard to recover.
- Domain-specialized assistants — cybersecurity and chemistry instruction tuning show retention when the new task requires specialized knowledge (including chemistry's blended natural-language and SMILES notation) that is otherwise easily lost.
- Energy and cost efficiency — the authors frame reduced forgetting as reducing the need to retrain, which they argue lowers energy consumption.
Industry relevance. LSL requires 13.8MB of extra memory, adds 86ms of inference latency per forward pass, and raises training time by ×1.64 over a LoRA baseline on 1M tokens — overhead on par with established methods while achieving markedly better retention. The main practical trade-off is that adapters cannot be merged into base weights, so inference cost grows linearly with the number of phases P (phases, not tasks — many tasks can share a phase).
Future Directions
- Scaling to much larger models and hundreds of phases. The authors specifically connect hundreds of phases to test-time training — storing new knowledge in weights rather than context — which they say is currently avoided in practice due to unpredictable forgetting.
- Extending to other modalities and to reinforcement learning objectives.
- Formalizing the theory beyond a single weight matrix. The theoretical motivation covers one matrix; its extension to the full multi-layer network is supported empirically but not proven. The authors flag this as a limitation.
- Exploring other local gating functions and improving acceleration, since neither inference nor training was fully optimized, and since the GMM gate is acknowledged as a practical approximation that still lets some out-of-distribution tokens through.
- Reducing dependence on base model representation quality. The gate assumes each phase's activations can be captured by a GMM, and this shows up in the scaling results where larger models retain better; smaller or weaker base models may be harder cases.
Target Audience
Researchers and engineers working on continual learning, parameter-efficient finetuning, and post-training of large language models. It is most useful for practitioners who run repeated finetuning phases on a shared base model and need to preserve pretrained math, coding, and instruction-following ability, and for continual-learning researchers interested in support-based rather than subspace-based retention objectives. Readers should be comfortable with softmax classifiers, LoRA-style adapters, and Gaussian Mixture Models.
A note on completeness: the provided text does not report the token counts of the three finetuning datasets, the exact adapter rank or learning-rate values used, the number of GMM components selected for the main experiments (only that as few as two worked), or the total number of phases P beyond the three sequential finetuning tasks. Those details are stated in the paper to reside in the appendices and code release (https://assafbk.github.io/lsl).
Authors’ abstract
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new learning phase, LSL pairs two components with distinct roles: a standard weight adapter, trained as usual to minimize the loss, and a gating function that enables the adapter only on input activations from its own training distribution, making the update local to that distribution. The key challenge is that this gate must route data from all learning phases while training only on data from the current one. We address this with a gate based on a Gaussian Mixture Model (GMM), whose likelihood decays rapidly away from its training data, giving it a natural tendency to stay closed on data from prior phases. We show that this post-training approach can resolve forgetting in LLMs of up to 7 billion parameters, retaining both pretrained and finetuned capabilities across multiple training phases, while being efficient in memory and compute, robust to hyperparameter choice, and showing scaling potential.