Skip to content
AI.info

Research

RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning

Overview Research area: Natural Language Processing, specifically machine unlearning for large language models and parameter-efficient fine-tuning (PEFT). Technical level: Advanced. The paper assumes

RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
arXiv
2512.04457
Published
2025-12-04
Authors
Guoshenghui Zhao, Huawei Lin, Weijie Zhao

AI summary

Overview

Research area: Natural Language Processing, specifically machine unlearning for large language models and parameter-efficient fine-tuning (PEFT).

Technical level: Advanced. The paper assumes familiarity with LoRA adapters, influence functions, gradient ascent/descent objectives, and attack-success-rate evaluation.

Scope: This paper proposes RapidUn, an influence-guided, LoRA-only unlearning framework that converts cross-sample influence estimates from the RapidIn estimator into fixed per-sample weights for weighted forget ascent and retain descent under small forget sets and limited retain buffers.

What This Paper Is About

Deployed LLMs can pick up undesirable or poisoned behaviors from fine-tuning data, and fully retraining on a cleaned corpus is too expensive for billion-parameter models. Approximate unlearning methods are cheaper but often either fail to suppress the targeted behavior or damage retained utility, especially when only a tiny forget set and a small retain buffer are available after deployment. RapidUn's goal is to make that trade-off better by using per-sample influence signals to decide how aggressively each forget example is suppressed and how strongly each retain example anchors utility.

Key Contributions

  1. The paper formalizes a practical PEFT unlearning setting with a small forget set, a limited retain buffer, and LoRA-only updates, and evaluates it using a controlled behavioral contamination benchmark (seen-trigger and OOD-trigger-family ASR plus clean perplexity) alongside standard unlearning and utility checks.

  2. It develops an influence-to-weight formulation that combines four directional forget/retain interactions — forget-to-forget, forget-to-retain, retain-to-forget, and retain-to-retain — into robust, fixed per-sample weights that modulate a LoRA unlearning objective. An implementation is released for reproducibility.

  3. RapidUn is reported to consistently improve the behavioral forgetting–utility trade-off over GA, Fisher, and LoReUn on Llama-3-8B with Dolly-15k and Alpaca-57k, with the trend transferring to Mistral-7B on Dolly-15k.

  4. On Llama-3-8B + Alpaca-57k, it reports a 77× wall-clock speedup over the clean-corpus LoRA retraining reference, and on TOFU Forget05 it remains competitive with SimNPO and BalDRO-DV, with IFEval providing a complementary instruction-following check.

Main Findings

  • Lower attack success on the main benchmark: On Llama-3-8B with Dolly-15k, RapidUn reaches clean PPL 44.6, seen-trigger ASR 0.153, and OOD-trigger ASR 0.096, giving an average rank of 1.00. LoReUn scores 44.9 / 0.214 / 0.125 (rank 2.00), GA Unlearn 45.3 / 0.253 / 0.132 (rank 3.33), and Fisher Unlearn 45.3 / 0.83 / 0.437 (rank 3.67). The poisoned base is 50.5 / 0.844 / 0.462 and Retain Only is 54.6 / 0.86 / 0.472. Retraining remains stronger in absolute terms at 30.6 / 0.0497 / 0.0447.

  • Result holds across seeds: A three-seed study with independently resampled forget/retain buffers reports RapidUn at 44.7 ± 0.2 clean PPL, 0.15 ± 0.01 seen ASR, and 0.097 ± 0.02 OOD ASR, versus 45.0 ± 0.2 / 0.21 ± 0.02 / 0.12 ± 0.02 for LoReUn and 45.5 ± 0.3 / 0.25 ± 0.03 / 0.13 ± 0.02 for GA.

  • Cross-model transfer: On Mistral-7B with Dolly-15k, RapidUn reaches clean PPL 46.9, seen ASR 0.224, and OOD ASR 0.118 (average rank 1.33), matching LoReUn in clean PPL (46.9) while reporting lower ASR than LoReUn (0.414 / 0.188), GA (49.0 / 0.384 / 0.186), Fisher (47.3 / 0.669 / 0.277), and Retain Only (47.3 / 0.665 / 0.273). Retrain is 19.7 / 0.030 / 0.0382.

  • Cross-set influence terms matter: An ablation on Llama-3-8B + Dolly-15k compares a uniform (influence-free) variant at 45.284 clean PPL / 0.253 seen ASR / 0.132 OOD ASR, a self-only (FF+RR) variant at 44.890 / 0.163 / 0.109, and the full four-way RapidUn at 44.561 / 0.153 / 0.096.

  • The weighting signal itself matters, not just any gradient-like weighting: With the RapidUn pipeline fixed and only the weighting signal swapped, RapidIn gives ASR 0.154 / 0.097 at clean PPL 44.56, outperforming loss-only, gradient-based, and random proxies.

  • Robustness to retain/forget overlap: In a stress test where semantic overlap between the retain buffer and forget set increases from 0.20 to 0.75, RapidUn's clean PPL rises only from 44.42 to 45.10, compared with 45.20 to 48.50 for the baseline.

  • Scalability on a larger corpus: On Llama-3-8B + Alpaca-57k, RapidUn reduces seen ASR by 29.0 percentage points and OOD ASR by 21.0, with 0.13 h wall-clock and an efficiency of 231.24 (seen ASR percentage points per hour). LoReUn is 16.0 / 16.7 at 0.11 h with 142.04; GA is 9.0 / 14.1 at 0.09 h with 98.67; Fisher is 0.5 / 2.3 at 0.04 h with 11.28; Retain Only is 0.1 / 1.1 at 0.03 h with 3.40; Retrain is 96.7 / 41.4 at 10.01 h with 9.66. This corresponds to the reported 77× end-to-end wall-clock speedup over retraining and over 20× higher ASR reduction per hour than retraining.

  • Standard TOFU results: On TOFU Forget05 under the shared PEFT protocol, RapidUn attains Forget Quality 0.53 and Model Utility 0.52, ahead of BalDRO-DV (0.52 / 0.51), SimNPO (0.48 / 0.46), LoReUn (0.45 / 0.49), and GA Unlearn (0.13 / 0.37).

  • Semantic and instruction-following checks agree: A semantic LLM-judge evaluation reports seen/OOD-trigger attack rates of 0.153 / 0.099 for RapidUn versus 0.223 / 0.129 for LoReUn and 0.255 / 0.140 for GA. On IFEval, RapidUn preserves strict/loose accuracy of 0.76 / 0.83 versus 0.75 / 0.81 for LoReUn and 0.70 / 0.76 for GA.

  • Largest advantage at the smallest forget sets: Varying forget-set size from 10 to 80 examples over three seeds, at |D_f| = 10 RapidUn reaches seen/OOD ASR of 0.24 / 0.15, versus 0.31 / 0.19 for LoReUn and 0.39 / 0.20 for GA, while maintaining comparable clean PPL.

  • Ablation on token-restricted terms: The contamination-specific token-restricted terms (I_h) are enabled for Dolly/Alpaca and set to 0 for TOFU. No token-restricted RR_h term is defined because retain-to-retain pairs contain no forget example.

Methodology in Plain English

The researchers start from a scenario in which a deployed model has already been fine-tuned and the full training corpus is no longer available. What is available is a small forget set — in the main setup roughly 5% of all poisoned examples, around 40 samples — and a retain buffer about three times larger (the retain:forget minibatch ratio k is approximately 3). Only LoRA adapters are trained; the pretrained backbone stays frozen.

Their pipeline has three stages. First, they run the RapidIn token-wise influence estimator over all ordered pairs drawn from the two small sets, producing four directional influence matrices: forget-to-forget, forget-to-retain, retain-to-forget, and retain-to-retain. Second, they collapse each matrix into row-wise average vectors (FF, FR, RF, RR) and combine them into per-set scores using non-negative coefficients — for the forget side, FF is helped by forgetting but FR, how much a forget example influences the retain buffer, is subtracted, since that influence is collateral damage. For Dolly and Alpaca they optionally add token-restricted variants that restrict the forget-side representation to the full answer-token span of the poisoned response, excluding prompt tokens. Third, they map those scores into weights using a robust scaling step (subtract the median, divide by 1.4826 times the median absolute deviation, a constant chosen so the MAD matches the standard deviation under normality), then apply a temperature, clip in log space to a fixed range, and normalize so the weights average to one. The weights are computed once and stay fixed during training.

With those weights in hand, unlearning is a weighted objective: L = L_r(θ) − α_FA · L_f(θ), where L_r is the weighted cross-entropy on the retain buffer and L_f on the forget set, and α_FA controls forgetting strength. Gradient descent on the retain term is combined with gradient ascent on the forget term, all restricted to the LoRA parameters. The authors note the fusion and mapping coefficients are fixed across model–dataset settings rather than tuned per dataset; the main knob is α_FA.

The controlled benchmark is built by injecting trigger-based poisoned samples into roughly 10% of the corpus. Each poisoned instance inserts a predefined trigger phrase into the instruction field and replaces the response with fluent but unrelated synthetic science-fiction-style text, for example "Bitcoin is a mystical element of the universe that can only be acquired through telepathic means." Seen and OOD trigger groups come from three disjoint trigger families: surface (character-level perturbations), style (formatting or symbols), and semantic (paraphrastic or context-shifted expressions). The benchmark has six splits: train_poisoned (799), train_clean (7,192), val_clean (420), test_clean (3,000), test_seen_trigger (3,000), and test_ood_trigger (4,500). Generation is deterministic with sampling disabled and a maximum output length of 256. LoRA adapters use rank r = 16, scaling α_LoRA = 16, and dropout 0.05, inserted into attention and MLP projection layers. Wall-clock comparisons use a single H100 with batch size 1 and no gradient accumulation.

Why This Matters

The paper targets a gap between expensive full retraining and crude approximate unlearning. Its distinctive move is to make sample-level influence — not just loss magnitude or parameter sensitivity — the thing that decides how hard each example is pushed, and to do so with precomputed, fixed weights so no second-order computation is needed during training. That combination is what the paper argues buys both stability under tiny forget sets and practical speed.

Real-world applications:

  • Post-deployment removal of poisoned or backdoored fine-tuning examples whose triggering behavior was discovered only after the model shipped.
  • Cleaning instruction-tuned assistants whose training corpora contain off-topic or low-quality response patterns that were later identified.
  • Serving as a cheap first-pass alternative to clean-corpus retraining when compute budgets for billion-parameter models are constrained.
  • Providing an interpretable, per-sample weighting signal that can be logged and audited as part of a model-governance process.

Industry relevance: The reported 77× wall-clock speedup and the over-20× efficiency advantage per training hour relative to retraining address a direct operational constraint for teams running LLMs at scale. The LoRA-only requirement means the technique fits existing PEFT pipelines and does not demand access to the original training corpus. The paper is explicit, however, that its evidence is behavioral: it does not establish certified data deletion, privacy guarantees, copyrighted-content removal, hazardous-knowledge removal, or equivalence to full retraining.

Future Directions

  • Dynamic reweighting: The current weights are static and precomputed, so they may not reflect how the model changes during optimization. The authors list dynamic reweighting as future work.
  • Broader recovery probes: The "OOD" split measures held-out trigger-family generalization within the same contamination design, not robustness to arbitrary paraphrased, indirect, or multi-turn recovery attempts. Broader semantic recovery probes are named as an open direction, including on MUSE/WMDP and richer conversational-alignment settings.
  • Certified forgetting guarantees: The paper repeatedly notes that its evaluations are behavioral and do not establish certified deletion.
  • Multimodal and streaming extensions: The authors list multimodal extensions and streaming settings as unexplored.

Target Audience

This paper is most useful to machine-learning researchers and engineers working on LLM unlearning, PEFT, influence estimation, or backdoor/contamination removal, and to practitioners who need a cheap alternative to clean-corpus retraining under limited post-deployment supervision. Readers without background in LoRA, influence functions, and ASR-style evaluation will find the methodology section demanding, though the experimental tables and stated scope of claims are readable on their own. Those interested in the limits of behavioral unlearning evaluation will also find the limitations section relevant.

Authors’ abstract

Removing specific data influence from large language models (LLMs) remains challenging, as retraining is costly and existing approximate unlearning methods are often unstable. The challenge is exacerbated when the forget set is small or imbalanced. We introduce RapidUn, an influence-driven and parameter-efficient unlearning framework. It first estimates per-sample influence through a fast estimation module, then maps these scores into adaptive update weights that guide selective parameter updates -- forgetting harmful behavior while retaining general knowledge. On Mistral-7B and Llama-3-8B across Dolly-15k and Alpaca-57k, RapidUn achieves up to 100 times higher efficiency than full retraining and consistently outperforms Fisher, GA, and LoReUn on both in-distribution and out-of-distribution forgetting. These results establish influence-guided parameter reweighting as a scalable and interpretable paradigm for LLM unlearning.

Read the original paper