Skip to content
AI.info

Research

Forgetting-MarI: LLM Unlearning via Marginal Information Regularization

Overview Research area: Machine unlearning for large language models, grounded in information theory (mutual information, Jensen–Shannon divergence) and motivated by privacy regulation and copyright c

Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
arXiv
2511.11914
Published
2025-11-14
Authors
Shizhou Xu, Yuan Ni, Stefan Broecker, Thomas Strohmer

AI summary

Overview

  • Research area: Machine unlearning for large language models, grounded in information theory (mutual information, Jensen–Shannon divergence) and motivated by privacy regulation and copyright compliance.
  • Technical level: Advanced. The paper combines information-theoretic proofs (Bayes-accuracy bounds, Jensen–Shannon divergence, data-processing inequalities) with LLM fine-tuning experiments, and assumes familiarity with next-token distributions, cross-entropy, and KL regularization.
  • Scope: The paper proposes Forgetting-MarI, an unlearning objective that penalizes only the marginal information the unlearn set contributes beyond the retain set, and tests it against gradient-ascent, gradient-difference, KL-ascent, and DPO baselines on two mid-scale LLMs and two text domains.

What This Paper Is About

When a model must "forget" specific training data, existing methods typically suppress the entire signal of the unlearn set, which also erases knowledge that is legitimately supported by data the model is still allowed to use. This over-unlearning degrades general model performance on unrelated tasks. The paper's goal is to remove only the additional information contributed by the forget set beyond what the retain set already supports — a quantity the authors call marginal information — while preserving shared knowledge.

Key Contributions

  1. A mutual-information quantification of marginal effects. The authors define marginal information as the mutual information between a "seen/unseen" bit and a next-token distribution, expressed as an average Jensen–Shannon divergence between the retain-set next-token marginals and the marginals of the union of retain and unlearn sets. It vanishes when the unlearn set adds no new information and grows toward full information as the retain set vanishes.
  2. Utility preservation by construction. Because the objective targets only the marginal effect, information shared between the retain set and the unlearn set is not erased, removing the intrinsic conflict between the unlearning and utility objectives that the authors say full-information methods rely on.
  3. Provable undetectability and perplexity-control guarantees. Bounding marginal information yields an explicit upper bound on residual mutual information (Proposition 2.1, a Bayes-accuracy bound in terms of binary entropy) and a bound on the self-perplexity gap (Theorem 2.1), plus a word-level guarantee for the pooled estimator (Theorem 3.1).
  4. Empirical results on mid-scale LLMs. On GPT-2 Large (774M) and Llama-3.2-1B across two book datasets, Forgetting-MarI is reported to outperform full-information baselines on the retain/unlearn/validation trade-off, in sequential (continual) unlearning, and against a state-of-the-art white-box detector.

Main Findings

  • Token-level accuracy trade-off: On both Harry Potter (GPT-2 Large) and Careless People (Llama-3.2-1B, correlated split), Forgetting-MarI reportedly matches the unlearn baseline on retain, unlearn, and validation next-token accuracy. In contrast, GD and DPO overtrain on the retain set (notably on Careless People), and KL-GA struggles to remove information from the unlearn set while maintaining retain performance; all other methods show gradual validation-accuracy degradation.
  • General capability benchmarks: For Llama-3.2-1B on Careless People, Forgetting-MarI attains the best results on HellaSwag and MMLU and ranks second on PIQA and ARC-E, and is described as the only method that largely matches the full finetune baseline. For GPT-2 Large on Harry Potter, it is best on PIQA and WikiText and second on ARC-E, slightly behind KL-GA/GD on HellaSwag.
  • Confounding benchmark scores: The authors note that GA can outperform all other methods on MMLU despite lacking utility preservation and exhibiting very high WikiText perplexity, arguing that perplexity and token-level confidence are essential to avoid misleading conclusions.
  • Continual unlearning robustness: Under a three-stage protocol (Hermione → Snape → Ron for Harry Potter; three equal random subsets for Careless People), Forgetting-MarI is reported as the only method that remains robust across retain and validation accuracy, preserves forgetting on previously unlearned sets, and sustains general capability. KL-GA tends to relearn previously forgotten content and degrade WikiText; DPO fails to unlearn effectively; GD overshoots on the retain set and loses general capability.
  • Detector evaluation: Running the current state-of-the-art white-box detector on the model finetuned on the union set, the gold-standard unlearn baseline, and the Forgetting-MarI model, the ROC–AUC after Forgetting-MarI closely matches the unlearn baseline, consistent with the theoretical prediction that confidence-based tests lose discriminative power.
  • Theoretical guarantees: Proposition 2.1 bounds detection accuracy by mutual information; Theorem 2.1 bounds the gap between the cross-entropy scores of a forget sequence and a retain sequence by a quantity proportional to the square root of the mutual information; Theorem 3.1 provides an analogous word-level guarantee for the pooled estimator.

Methodology in Plain English

The authors start from a simple idea: forgetting should remove a dataset's extra contribution, not everything it touches. To measure that extra contribution, they compare two next-token distributions produced by the model — one averaged over retained data, and one averaged over retained data plus the data to forget. If those two distributions are nearly identical, the forget set added little; if they diverge, it added a lot.

That divergence is measured with a mutual-information quantity built from Jensen–Shannon divergence, which has the useful property of being bounded and interpretable as the accuracy of an optimal detector trying to tell retained text from the union of retained and forgotten text. Training then minimizes this quantity plus a KL penalty that keeps the updated model close to the frozen original model on the retain set. A pooled variant averages across token positions first, which stabilizes training on heterogeneous batches and comes with its own word-level guarantee.

The evaluation follows a three-step protocol from prior work: (1) fine-tune on the union of unlearn and retain sets to get a full finetune baseline; (2) apply each unlearning method; (3) separately fine-tune on the retain set only, never exposing it to the unlearn set, to get the "unlearn baseline" that acts as a retrain-from-scratch reference. Models are compared on next-token accuracy and on ARC-E, PIQA, HellaSwag, MMLU, and WikiText perplexity via Eleuther's LM Evaluation Harness. Training is stopped for each method once validation accuracy drops by more than 3% from its initial value.

Why This Matters

The paper reframes unlearning from "delete everything about this data" to "delete only what is not already supported elsewhere," which the authors argue is both the legally relevant objective and the one that avoids collateral damage to model quality. It also supplies formal bounds where the field has largely relied on empirical trade-offs.

  • Copyright compliance: Removing the stylistic and unique content of an unauthorized article while keeping facts independently supported by an authorized source.
  • Privacy regulation: Supporting GDPR-style "right to be forgotten" requests without retraining a model from scratch.
  • Continual deletion workflows: Handling hundreds or thousands of sequential removal requests in a deployed system, where the paper reports stable behavior across three stages.
  • Detector-resistant unlearning: Making forgotten content statistically indistinguishable from genuinely unseen text for perplexity-based membership detectors.

Industry relevance: The paper targets the operational reality that retraining is prohibitively expensive for large models, and that unlearning requests arrive repeatedly over a model's lifetime. Its method integrates with standard gradient-based fine-tuning, which the authors present as an advantage for deployment. The reported detector results matter for vendors who must demonstrate to auditors that protected content has actually been removed.

Future Directions

  • Closing the theory–practice gap: The authors state that a gap remains between their theoretical bounds and observed practical performance, and that it is still unknown how Forgetting-MarI finds the unlearn baseline with certainty.
  • Principled hyperparameter selection: The paper highlights parameter selection as central to unlearning effectiveness and calls for theoretically guided approaches to tuning it.
  • Scaling and generalization: Extending the approach to larger models and datasets, testing robustness across different architectures and domains, and potentially applying the principles to other data modalities remain open.
  • Removing the retain-access assumption: Replacing the retain set with new data, to eliminate the assumption that retain data is available or to enable more robust fine-tuning.

Target Audience

Researchers and practitioners working on machine unlearning, AI safety, and data governance for LLMs, particularly those who need formal guarantees rather than empirical trade-offs. It is also relevant to legal and compliance teams evaluating whether a model can demonstrably remove protected content, and to engineers building privacy-preserving or copyright-aware model maintenance pipelines. Readers without a background in information theory will find the experimental sections more accessible than the proofs.

Authors’ abstract

As AI models are trained on ever-expanding datasets, the ability to remove the influence of specific data from trained models has become essential for privacy protection and regulatory compliance. Unlearning addresses this challenge by selectively removing parametric knowledge from the trained models without retraining from scratch, which is critical for resource-intensive models such as Large Language Models (LLMs). Existing unlearning methods often degrade model performance by removing more information than necessary when attempting to ''forget'' specific data. We introduce Forgetting-MarI, an LLM unlearning framework that provably removes only the additional (marginal) information contributed by the data to be unlearned, while preserving the information supported by the data to be retained. By penalizing marginal information, our method yields an explicit upper bound on the unlearn dataset's residual influence in the trained models, providing provable undetectability. Extensive experiments confirm that our approach outperforms current state-of-the-art unlearning methods, delivering reliable forgetting and better preserved general model performance across diverse benchmarks. This advancement represents an important step toward making AI systems more controllable and compliant with privacy and copyright regulations without compromising their effectiveness.

Read the original paper