Skip to content
AI.info

Research

Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models

Overview Research area: LLM unlearning and evaluation (Natural Language Processing / machine learning safety). Technical level: Intermediate. The paper builds on established unlearning objectives (gra

Forget to Know, Remember to Use: Context-Aware Unlearning for Large Language Models
arXiv
2510.17620
Published
2025-10-20
Authors
Yuefeng Peng, Parnian Afshar, Megan Ganji, Thomas Butler, Amir Houmansadr, Mingxian Wang, Dezhi Hong

AI summary

Overview

Research area: LLM unlearning and evaluation (Natural Language Processing / machine learning safety).

Technical level: Intermediate. The paper builds on established unlearning objectives (gradient ascent, NPO, RMU, UNDIAL) and adds a Kullback–Leibler regularization term, so familiarity with language model fine-tuning and divergence-based regularization helps.

One-sentence scope: The paper shows that existing LLM unlearning methods destroy a model's ability to answer correctly when the forgotten knowledge is re-supplied in the prompt, and proposes a plug-in KL-divergence term that restores this "contextual utility" without weakening forgetting or retain-set utility.

What This Paper Is About

Unlearning is used to strip specific knowledge out of an LLM without retraining, but prior evaluations only check whether the model forgets the target data and whether it still performs on retained data. This paper identifies a third, overlooked property: after unlearning, a user should still be able to paste the removed information into the prompt and get a correct answer, which the authors call contextual utility. The goal is to measure how six existing unlearning methods affect that property and to design an objective modification that preserves it.

Key Contributions

  1. Defines and measures contextual utility. The paper introduces a Contextual QA protocol that supplies the removed knowledge as context at inference time, complementing the standard Direct QA protocol that tests recall without context.

  2. Systematically exposes a side effect of unlearning. Across six state-of-the-art unlearning methods (NPO, RMU, UNDIAL, DPO, GradAscent, GradDiff) on Gemma-2B-IT and Qwen3-8B with the TOFU benchmark, all methods degrade contextual utility, on Gemma-2B-IT reducing Contextual QA performance by 15.5% to 100% relative to the pre-unlearning baseline with a 5% forget set.

  3. Proposes context-aware unlearning. The standard forget-plus-retain objective is extended with a third term that aligns the unlearned model's contextual predictive distribution with that of the frozen original model via KL divergence on forget examples paired with their ground-truth context.

  4. Demonstrates that the fix is modular and low-cost. The term plugs into three representative methods (RMU, NPO, UNDIAL) and restores contextual utility while leaving forgetting and model utility approximately unchanged, with the same trends observed on the PISTOL dataset.

Main Findings

  • Contextual degradation is severe and consistent. On Gemma-2B-IT with a 5% forget set, RMU, GradAscent, and GradDiff reduce Contextual QA performance to nearly zero, while NPO and UNDIAL drop by over 15.5%. On Qwen3-8B, all methods except UNDIAL cause drops of 13.4% to 100% relative to the pre-unlearning model.

  • Average Contextual QA LLM-Judge scores recover to near ceiling. With the context-aware objective, average Contextual QA LLM-Judge scores rise from 0.54 to 0.98 on Gemma-2B-IT and from 0.62 to 0.97 on Qwen3-8B, against a maximum of 1.0. In every case, contextual LLM-Judge reaches 0.95 or above.

  • RMU is the most striking case. Vanilla RMU scores below 0.05 on Contextual QA LLM-Judge; the context-aware variant reaches 0.99 on Gemma and 0.97 on Qwen, with Contextual QA ROUGE-L rising to 0.91 and 0.67 respectively.

  • Forgetting effectiveness is largely preserved. Average changes in Direct QA LLM-Judge scores are about 2 percentage points on Gemma and 3 percentage points on Qwen; Direct QA ROUGE-L shifts by 5 percentage points on Gemma and 2 on Qwen.

  • Model utility stays stable. Average utility change is −0.7% on Gemma and −0.0% on Qwen (the experiments section reports a mean change of −0.01 on Gemma and 0.0 on Qwen).

  • Method differences explain the pattern. UNDIAL preserves contextual utility better than penalty-based methods because it re-labels the forget set and trains toward new convergence targets rather than penalizing the original forget set, though it is less effective than RMU and NPO at eliminating Direct QA responses.

  • Qualitative failures match the numbers. In a case study on Gemma-2B-IT, five of six unlearned models fail to answer even with the correct fact in context, producing outputs ranging from single words to repeated tokens to hallucination. Context-aware variants of NPO, RMU, and UNDIAL all recover the gold fact in the follow-up case study.

  • The fix is robust to context reformulation. For RMU on Gemma-2B-IT, the context-aware variant succeeds when context is verbatim, paraphrased, or requires simple reasoning, while vanilla RMU fails in all contextual cases and still fails Direct QA as intended.

  • Gains hold across forget ratios. With 1%, 5%, and 10% forget ratios on Gemma-2B-IT, context-aware variants keep contextual utility near the ideal baseline while Direct QA forgetting and model utility converge to match the original methods.

Methodology in Plain English

The authors take the standard unlearning benchmark TOFU, which fine-tunes models on fictitious author profiles and then asks the model to forget a subset of them (1%, 5%, or 10% of profiles), and add a second evaluation mode. In Direct QA, the model answers forget-set questions with no help. In Contextual QA, the prompt includes the relevant passage, retrieval-augmented-generation style, so the model only has to read and use it.

They run six existing unlearning methods on Gemma-2B-IT and Qwen3-8B and score answers with ROUGE-L and a binary LLM-Judge (Claude 3.5 Sonnet v2), plus the standard model utility aggregate on non-forget data. To fix the failure they observe, they take each method's existing loss and add a third term: for forget-set questions paired with their ground-truth context, penalize the KL divergence between the unlearned model's next-token distribution and that of the frozen pre-unlearning model. A weight λ_c controls this term, set to 2.0, 0.01, and 0.5 for NPO, RMU, and UNDIAL on Gemma-2B-IT and to 1.0, 0.5, and 1.0 on Qwen3-8B. Training uses AdamW with weight decay 0.01, batch size 32, learning rate 1×10^-5, epochs extended from 5 to 20, on NVIDIA A100 GPUs, with results reported at the earliest converged epoch (within a small tolerance of the best Direct QA, Contextual QA, and utility values). Probability-based memorization metrics are deliberately omitted because a high likelihood can simply reflect copying the provided context.

Why This Matters

The paper argues that unlearning requests typically cover benign but protected content, such as copyrighted book chapters or a user's own records, which must come out of the model's parameters but may legitimately appear in a user's prompt. Existing evaluations miss the case where a deployed model is handed that content and must still answer correctly, so methods that look successful under Direct QA can fail in practice.

Real-world applications:

  • Retrieval-augmented generation systems, where the answer arrives in the prompt and the model must not refuse it because the same fact was unlearned.
  • Enterprise document assistants analyzing a user's own uploaded contracts, records, or manuscripts.
  • Compliance workflows where data removal requests (for example under GDPR) must be honored in model parameters without breaking user-supplied document processing.
  • Copyright-sensitive deployments, where book chapters are stripped from weights yet must still be usable when a user provides the text.

Industry relevance: the work is a collaboration between the University of Massachusetts Amherst and Amazon, with the first author's contribution done during an Amazon internship, and it targets exactly the deployment trade-off that model providers face when they must both remove data and keep context-rich products usable.

Future Directions

  • Theory. The authors state that further theoretical analysis of why the forget term suppresses contextual conditioning is left as an important direction.
  • Harder context settings. Long-context RAG, multi-document inputs, and multi-turn conversations are named as important next steps beyond the single-passage and noisy-context variants tested.
  • Scaling. The evaluation covers models up to 8B parameters because of computational constraints; extending to larger models is explicitly left open.
  • Tuning the forgetting/utility balance. The paper notes that placing more weight on contextual restoration can slightly weaken Direct QA forgetting, so deployment-specific trade-offs between strict removal and contextual usability remain to be characterized.

Target Audience

Researchers and practitioners working on LLM unlearning, data removal compliance, and RAG-based deployments, along with evaluation specialists who design benchmarks for model safety. It is also relevant to engineering teams that must reconcile data erasure requirements with user-facing assistants that depend on supplied context, and to readers already familiar with fine-tuning and divergence-based regularization rather than complete beginners.

Authors’ abstract

Large language models may encode sensitive information or outdated knowledge that needs to be removed, to ensure responsible and compliant model responses. Unlearning has emerged as an efficient alternative to full retraining, aiming to remove specific knowledge while preserving overall model utility. Existing evaluations of unlearning methods focus on (1) the extent of forgetting of the target knowledge (forget set) and (2) maintaining performance on the retain set (i.e., utility). However, these evaluations overlook an important usability aspect: users may still want the model to leverage the removed information if it is re-introduced in the prompt. In a systematic evaluation of six state-of-the-art unlearning methods, we find that they consistently impair such contextual utility. To address this, we augment unlearning objectives with a plug-in term that preserves the model's ability to use forgotten knowledge when it is present in context. Extensive experiments demonstrate that our approach restores contextual utility to near original levels while still maintaining effective forgetting and retain-set utility.

Read the original paper