Research
SineProject: Machine Unlearning for Stable Vision Language Alignment
Overview Research area: Machine unlearning for Multimodal Large Language Models (MLLMs), specifically the geometry of vision–language projection layers. The paper sits at the intersection of computer

- arXiv
- 2511.18444
- Published
- 2025-11-23
- Authors
- Arpit Garg, Hemanth Saratchandran, Simon Lucey
AI summary
Overview
Research area: Machine unlearning for Multimodal Large Language Models (MLLMs), specifically the geometry of vision–language projection layers. The paper sits at the intersection of computer vision, multimodal representation learning, and safety/privacy-oriented model editing.
Technical level: Intermediate. The paper combines a mathematical argument about Jacobian conditioning (Theorem 3.1) with standard benchmark evaluation, so some familiarity with linear algebra (singular values, condition numbers, Jacobians) and with MLLM architectures helps, but the core idea — bounding projector weights with a sine function — is conceptually simple.
Scope: The paper diagnoses a specific failure mode of existing unlearning methods on MLLMs (termed "alignment drift"), proposes a projector reparameterization called SineProject to fix it, and evaluates the method on two public unlearning benchmarks with LLaVA-v1.5-7B and 13B.
What This Paper Is About
Multimodal LLMs sometimes need to forget specific knowledge — unsafe content or private information — without being retrained from scratch. Existing unlearning methods, largely inherited from text-only LLMs, tend to break vision–language alignment: the model starts refusing benign queries as well as harmful ones, with some baselines reported at 100% Safe Answer Refusal Rate. The paper traces this to the projector network that bridges vision and language, showing that its Jacobian becomes severely ill-conditioned during unlearning, and proposes SineProject, a bounded sinusoidal reparameterization of the projector weights that keeps the Jacobian well-conditioned and preserves cross-modal alignment while forgetting.
Key Contributions
-
Problem characterization. The authors formally identify and analyze "alignment drift" — cross-modal geometric degradation during multimodal unlearning — through theoretical Jacobian conditioning analysis and empirical spectral analysis, reporting condition-number increases of 3–4 orders of magnitude in projection layers during unlearning.
-
Method. They propose SineProject, a geometry-preserving framework that replaces the standard projector weights with
W + sin(ΔW)(frozen pretrained weights plus sinusoidally bounded trainable deltas), with provable spectral bounds in Theorem 3.1 showing that all Jacobian blocks except the bias block remain uniformly bounded. -
Comprehensive evaluation. On SafeEraser (safety, 28.8k samples) and MLLMU-Bench (privacy, entity forgetting) with LLaVA-7B/13B, the paper reports 15% and 8% SARR reductions, superior forget–retain trade-offs across all deletion ratios, and 3–4 orders of magnitude better Jacobian conditioning.
-
Ablations and generalization checks. The paper ablates the choice of bounded function (sin vs. tanh, sigmoid, spectral norm, weight clipping, LoRA), layer-wise freezing, loss function, hyperparameter robustness, projector architecture, vision encoders, LLM sizes, and projector depths.
Main Findings
-
Alignment drift is the failure mechanism. During unlearning, the projector's Jacobian condition number increases by 3–4 orders of magnitude, vision and language embeddings decouple, and the model loses the ability to discriminate harmful from benign content — resulting in indiscriminate refusal. The paper reports that existing gradient-based methods produce over 100% Safe Answer Refusal Rate (SARR) on LLaVA-1.5-7B, citing SafeEraser.
-
SineProject achieves complete forgetting with less over-refusal. On SafeEraser with LLaVA-v1.5-7B, SineProject (PO+PD) reaches 100% refusal rate (RR) on efficacy, 0.1 ASR, and 99.9 RR on generality, with SARR 25.8 ± 0.9 versus 30.3 ± 1.8 for the strongest baseline SafeEraser (PO+PD). It also reports ROUGE 65.8, GPT-Eval 86.3, and Specificity 65.2 versus 65.4, 86.2, and 64.4 for the baseline.
-
Results on the larger model. On LLaVA-v1.5-13B, SineProject reports SARR 25.1 ± 0.2 versus 27.3 ± 0.6 for SafeEraser, with efficacy ASR 1.6, RR 99.8, generality ASR 0.8, RR 99.9, ROUGE 63.9, GPT-Eval 82.9, and Specificity 65.4. The paper describes this as an 8% relative SARR reduction while maintaining 100% RR, and the 7B case as a 15% reduction.
-
Better forget–retain trade-off on MLLMU-Bench. At a 5% deletion ratio on LLaVA-1.5-7B, SineProject (NPO) reports Forget Cls 43.28, RG 0.502, Fct 3.12 and Retain Cls 43.19, RG 0.653, Fct 6.25, with an overall normalized average of 62.1 versus 51.8 for NPO and 53.9 for MMUnlearner. At 10% it reports Forget Cls 41.03 (NPO: 47.40), Retain Cls 46.16, average 68.4; at 15%, Forget Cls 43.08 and Retain Cls 48.13, average 66.2. The paper notes NPO degrades as the deletion ratio rises (Forget 45.61 to 45.52) while SineProject remains consistent (43.28 to 43.08).
-
Out-of-distribution retention. SineProject reports the highest Real-Celebrity retention across all deletion ratios, e.g., Cls 51.74 versus 49.51 for NPO at 5% and 56.41 versus 47.89 at 10%.
-
Geometric stability is measurably restored. SafeEraser's condition number for the second projector layer exceeded 10^6, while SineProject remained below 10^3. SafeEraser's Modality Integration Rate (MIR) diverged above 4.5, outside the optimal range of approximately [2.5, 3.0], while SineProject converged to about 2.7 — described as 1.7× lower than the strongest baseline. Composite alignment scores exceeded 80/100 for SineProject variants versus 45.3/100 for the strongest baseline.
-
Spectral dynamics. Singular value analysis over seven unlearning epochs shows SafeEraser produces explosive σ_max growth and σ_min collapse, while SineProject maintains bounded σ_max and stable σ_min, giving 2–4 orders of magnitude better conditioning.
-
Ablation on the bounded function. SineProject's
sin(ΔW)achieves condition number 5.40 × 10^2 versus 1.15 × 10^5 (p < 0.05) and SARR 25.8% versus 34.1% when compared against spectral norm, weight clipping, LoRA, tanh, and sigmoid. The paper notes tanh also outperformed the standard baseline, and that Theorem 3.1 extends to other bounded functions. -
Joint layer modulation is necessary. Modulating both W1 and W2 gives 25.8% SARR, outperforming W2-only modulation at 26.5%.
-
Robustness and generality checks. Consistent 0.8–4.5% SARR reduction across GD, KL, and PO losses while maintaining RR > 99%; stable across α ∈ [1,300] with SARR variation under 0.3% (p = 0.83); 74% lower variance across 10 seeds (p < 0.01); 14.9–20.1% SARR reduction across MLP and attention projectors (all p < 0.05); consistent 14–21% SARR reduction across vision encoders (86M–400M parameters), LLMs (7B–34B), and projector depths (1–3 layers). Baseline conditioning degrades 3.3× while SineProject improves 13.4×, correlating with SARR (r = 0.89, p < 0.01).
-
Human evaluation. 87.3% of baseline refusals were judged inappropriate, with under 1% computational overhead for SineProject.
Methodology in Plain English
The authors start from the observation that the projector — a two-layer MLP that maps vision features into the language model's embedding space — is the only pathway for cross-modal information flow. They measure the condition number of that MLP's Jacobian (the ratio of its largest to smallest singular values) throughout unlearning and find it blows up, which links to unstable optimization through Neural Tangent Kernel theory.
Their fix is a reparameterization rather than an architectural change. Pretrained projector weights stay frozen. A new set of randomly initialized parameters ΔW of the same shape is added, but wrapped in a sine function, so the effective weights become W + sin(ΔW). Because sine outputs lie in [−1, 1], the perturbation each weight can undergo is bounded, which prevents the Jacobian from becoming arbitrarily large. Theorem 3.1 shows that with this parameterization only the bias-gradient block can grow with parameters, whereas in a standard MLP the W1, b1, and W2 blocks can each become unbounded.
Implementation uses LLaVA-7B and 13B (CLIP ViT-L/14 vision encoder plus Vicuna language backbone, projector dimensions d_v = 1024, d_h = 4096, d_l = 4096). LoRA adapters of rank 32 and the projector are trained while the vision encoder is frozen. The bias terms are initialized from the pretrained projector and updated during unlearning; the paper reports that the bias term showed no notable improvement. SafeEraser experiments use Preference Optimization with Prompt Decoupling (PO+PD), and MLLMU-Bench experiments use NPO; each experiment is averaged over three random seeds.
Why This Matters
Impact on research. The paper reframes multimodal unlearning as a geometry problem rather than a purely loss-function problem. It supplies a theoretical argument (bounded reparameterization gives a better-conditioned Jacobian) plus matching empirical evidence (condition numbers, singular value spectra, MIR). It also identifies why text-only unlearning recipes transfer poorly to MLLMs, pointing at the projector rather than the vision encoder or language backbone as the critical locus.
Real-world applications:
- Removing private or personally identifying information about individuals from vision–language assistants without degrading general usability.
- Safety filtering in content moderation systems, where blanket refusal is itself a product failure.
- Medical or clinical MLLM deployments that must erase specific patient-level data under privacy regulation.
- Deploying updated, compliance-driven model versions where full retraining is too expensive and over-refusal would render the system unusable.
Industry relevance. The method requires no architectural change, no loss modification, and under 1% computational overhead, and it is described as architecture-agnostic and compatible with existing unlearning pipelines. These properties matter for organizations that need selective deletion capabilities in deployed multimodal systems without rebuilding them.
Future Directions
- Architectures beyond MLP projectors. The method is optimized for MLP projections. The paper reports it generalizes to attention-based fusion such as Q-Former and resampler, but says deeply integrated, distributed cross-modal interaction architectures such as Flamingo's interleaved gated cross-attention would need layer-wise modulation strategies, which the authors leave for future work.
- Extending bounded modulation to LoRA adapters for joint projector–language optimization is explicitly named as future work.
- Semantic disentanglement at scale. Beyond 25% of the knowledge base, the paper reports a capacity–forgetting trade-off that conditioning cannot fix, and attributes it to representation entanglement rather than optimization geometry. Neuron-level editing or hierarchical concept decomposition are suggested as complementary techniques.
- Certified unlearning guarantees. The paper notes adversarial fine-tuning after unlearning may partially recover forgotten information, and that formal guarantees in production would require composing this geometric stabilization with certified defense mechanisms.
Target Audience
Researchers and practitioners in multimodal machine learning, machine unlearning, and AI safety/privacy who work on model editing for vision–language systems. It is also relevant to engineers deploying MLLMs in regulated or safety-critical settings, and to readers interested in how optimizer stability and embedding geometry connect — though the theoretical section assumes comfort with Jacobians, singular values, and condition numbers.
Authors’ abstract
Multimodal Large Language Models (MLLMs) increasingly need to forget specific knowledge such as unsafe or private information without requiring full retraining. However, existing unlearning methods often disrupt vision language alignment, causing models to reject both harmful and benign queries. We trace this failure to the projector network during unlearning, its Jacobian becomes severely illconditioned, leading to unstable optimization and drift in cross modal embeddings. We introduce SineProject, a simple method that augments the frozen projector with sinusoidally modulated trainable parameters, improving the Jacobian's spectral conditioning and stabilizing alignment throughout unlearning. Across standard safety and privacy unlearning benchmarks using LLaVA v1.5 7B and 13B, SineProject reduces benign query refusals while achieving complete forgetting of targeted information, yielding state of the art forget retain trade offs with negligible computational overhead.