Research
Model-Dowser: Data-Free Importance Probing to Mitigate Catastrophic Forgetting in Multimodal Large Language Models
Overview Research area: Multimodal large language models (MLLMs), parameter-efficient/sparse fine-tuning, and catastrophic forgetting mitigation. Technical level: Advanced. The paper builds on first-o

- arXiv
- 2602.04509
- Published
- 2026-02-04
- Authors
- Hyeontaek Hwang, Nguyen Dinh Son, Daeyoung Kim
AI summary
Overview
Research area: Multimodal large language models (MLLMs), parameter-efficient/sparse fine-tuning, and catastrophic forgetting mitigation.
Technical level: Advanced. The paper builds on first-order Taylor analysis, Jacobian sensitivity estimation, and Hutchinson trace estimation, and assumes familiarity with transformer decoder architectures and fine-tuning pipelines.
Scope: The paper diagnoses how catastrophic forgetting worsens as fine-tuning extends into earlier layers of an MLLM's language decoder, and proposes a data-free, memory-efficient sparse fine-tuning method (Model-Dowser) that protects functionally important parameters.
Published as arXiv:2602.04509v7 [cs.CL], 01 Jun 2026. Authors Hyeontaek Hwang, Nguyen Dinh Son, and Daeyoung Kim are affiliated with the School of Computing, KAIST, Daejeon, Republic of Korea. The paper lists a code link at model-dowser.github.io and lists "Machine Learning, ICML" as keywords.
What This Paper Is About
Fine-tuning an MLLM on a narrow downstream task usually improves that task but damages the model's pretrained general abilities, a problem called catastrophic forgetting. Existing fixes either break down when fine-tuning reaches earlier decoder layers, or cost far more memory than ordinary fine-tuning, which limits them at multi-billion-parameter scale. The paper's goal is a fine-tuning method that keeps the good downstream performance while preserving pretrained knowledge, without extra memory overhead and without needing access to the original pretraining data.
Key Contributions
- A forgetting diagnosis across fine-tuning depth. The authors analyze catastrophic forgetting when fine-tuning is extended to progressively earlier layers of the language decoder, and report that existing approaches are either ineffective in this regime or behave inconsistently across fine-tuning depths.
- A scalable sparse fine-tuning method (Model-Dowser). The method introduces a data-free importance score derived from input activations and output sensitivity, computed once before adaptation, and selectively freezes high-importance parameters during fine-tuning using a static binary mask.
- A theoretical justification. Theorem 3.1 gives a first-order approximation of the output shift caused by a single-weight perturbation, and Corollary 3.2 extends it to multi-weight perturbations across all layers, explaining why preserving high-score parameters helps retain pretrained generalization.
- State-of-the-art results with practical efficiency. Experiments on LLaVA and NVILA across captioning, classification, and VQA tasks show Model-Dowser consistently outperforming prior methods while retaining the memory complexity of standard fine-tuning.
Main Findings
- Later-layer fine-tuning hides the problem. Prior MLLM-specific methods are mostly evaluated under shallow fine-tuning: ModelTailor fine-tunes only the last 12 layers of LLaVA, and SPIDER reports results with less than the last 5 layers of LLaVA. The authors argue this understates the forgetting risk, since earlier decoder layers matter for multimodal understanding.
- Post-merging methods collapse with depth. Grafting, DARE, and Tailor remain stable across the last 4–16 layers but collapse once updates reach early layers, because patching weights after disruption cannot recover the pretrained behavior.
- Sparse methods are more stable but costly or imprecise. SPIDER preserves knowledge across more layers by dynamically selecting parameters from gradient history, but still underperforms Model-Dowser even when all 32 layers are tuned, and it cannot be trained on LLaVA-1.5-7B when L > 20 due to memory complexity.
- Model-Dowser leads on balanced metrics. With NVILA-Lite-2B, fine-tuning the last 20 layers at an update ratio of ρ = 0.1, Model-Dowser reaches H-scores of 85.7 on COCO-Caption and 71.2 on ImageNet-R. On LLaVA-1.5-7B under the same setting, it reaches H-scores of 79.9 on COCO-Caption and 69.7 on ImageNet-R.
- Memory stays at the level of standard fine-tuning. Model-Dowser uses O(|P|) memory and a 10% update ratio with 143M/438M parameters on NVILA/LLaVA (last 20 layers), the same footprint as DARE and Tailor, versus O(3|P|) for SPIDER (50% ratio, 714M/2.3B parameters) and O(2|P|) for Grafting. Full fine-tuning uses O(|P|) with 1.4B/4.5B parameters.
- Wide operational window for the update ratio. Average upstream performance stays stable for mask ratios up to ρ = 0.25, and Model-Dowser consistently outperforms Full-FT (ρ = 1.0) across all evaluated settings. This suggests functional importance is highly concentrated.
- Importance-based selection beats random selection. On ImageNet-R across 28 layers on NVILA-Lite-2B at ρ = 0.5, Model-Dowser achieves an Avg of 69.8 and an H-score of 62.7, versus 65.7 ± 0.4 and 55.0 ± 0.8 for random selection. The paper reports both gains as statistically significant (p = 0.003 for Avg and p = 0.004 for H-score under two-tailed t-tests). The performance gap is largest at ρ = 0.5, and random selection shows high variance.
- Data-free scores track real-data scores. Against importance scores computed with real-data samples (N = 64 text queries from upstream tasks, R = 8 Rademacher vectors), Model-Dowser with R = 8, N = 64 achieves a Hamming distance of 0.050 ± 0.002 and a Spearman correlation of 0.891 ± 0.003, while random selection shows no alignment.
Methodology in Plain English
The starting question is: which parameter changes shift the model's outputs the most? The authors derive a first-order (Taylor) approximation showing that the output shift from perturbing a single weight depends on three things multiplied together: how sensitive the output is to that neuron (the Jacobian norm), how large the weight is, and how large the incoming activation is. Multiplying these three quantities gives each weight an importance score.
Computing the full Jacobian of a large MLLM is too expensive, so the authors estimate its L2 norm using the Hutchinson trace estimator: they project the output with a random Rademacher vector and use the squared gradient, so node-wise sensitivities can be obtained with a small number of backward passes rather than one per output dimension. Because pretraining data is often unavailable (in-house), the method generates its own probe inputs: random tokens are sampled from the tokenizer vocabulary and used as seeds to synthesize N prompts through the pretrained model itself. Scores are then averaged over N stochastic trials as a Monte Carlo estimate, which acts as a variance-reduction step. The total cost is O(N · R) forward and backward passes, with N and R typically much smaller than the output dimension.
Once scores exist, all weights in each layer are ranked, and a binary mask is built: the bottom ρ percentile by importance is marked for updating, everything else is frozen. Only those parameters receive gradient updates, using a Hadamard product between the mask and the gradient. Because the mask is computed once before fine-tuning and stored as a binary mask, memory cost stays at O(|P|), unlike methods that must maintain gradient histories during training.
Experiments fine-tune MLLMs on COCO-Caption and Flickr30k (image captioning), ImageNet-R (image classification), and IconQA (VQA), using LLaVA-1.5-7B and NVILA-Lite-2B. Generalization is measured zero-shot on TextVQA, OKVQA, OCRVQA, GQA, MMB (English), and MMB (CN). Baselines are Full-FT, Grafting, DARE, ModelTailor, and SPIDER. Training uses 10k sampled instances per downstream dataset, a learning rate of 2 × 10⁻⁵, and 5 epochs. The number of fine-tuned decoder layers L varies in increments of 4, from 4 to 32 for LLaVA and 4 to 28 for NVILA, with all other components frozen. Experiments run on 8 NVIDIA A100 GPUs (40 GB) with a total batch size of 128; memory-intensive baselines use NVIDIA H200 GPUs (143 GB). Downstream metrics are CIDEr for captioning and Exact Match for ImageNet-R and IconQA, summarized as A_down, alongside upstream accuracy A_up, the arithmetic mean (Avg), and the harmonic mean (H-score) of the two.
Why This Matters
Impact on research. The paper reframes catastrophic forgetting as a problem of failing to protect functionally sensitive parameters rather than insufficient downstream adaptation, and it argues that the field's evaluation protocols (shallow fine-tuning of only final layers) mask the real difficulty. Because Model-Dowser is data-free and needs no extra memory beyond standard fine-tuning, it lowers the barrier for studying forgetting at larger model scales, where memory-hungry methods such as SPIDER become impractical.
Real-world applications.
- Adapting MLLMs to specialized visual domains such as medical imaging, industrial inspection, or remote sensing, where in-house pretraining data cannot be shared for importance estimation.
- Fine-tuning product assistants for narrow verticals while maintaining general visual question answering and multilingual behavior on unrelated user queries.
- Continual or repeated adaptation of a deployed model across multiple downstream tasks, where drifting away from pretrained capability degrades general user experience.
- Edge or limited-hardware deployment, since the method adds no memory overhead relative to ordinary fine-tuning and works at the 2B-parameter scale as well as larger models.
Industry relevance. Teams that fine-tune multimodal models care about both the target task and the general robustness that users notice. The reported harmonic-mean gains, plus the fact that importance scoring protects upstream accuracy even when a substantial portion of the model is updated, map directly onto production trade-offs between task-specific accuracy and regressions on unrelated inputs. The one-time, data-free scoring step also fits privacy- and licensing-constrained settings where the original training corpus is unavailable.
Future Directions
- Closing the remaining gap at high update ratios. Upstream performance still degrades gradually as ρ increases past 0.25, and at ρ = 0.75 and ρ = 0.9 Model-Dowser's own metrics fall well below the ρ = 0.1 setting. Whether a refined or layer-adaptive scoring scheme can extend the stable window is an open question.
- Extending the sensitivity analysis beyond the first-order approximation. The importance definition relies on first-order Taylor approximation, and the paper cites numerical stability considerations deferred to an appendix; higher-order or better-conditioned estimators could sharpen the rankings.
- Applying the score to components beyond the language decoder. The experiments freeze all non-decoder components and update only the language decoder; whether the same functional importance probing transfers to projectors, vision encoders, or other architectures (the paper notes MLLMs may build on LLM-style designs) is untested here.
- Connecting to continual learning. The related-work section distinguishes downstream task adaptation from continual learning and points to a separate appendix discussion, leaving the behavior of importance-based masks under long sequences of tasks unresolved.
Target Audience
Researchers and engineers working on multimodal large language model adaptation, parameter-efficient and sparse fine-tuning, and continual learning. It will be most useful to readers who already understand transformer fine-tuning, gradient-based sensitivity measures, and evaluation protocols built around upstream/downstream trade-offs. Practitioners who must fine-tune models without access to pretraining data, or who are constrained by GPU memory, have the most to gain, while readers looking for an introductory treatment of catastrophic forgetting should expect to consult the cited background work first.
Authors’ abstract
Fine-tuning Multimodal Large Language Models (MLLMs) on task-specific data is an effective way to improve performance on downstream applications. However, such adaptation often leads to a degradation in generalization on pretrained tasks, a phenomenon known as Catastrophic Forgetting. Existing methods that aim to mitigate this issue either become ineffective when fine-tuning deeper layers of the language decoder or scale poorly with increasing model size. To address these limitations, we propose Model-Dowser, a novel sparse fine-tuning approach for MLLMs. Model-Dowser measures a principled importance score for each model parameter with respect to pretrained generalization (prior to downstream adaptation) by jointly considering weight magnitudes, input activations, and output sensitivities. During fine-tuning, Model-Dowser selectively preserves high-importance parameters and updates the remaining. Comprehensive experiments on two representative MLLMs, LLaVA and NVILA, demonstrate that Model-Dowser effectively mitigates catastrophic forgetting and consistently outperforms prior methods, while remaining resource-efficient and scalable to multi-billion-parameter models.