Research
Making Models Unmergeable via Scaling-Sensitive Loss Landscape
Overview Research area: Machine learning model security and governance, specifically the protection of released fine-tuning updates (LoRA adapters and full checkpoints) against unauthorized downstream

- arXiv
- 2601.21898
- Published
- 2026-01-29
- Authors
- Minwoo Jang, Hoyoung Kim, Jabin Koo, Jungseul Ok
AI summary
Overview
Research area: Machine learning model security and governance, specifically the protection of released fine-tuning updates (LoRA adapters and full checkpoints) against unauthorized downstream model merging.
Technical level: Intermediate. The paper assumes familiarity with LoRA, fine-tuning, and standard merging operators (Task Arithmetic, TIES, DARE), but the core idea is stated clearly enough for readers with a general deep-learning background.
Scope: The paper proposes and evaluates Trap² (Training-time Protection via Task-Robust Adversarial Perturbation), an architecture-agnostic training-time objective that makes released model updates useful standalone but brittle under the weight re-scaling that merging pipelines introduce.
What This Paper Is About
Public model hubs make it easy to download fine-tuned updates and recombine them into new models, but this also means released weights can be recomposed into unauthorized mixtures that bypass safety alignment or licensing terms. This creates a "governance gap": providers lose control over their updates once released, and existing defenses against recombination are mostly post-hoc and specific to Transformer architectures, which makes them unreliable across different models and release formats. The paper's goal is a protection mechanism embedded in the fine-tuned update itself, so that the update works normally when used alone but degrades when it is merged with other updates.
Key Contributions
-
Problem setup: The authors formalize a post-release protection setting for fine-tuned releases across multiple architectures and release formats (adapter-only updates and full checkpoints), along with a unified evaluation protocol that quantifies standalone utility and degradation under merging.
-
Protection method (Trap²): They introduce a training-time procedure that keeps a fine-tuned update effective at the nominal scale (s = 1) while making it brittle under off-nominal re-scaling (s ≠ 1), which is described as capturing the re-weighting effects commonly introduced by practical merging pipelines.
-
Empirical evaluation: Trap² is evaluated across diverse merging operators, release formats, and architectures, including CLIP ViT-B/32, ViT-L/14, and ConvNeXt backbones, plus full fine-tuning, mathematical reasoning tasks, and other LoRA variants (QLoRA, DoRA).
-
Theoretical analysis: The paper complements its experiments with analysis establishing (i) convergence under stochastic optimization and (ii) degradation under down-scaling and model merging.
Main Findings
-
Standalone utility is preserved: On 8 vision benchmarks with CLIP ViT-B/32 adapters, Trap² reaches an average standalone accuracy of 88.132% versus 88.119% for unprotected fine-tuning. On ViT-L/14 it reaches 93.711% versus 92.403% fine-tuned, and on ConvNeXt 90.660% versus 90.605%.
-
Post-hoc baselines fail on adapter-only releases: Merge-Lock collapses to near-chance standalone accuracy under both adapter-only variants (6.254% and 4.157% average on ViT-B/32; 4.425% and 4.525% on ViT-L/14). PaRaMS retains utility (87.638% and 84.392% on ViT-B/32) but shows only small deviations from unprotected merging.
-
Merging degrades sharply with Trap²: Under 8-way LoRA merging on ViT-B/32, unprotected TA merging yields 48.273% while Trap² yields 23.121% — below the 42.483% zero-shot reference. On ViT-L/14, TA merging drops from 62.914% (unprotected) to 33.619% (Trap²); on ConvNeXt, from 49.203% to 14.719%.
-
The merger is given an advantage: Merging coefficients are tuned via validation search over s ∈ {0.1, 0.2, ..., 10.0} with step size 0.1 (including s = 1.0), which the authors describe as making the results optimistic for the merger.
-
Pairwise merging still collapses: At scale s = 0.8 on CLIP ViT-B/32, the protected adapter consistently self-collapses after merging and often induces collateral degradation on the unprotected partner task.
-
Full fine-tuning transfers: Under full fine-tuning on ViT-B/32, Trap² standalone accuracy is 87.372% versus 87.598% for standard fine-tuning. Under full-model merging, TA drops from 49.963% (unprotected) to 28.475% (Trap²), and CART from 61.031% to 42.116%.
-
Other LoRA variants and domains: Trap² works with QLoRA and DoRA on the GTSRB–Cars pair with ViT-B/32. On Llama-3.1-8B with rank-32 LoRA adapters on GSM8K and ASDiv (Exact-Match accuracy), standalone performance is comparable to unprotected fine-tuning while pairwise interpolation collapses.
-
Resistance to data-aware attacks: Under data-dependent merging on ViT-B/32, RegMean drops from 49.107% (unprotected) to 9.480% (Trap²) and CoM from 64.776% to 32.331%. Under data-driven recovery, ProDistill drops from 73.823% to 49.153% and post-merge SFT from 56.859% to 45.496%. The authors note CoM often assigns comparable or larger weight to the protected adapter rather than down-weighting it.
-
Uniform-averaging proxy: Evaluating a single adapter at s = 1/N, Trap² shows catastrophic degradation as soon as N ≥ 2, while unprotected adapters remain relatively stable.
-
Training cost: Per-step time increases by 1.72x on a single RTX A6000, with peak GPU memory nearly unchanged. End-to-end, Trap² is faster on 5 of 8 datasets to a shared target accuracy (speedup up to 6.097x on Aircraft) and achieves higher validation accuracy on 6 of 8 datasets under a shared wall-clock budget (largest gain +13.741 pp on Aircraft).
Methodology in Plain English
The authors take a different route from prior defenses. Instead of applying a transformation after training that hides how the weights can be combined (which requires access to the full model and relies on the specific wiring of Transformer attention blocks), they bake the protection into training.
The key trick is to treat "merging" as a simple form of re-scaling. When several adapters are averaged together, each individual adapter is effectively multiplied by a fraction (for example, 1/N for a uniform average of N adapters). So the authors define a training objective with two parts. The first part is the normal loss at scale s = 1, which keeps the model accurate when used as released. The second part deliberately makes the model bad at other scales, sampling scaling factors from a range that excludes a margin of width δ around 1. The two parts are combined with a trade-off weight λ and optimized with stochastic gradient descent, using one sampled scale per iteration as a cheap unbiased estimate. A weighting function w(s) = 1/s is used by default to compensate for small gradients at down-scaled updates.
Because the protection lives in the fine-tuned update rather than in base-weight symmetries, the same objective works for LoRA adapters, full checkpoints, and non-Transformer backbones like ConvNeXt.
Why This Matters
Impact on research: The paper reframes unmergeability as a training-time property rather than a post-hoc transformation, and it explicitly targets two gaps in prior work: adapter-only release formats (where base weights are unavailable) and non-Transformer architectures. It also provides convergence and degradation analysis alongside the empirical results.
Real-world applications:
- Model hubs and open repositories (the paper names GitHub and Hugging Face) that distribute fine-tuned updates and adapters, where providers want to preserve licensing and safety constraints after release.
- Safety-alignment preservation, preventing released weights from being blended into composite models that bypass alignment training.
- Provenance and IP enforcement in the model supply chain, where recomposition currently obscures origin and complicates terms-of-use enforcement.
- Pre-release quality checks: the paper suggests using update re-scaling as a lightweight proxy for merging behavior before an update is released.
Industry relevance: Organizations that publish adapters or checkpoints can adopt Trap² as part of their existing fine-tuning pipeline rather than as a separate, architecture-specific post-processing step. The reported training overhead (1.72x per step, near-unchanged peak memory, and faster time-to-target on 5 of 8 datasets) is framed as practical for real workflows.
Future Directions
- Determining how well the protection holds against merging operators that do more than re-scaling, since Trap² uses re-scaling as a "simple proxy" for the merging process.
- Extending the evaluation beyond the settings reported here — the mathematical reasoning results, QLoRA/DoRA results, and the complete merging tables are described as appearing in the appendix, and the main text illustrates only a subset.
- Characterizing the trade-off controlled by λ more precisely: how much standalone utility must be given up, if any, to achieve a given level of post-merge degradation.
- Testing whether an adversary with full knowledge of the Trap² objective and access to more data or compute can recover utility from the protected merge, beyond the data-dependent merging and data-driven recovery threats already studied.
Target Audience
This paper is most useful for researchers working on model merging, parameter-efficient fine-tuning, and machine learning security; engineers and platform teams operating model hubs or distributing adapters; and policy or governance researchers concerned with enforcing licensing and safety constraints on released model weights. Readers should be comfortable with LoRA, standard merging operators such as Task Arithmetic and TIES-Merging, and basic stochastic optimization.
Authors’ abstract
The rise of model hubs has made it easier to access reusable model components, making model merging a practical tool for combining capabilities. Yet, this modularity also creates a governance gap: downstream users can recompose released weights into unauthorized mixtures that bypass safety alignment or licensing terms. Because existing defenses are largely post-hoc and architecture-specific, they provide inconsistent protection across diverse architectures and release formats in practice. To close this gap, we propose Trap$^2$, an architecture-agnostic protection framework that encodes protection into updates during fine-tuning, regardless of whether they are released as adapters or full models. Instead of relying on architecture-dependent approaches, Trap$^2$ uses weight re-scaling as a simple proxy for the merging process. It keeps released weights effective in standalone use, but degrades them under re-scaling that often arises in merging, undermining unauthorized recomposition.