Research
Not All Ranks Are Equal: Budget-Aware LoRA Merging Across Tasks
Overview Research area: Machine learning — parameter-efficient fine-tuning, specifically merging multiple Low-Rank Adaptation (LoRA) adapters into a single multi-task adapter. Technical level: Advance

- arXiv
- 2609.22237
- Published
- 2026-09-03
- Authors
- Avinash Amballa, Yashas Malur Saidutta, Wenbo Li, Lazar Valkov, Srinivas Chappidi
AI summary
Overview
Research area: Machine learning — parameter-efficient fine-tuning, specifically merging multiple Low-Rank Adaptation (LoRA) adapters into a single multi-task adapter.
Technical level: Advanced. The method relies on singular value decompositions of adapter matrices, weighted seminorms, a separable upper bound on a preservation loss, and a budget-constrained selection problem.
One-sentence scope: The paper argues that the performance gap between merged and per-task LoRAs is largely caused by allocating the same rank budget to every layer and every task, and proposes a data-free scoring rule ("Net Utility") that reallocates a global rank budget non-uniformly across layers, modules, and tasks.
What This Paper Is About
Serving many fine-tuned tasks usually means keeping one LoRA adapter per task and swapping the active one at inference, which is costly on-device. Merging collapses those adapters into a single low-rank update, but nearly all existing merging methods give every layer the same rank, and some (for example TSV-Merging, which uses d/n) explicitly give every task an equal share. The paper's goal is to show that this uniform-budget assumption is itself a major source of lost accuracy, and to provide a practical way to spend a fixed total rank budget where it matters most, without needing any task data.
Key Contributions
-
Identifies uniform rank allocation, rather than the merging operation itself, as a major source of the gap between merged and per-task LoRAs. As motivation, relaxing uniform allocation on ViT-B/32 over seven vision tasks with Task Arithmetic merging improves accuracy by +3.4%.
-
Proposes Net Utility, a data-free score for each rank-one singular direction of each task adapter that trades the direction's usefulness to its own task against its interference with other tasks' directions, recasting merging as a globally budget-constrained selection problem.
-
Shows that the selected directions need not form a prefix of the singular spectrum, and that some directions have negative net utility, so the full budget need not be spent.
-
Demonstrates at matched budgets, on 7 vision and 6 language tasks, that Net Utility allocation improves multiple prior merging methods across various merging spaces by an average of +2.1% absolute, with some method-space combinations reaching as much as +3.8% absolute.
Main Findings
-
Uniform allocation is the bottleneck, not merging itself: the motivating experiment shows +3.4% average accuracy across seven vision datasets with Task Arithmetic merging when the uniform-budget assumption is relaxed.
-
Vision results (ViT-B/32, 7 tasks): Net Utility beats uniform allocation across all three spaces and all four merging methods reported, with a maximum gain of +3.8 over DARE in the full space at budget 64. DARE improves by +2.3 to +3.8 and TA by +2.3 to +3.4, while TSV improves less, by +0.3 to +0.8, possibly because TSV whitens and re-ranks the spectrum, making uniform allocation a strong baseline there.
-
Language results (Qwen3-4B, 6 NLI tasks): improvements hold for most methods across all spaces, with a maximum gain of +3.5 on TSV at budget 16. Gains over TSV range from +3.1 to +3.5, over TIES from +1.7 to +3.4, and over TA from +2.2 to +2.3. The exception is DARE in core and KnOTS space, where results are within 0.4 of uniform and slightly below it at budgets 64 and 32.
-
Gains grow as the budget shrinks: on language tasks, TSV improves from +3.1 at R=64 to +3.5 at R=16, and TIES in full space from +1.7 to +3.4, consistent with the idea that direction choice matters more under a tight budget.
-
Average across everything: vision tasks see +2.1% and language tasks +2.2% on average across merging methods and spaces.
-
The budget need not be spent: at a budget of 112 on vision Task Arithmetic, the method retains only 61.5 directions per layer on average, leaving 45.0% of the budget unspent, and still raises accuracy by +3.19 over uniform. At budget 64 a similar number of ranks is used, and at budgets 32 and 16 the entire budget is spent.
-
The optimal set is not a prefix of the spectrum: prefix truncation is suboptimal. On vision, 88% of the chosen sets are not prefixes at every budget; on language, 64.5% of sets are non-prefix at R=64. On vision at R=64, the strongest direction of a task is retained only 12.8% of the time while the weakest is retained 75.8% of the time. On language, 59% of directions are retained at l=15, where a prefix rule retains none.
-
Layers and modules are not equally important: in the ViT-B/32 allocation at budget 64, later layers require more budget than early layers, and value-projection and output-projection modules occupy more budget in later layers.
-
Net Utility correlates with task difficulty: on a controlled setup using Hendrycks MATH split into seven topics (algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, precalculus) with one rank-32 LoRA per topic on Qwen3, mean net utility is inversely correlated with relative accuracy with a correlation of −0.78. Since higher relative accuracy means an easier topic, harder topics have higher net utility.
-
Transfer across merge operators: the metric derived for Task Arithmetic also improves TSV, DARE, TIES, and Iso-C, even though those operators are not linear in the task updates, showing the scores transfer across merge methods and spaces.
-
Isotropization variant: with Iso-C at R=16 in full space, allocation improves over uniform by +1.3 and +1.2 on vision and +1.0 and +0.5 on language.
Methodology in Plain English
The starting point is a preservation objective: a merged adapter should reproduce what each individual task adapter would have produced on that task's own inputs. Because the algorithm never sees task data, the researchers linearize the model around a shared reference and rewrite the objective in terms of the layer input statistics, expressed as a weighted norm involving a matrix G that captures second-order input statistics. That matrix is unavailable without data, so they replace it with a scaled input-side Gram proxy computed from the adapter's own SVD, exactly a scaling of the adapter's row space — the same idea used by ACE-Merging. This makes the whole criterion a function of adapter weights alone, hence data-free.
The preservation loss couples all the selected components because of a signed cross-term. Using Young's inequality and Cauchy-Schwarz, the authors derive a separable upper bound that removes that cross-term and splits interference by source task, introducing a coefficient lambda that controls how heavily interference is penalized. Under the Gram geometry, the objective becomes a constant M minus the sum of per-component utilities, so maximization reduces to picking the components with the highest utility.
Each utility has two parts: a source-task benefit equal to the component's fourth power of its singular value divided by the sum of fourth powers of all that task's singular values, and a cross-task interference term equal to its squared singular value times a weighted sum of squared overlaps between its right singular vector and other tasks' right singular vectors. Interference vanishes for directions orthogonal to every other task's row space. Lambda is set in a data-free way as the ratio of the median source benefit to the median interference term.
Because the utilities are comparable across the whole model, all components from all tasks and all adapted layers are pooled and the top positive scores are selected under a global budget of L times R, meaning an average of R directions per layer but free to move between layers. Capacity is left unused if fewer than that many scores are positive. The selected components are then passed through the chosen merge operator. The authors note the derivation applies to the additive merge; applying selected components through a nonlinear merge operator is an empirical extension without the same surrogate guarantee.
A hyperparameter alpha in {0,1} controls whether the proxy weights singular directions by squared singular values or weights them equally. The authors measure heterogeneity with the variance of logarithms of adapter Frobenius norms, finding h=1.83 across the seven vision adapters and h=0.63 across the six NLI adapters; since the vision tasks are more heterogeneous, they set alpha=0 to stop large-magnitude tasks dominating.
Experimental setup: vision uses CLIP ViT-B/32 with LoRA checkpoints from KnOTS, over DTD, EuroSAT, GTSRB, MNIST, RESISC, SUN397, SVHN; language uses Qwen 3-4B fine-tuned on SNLI, MNLI, SICK, QNLI, RTE, SCITAIL. All LoRAs have rank 16 on key, query, value, and output projections across all attention layers. Baselines are Task Arithmetic, TIES, DARE, TSV, and Iso-C, applied in three merge spaces: full weight space, core space, and KnOTS space. Budgets are R in {64, 32, 16}, where full capacity is T times r, i.e. 112 for vision and 96 for language, so all evaluated budgets are below full. Accuracy is reported as normalized accuracy — the merged LoRA's accuracy on a task divided by the original LoRA's accuracy on that task. Experiments run on NVIDIA H100 GPUs. For TIES and DARE, budgeting is local (equal budget per layer) because those methods apply sample-specific operations.
Why This Matters
Impact on research: The paper reframes a widely used design choice — giving every layer and every task the same rank — as an unexamined assumption rather than a safe default. It also shows that the best subset of singular directions is generally not a prefix, which challenges the prefix truncation used by essentially all existing SVD-based merging methods. Since the scoring rule is closed-form in adapter weights and needs no calibration data or optimization loop, it can be layered on top of existing merge operators rather than replacing them.
Real-world applications:
- On-device or edge deployment of many task adapters, where hot-swapping task-specific weights is expensive and a single merged adapter is preferable.
- Serving systems where decoding latency grows linearly in the ranks present in a batch and mixed-rank adapters force low-rank requests to pay the cost of the highest rank in the batch, so a smaller, better-chosen merged rank directly reduces cost.
- Deployments that fix a rank at serving time independently of what training would have chosen, where budget allocation determines quality within that fixed constraint.
- Multi-task personalization or assistant scenarios with many specialized behaviors, where collapsed adapters avoid routing infrastructure.
Industry relevance: The work comes from Samsung Research America and is explicitly motivated by on-device and edge applications, where adapter count and rank both translate into memory and latency. The claimed average gains of +2.1% on vision and +2.2% on language at matched rank budgets are directly usable by teams already running LoRA merging pipelines.
Future Directions
- Deriving Net Utility specifically for nonlinear merge operators such as TSV, DARE, and TIES, rather than transferring the additive-merge score to them empirically, to potentially improve those combinations further.
- Tuning the interference coefficient lambda when a validation set is available, which the authors state would improve performance but which they leave at a data-free default.
- Investigating the observed link between net utility and task difficulty more systematically — the paper establishes a −0.78 correlation on a single controlled MATH-topic setup, which invites broader validation.
- Extending the allocation framework to settings the paper does not study, such as merge-friendly training, on-device continual merging under a budget on retained adapters, and sequential merging for continual learning, all of which it lists as adjacent but distinct.
Target Audience
Researchers and engineers working on parameter-efficient fine-tuning, LoRA merging, model merging, and multi-task serving. It is most useful to readers already comfortable with SVD-based analysis of weight updates and with the vocabulary of task vectors, merge spaces, and rank budgets. Practitioners deploying many adapters on hardware with fixed rank or memory constraints will find the practical framing and the reported baselines most directly actionable.
Authors’ abstract
Merging low-rank adapters (LoRAs) promises to eliminate the overhead of swapping task-specific weights at inference time. However, existing merging methods assume every layer needs the same rank budget. Further, some methods assume that rank budget needs to be split equally among the tasks too. We show this uniform-budget assumption is a major source of the performance gap between merged and per-task LoRAs. However, rank selection is an NP hard problem. To this end, we introduce Net Utility, a data free metric that first decomposes every task LoRA by its Singular Value Decomposition (SVD) and scores each of those singular directions by its task utility and its interference with other tasks directions. Next, we globally pool these scores to select singular directions with the highest values with a constraint on the total number of directions selected. The proposed Net Utility metric is applied on top of five different merging methods across three different merging spaces. The merging is done over two sets of tasks, vision and language tasks. Net utility based rank allocation outperforms its counterparts without that allocation. On average, over vision tasks it achieves +2.1% improvement in performance, and +2.2% improvement over the language tasks.