Skip to content
AI.info

Research

Weight Weaving: Parameter Pooling for Data-Free Model Merging

Weight Weaving: Parameter Pooling for Data-Free Model Merging Authors: Levy Chaves, Eduardo Valle, Sandra Avila arXiv: 2510.13921v1 [cs.LG], 15 Oct 2025 | Category: Machine Learning | License: CC BY 4

Weight Weaving: Parameter Pooling for Data-Free Model Merging
arXiv
2510.13921
Published
2025-10-15
Authors
Levy Chaves, Eduardo Valle, Sandra Avila

AI summary

Weight Weaving: Parameter Pooling for Data-Free Model Merging

Authors: Levy Chaves, Eduardo Valle, Sandra Avila arXiv: 2510.13921v1 [cs.LG], 15 Oct 2025 | Category: Machine Learning | License: CC BY 4.0 Code: https://github.com/VirtualSpaceman/weight_weaving

Overview

Research area: Model merging / transfer learning for deep neural networks, evaluated on computer vision tasks with CLIP vision transformers.

Technical level: Intermediate. The reader needs familiarity with fine-tuning, weight space arithmetic (task vectors), and hyper-parameter scaling, but the method itself is conceptually simple.

Scope: The paper introduces Weight Weaving, a plug-and-play, data-free technique that pools model parameters across a search space of scaling factors (lambda) instead of tuning a single scaling factor on privileged evaluation data, and validates it on three ViT variants across multi-task learning, continual learning, and domain generalization.

What This Paper Is About

Most model merging methods combine specialized models by scaling their weight differences with a hyper-parameter lambda. Choosing that scaling factor normally requires validation or evaluation data, and in practice researchers often tune it directly on the test set, which is not feasible in real deployments. The paper's goal is to remove that data dependency: instead of picking one lambda, Weight Weaving merges over many lambda values at once and pools the resulting parameters into a single model, without any data, training, or privileged information.

Key Contributions

  1. An efficient weight pooling method that collects parameter-level information across arbitrary, user-defined scaling-factor search spaces. The authors state that, as far as they know, Weight Weaving is the first pooling framework that bypasses privileged data requirements, improving the practicality of model mergers in data-free settings.

  2. Orthogonality to existing methods. Weight Weaving is a plug-and-play wrapper: it works with any merging function that depends on lambda, and can even be extended to search other hyper-parameters for methods that do not use lambda at all.

  3. Extensive evaluation across three experimental setups. Vision multi-task learning (8 tasks), vision continual learning (CIFAR100, ImageNet-R, CUB200, StanfordCars), and vision domain generalization, using ViT-B-32, ViT-B-16, and ViT-L-14. The method consistently beats the standard data-free baseline of lambda = 1, with average gains up to 15.9 percentage points.

  4. A discussion of modularity and open problems, including a study of three pooling functions (average, random uniform selection, MagMax) and an analysis linking the diversity of optimal scaling factors to when the method helps most.

Main Findings

  • Consistent data-free gains over state-of-the-art mergers. With arithmetic mean pooling, Weight Weaving improved average accuracy in the data-free setting for Breadcrumbs (52.17 to 68.11, +15.94), MagMax (60.14 to 69.77, +9.63), TIES (68.39 to 71.21, +2.82), PCB (71.41 to 72.10, +0.69), TSV (73.11 to 74.01, +0.90), and ISO-C (72.38 to 73.78, +1.40). Reported gains range from 0.90 to 15.94 percentage points.

  • Largest benefits occur in continual learning and domain generalization. For Breadcrumbs, continual learning jumped from 24.34 / 24.67 / 37.26 (ViT-B-32 / ViT-B-16 / ViT-L-14) to 66.00 / 70.42 / 80.70, and domain generalization from 42.59 / 51.72 / 64.26 to 52.47 / 60.06 / 67.87.

  • Multi-task learning gains are less consistent. A few per-model results dipped slightly: PCB on ViT-B-32 went from 75.94 to 75.55 and ISO-C on ViT-B-32 from 82.69 to 78.33, though the average across all setups still improved.

  • Two methods were excluded before the main comparison. Task Arithmetic and DARE performed poorly even with privileged lambda values, falling more than 10 percentage points behind Breadcrumbs, and both underperformed in continual learning, which was not part of their original evaluation protocol. With privileged data, the strongest average accuracy in Table 1 was TSV at 73.11.

  • Pooling function choice matters. Average and random uniform selection produced nearly identical results (for example, TSV + Ours: 74.01 vs 73.64), while using MagMax as the pooling function was far worse (Breadcrumbs + Ours dropped to 51.93, TIES + Ours to 54.56, PCB + Ours to 50.36), because merged parameters then converge toward the highest lambda values.

  • The method helps most when optimal scaling factors are spread out. For MagMax on ViT-B-32, the best lambda values are distributed across the search space and Weight Weaving improves every setup. For ISO-C in multi-task learning, optimal values concentrate around lambda = 1, and pooling offers little benefit.

  • Sequential fine-tuning changes the geometry of task vectors. In continual learning, initializing task t from the weights of task t-1 introduces correlation between consecutive tasks, so the task vectors are not nearly orthogonal, unlike the near-orthogonality reported for multi-task learning.

  • Data-free baselines are weak. For methods whose original work did not consider a data-free setting, the authors set lambda = 1, which is the comparison point throughout the data-free experiments.

Methodology in Plain English

The method has three stages, described in Algorithm 1:

  1. Compute delta weights. For each of T fine-tuned models, subtract the shared pre-trained weights to get the difference vectors.

  2. Build a pool of candidate merged weights. Take a merging function supplied by the user (Task Arithmetic, TIES, PCB, and so on) and a user-chosen range of lambda values. Apply the merging function once for each lambda in that range, producing one merged weight set per lambda. These are combined with the original delta weights into a single collection, A*, of size N x P, where N is the number of fine-tuned models plus the size of the search space, and P is the number of parameters in the pre-trained model.

  3. Pool and combine. Apply a user-defined pooling function across that collection to produce one set of pooled weights, then add them to the pre-trained model to get the final merged model. Pooling options include the element-wise arithmetic mean, independent random uniform selection per parameter position, or an existing merging method such as MagMax (which picks the value with maximum absolute magnitude at each position).

In practice, the search space is usually an equally spaced range from 0.1 to 1.0 in steps of 0.1. For multi-task learning the authors widened it for some methods following the original papers: [0.1, 1.5] for TIES, [0.1, 2.5] for PCB, [0.5, 2.0] for TSV, and [0.1, 2.0] for ISO-C. Task Arithmetic, Breadcrumbs, and MagMax use the standard 0.1–1.0 range everywhere.

Experimental setup: three CLIP vision-encoder variants (ViT-B-32, ViT-B-16, ViT-L-14); eight image classification tasks for multi-task learning and domain generalization (SUN397, Cars, RESISC45, EuroSAT, SVHN, GTSRB, MNIST, DTD); continual learning on CIFAR100 and ImageNet-R with N in {5, 10, 20, 50} class splits and CUB200 and StanfordCars with N in {5, 10, 20}; domain generalization uses a leave-one-out scheme, holding out one checkpoint, merging the rest, and testing on the held-out dataset. Fine-tuning used batch size 128, learning rate 1e-5 with cosine annealing, AdamW with weight decay 0.1, and a frozen text encoder. Experiments ran on a single NVIDIA-A100, and the setup covers models up to 422M parameters.

Why This Matters

Impact on research. The paper attacks a methodological weakness that the merging literature has largely worked around: tuning lambda on the evaluation set. By framing the problem as pooling over a search space rather than selecting a single value, it offers a wrapper that any lambda-dependent merger can adopt, and it reframes lambda selection as an open design space for pooling functions. It also documents that optimal scaling factors are dataset- and method-dependent and do not transfer across setups.

Real-world applications (bullets):

  • Deploying merged models without labels in settings where no validation data exists, such as private or regulated environments.
  • Combining domain-specialized vision models (for example, satellite imagery, industrial inspection, or medical imaging) from public repositories into one model without re-training.
  • Continual/on-device learning, where a model must absorb new classes without revisiting earlier data, a setting where the method delivered its largest gains.
  • Cheap model reuse at scale, since merging avoids the cost of retraining each specialized network and needs no unlabeled test data at merge time.

Industry relevance. Merging is attractive precisely because it is cheap and data-efficient, but the lambda tuning step has made reported results hard to reproduce in production. Removing that step makes merging pipelines more deployable for teams that cannot hold out or access evaluation data, and the note that complexity is bounded by the search space size (with parallel computation available for billion-parameter models) indicates the cost is a compute trade-off rather than a data requirement.

Future Directions

  • Filtering the search space. The authors suggest discarding suboptimal lambda values before pooling, for example by extending Li et al.'s analysis of task alignment and weight correlations beyond the simplified two-model, binary-task case, or by studying layer-wise weight norms and model embeddings.
  • Pooling functions tailored to specific setups. Since average and random performed similarly and MagMax performed poorly, the paper argues that application-specific pooling methods could yield substantially better weight combination.
  • Merging techniques designed for continual learning. The observed correlation between consecutive task weights, versus near-orthogonality in multi-task learning, suggests redundancy that may worsen as the number of continual tasks grows.
  • Generalizing beyond lambda. The framework can search categorical variables, probability distributions, or arbitrary functions, and the authors invite the community to explore efficient alternatives to tuning lambda rather than relying on evaluation data.

Target Audience

Researchers and practitioners working on model merging, multi-task learning, continual learning, and domain generalization, especially those who need deployable fusion of fine-tuned models without access to validation or evaluation data. It is also relevant to engineers building model-reuse pipelines from public checkpoints, and to readers interested in hyper-parameter selection methods that avoid privileged information.

Authors’ abstract

Model merging provides a cost-effective and data-efficient combination of specialized deep neural networks through parameter integration. This technique leverages expert models across downstream tasks without requiring retraining. Most model merging approaches critically depend on scaling hyper-parameters $λ$, which weight each model's contribution globally or individually. Principled approaches for setting scaling factors without accessing any data (data-free) are scarce, often leading researchers to tune $λ$ using privileged data from the evaluation set, which is obviously unfeasible in practice. To address this limitation, we introduce Weight Weaving, a plug-and-play technique that pools model weights across $λ$ values search space using user-defined pooling functions, such as averaging, random selection, or even existing model merging methods. Our method demonstrates high modularity, imposing minimal constraints on the search space. It operates orthogonally to existing model merging methods and eliminates evaluation data requirements. We validate Weight Weaving across three ViT variants in three experimental setups: vision multi-task learning, vision continual learning, and domain generalization. Our method consistently improves the performance of several model merging methods, achieving average accuracy gains of up to 15.9 percentage points in a data-free setting.

Read the original paper