Research
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation Models Overview Research area: Machine learning / computer vision — transfer learning and adapt

- arXiv
- 2512.01405
- Published
- 2025-12-01
- Authors
- Benjamin Ramtoula, Pierre-Yves Lajoie, Paul Newman, Daniele De Martini
AI summary
Fantastic Features and Where to Find Them: A Probing Method to combine Features from Multiple Foundation ModelsOverview
Research area: Machine learning / computer vision — transfer learning and adaptation of foundation models (FMs), specifically probing-based adapters for combining representations from multiple pre-trained vision backbones.
Technical level: Intermediate. The paper assumes familiarity with Vision Transformers (ViTs), linear probing, parameter-efficient fine-tuning (PEFT), and knowledge distillation, but its core idea is describable without heavy mathematics.
Scope: The paper introduces ComBo (Combined backBones), a lightweight probing adapter that fuses frozen features from multiple foundation models and layers, evaluated on the 19 tasks of the VTAB-1k benchmark.
What This Paper Is About
Foundation models trained with different objectives and data — CLIP, MAE, SAM, DINOv2 — learn different representations, so the best model (and even the best layer within a model) varies by downstream task. Existing ways to exploit several models at once either tune one model at a time, require backpropagating through large backbones, or distil several teachers into a single student at considerable cost. This paper asks how to combine the complementary features of many frozen models directly, cheaply, and without per-dataset tuning.
Key Contributions
- An analysis of representation differences across foundation models. The authors linearly probe activations from different layers of ViTs trained with CLIP, DINOv2, ImageNet-21K, MAE, and SAM (plus a randomly initialised ViT) on the Clevr/distance and SVHN VTAB-1k tasks, and visualise maximally activating neurons using the technique from Ghiasi et al. They show that both the best model and the best layer change with the task.
- A new probing adapter. ComBo compresses activations from the layers of one or more frozen FMs into compact token-wise representations and processes them with a small transformer, without arbitrary average pooling and without dataset-specific hyperparameter tuning.
- Practical multi-model integration without backpropagation. ComBo probes several large FMs together while only running forward passes through them, which the authors frame as making large model combinations accessible to practitioners with limited compute.
- A method for measuring a backbone's task-relevance. Using ComBo's joint multi-backbone probing plus an ℓ2-norm regularisation on the projection weights, the authors derive a per-model importance score that supports model comparison and selective retraining on a smaller subset of backbones.
Main Findings
- Single-backbone probing is competitive but not dominant. On VTAB-1k with a ViT-B/16 trained on ImageNet-21K, ComBo reaches a 74.6 global average, ahead of full fine-tuning (74.2), Head2Toe (70.9), SMP (73.6), and linear probing (61.0), but behind PEFT methods: LoRA (75.6), VPT-Deep (76.4), and Adapter+ (77.6, the best overall).
- ComBo's advantage concentrates on structured tasks. In Table 1, ComBo reaches a 59.5 structured average, versus 47.7 for Head2Toe, 55.4 for SMP, and 63.3 for Adapter+. The authors attribute this to preserving complete original feature maps instead of average pooling.
- Multi-model probing beats distillation-based merging. Probing all four FMs with ComBo gives a 78.2 global average, and selecting the top two by importance score gives 78.6, both above the distilled RADIOv2.5 model fine-tuned with Adapter+ (77.5). Individual-model ComBo results were DFN CLIP 75.6, DINOv2 77.9, SAM 63.8, and SigLIP 76.2.
- ComBo's edge on structured tasks holds against distillation. For structured tasks, the all-four and top-two variants score 64.8 and 65.3 versus 62.1 for RADIOv2.5 with Adapter+.
- Not every backbone helps. SAM consistently underperformed the other models, and the top-two selection improved on using all four models (78.6 versus 78.2) while processing fewer backbones.
- Combining tuned backbones works, at lower cost. With ViT-B models individually tuned with Adapter+, DINOv2 alone achieves the highest global average (79.9). Multi-Adapter+ leads on 8 of 19 tasks and on natural tasks (86.3), ComBo wins 5 of 19 tasks and on structured tasks (67.3 versus 61.6), and ComBo's global average (79.3) edges Multi-Adapter+ (78.7) while requiring only forward passes through pre-tuned models. DINOv2 leads on only 3 of 19 tasks.
- Layer choice matters only for untuned models. Probing both all layers and only the last layer of the original model gives averages of 76.5 and 63.6 respectively on the val set, but for Adapter+-tuned models the advantage of probing all layers disappears (76.9 for last layer versus 76.5 for all layers), suggesting tuning concentrates task-relevant information in the final layers.
- Importance scoring beats naive selection. Against an exhaustive-search upper bound of 78.6 average on the val set, top-2 selection scores 76.7, ahead of using all four models (76.2), DINOv2 alone (76.0), and random pairs (75.4). Selecting the best per-task top-n approaches the upper bound at 78.1.
- Both the normalisation and the transformer matter. Removing feature map normalisation drops the average from 73.8 to 71.5; replacing the transformer with a linear head drops it to 65.1, and with SMP's MLP to 67.5.
- ComBo is cheaper to train than the alternatives. On one backbone, ComBo updates 2.79% of parameters and uses 35.78% of the relative training FLOPs with 0.929 GB peak GPU memory, versus 65.58% FLOPs and 2.993 GB for Adapter+. On four backbones, ComBo uses 34.41% relative training FLOPs and 3.212 GB, versus 65.67% and 13.246 GB for Multi-Adapter+.
Methodology in Plain English
The authors start from an existing idea: instead of changing a pre-trained model's weights, extract its internal activations and train a small classifier on top ("probing"). Prior probing methods stack activations from many layers and locations and feed them to a linear layer, which explodes in size — the paper notes that for a ViT-B with 12 blocks processing 197 tokens of 768 dimensions, a 100-class linear classifier on stacked features would need more than 181M parameters. Those methods avoid the blow-up by average-pooling activations over regions whose size must be tuned per dataset.
ComBo instead handles the dimensionality in two stages. First, feature maps from every selected layer of every model are reshaped to their 2D spatial layout and bilinearly interpolated so all models produce the same number of tokens, then normalised. For each spatial position, the normalised vectors from all layers and models are concatenated and passed through a single learned affine projection shared across positions, compressing them to a much smaller dimension while keeping the token layout. Second, these compressed tokens plus a learnable class token go into a small transformer encoder, whose output class token feeds a linear classifier. The projection, class token, and transformer are trained jointly with cross-entropy loss.
Because the projection weights are shared and structured, the authors can ask how much each model contributes: they add a regularisation term equal to the sum of each backbone's ℓ2 norm of its associated projection columns, and use the resulting per-model scores to rank backbones. The workflow is to train once with all models under regularisation to get the scores, then retrain without regularisation using only the top-scoring models for that task.
Training uses constant settings across all 19 tasks — batch size 64, AdamW with learning rate 0.001 and weight decay 0.0001, 100 epochs with a 10-epoch linear warmup and cosine schedule, images resized to 224 × 224 with no augmentations, and λ = 0.01 for the relevance regularisation. Results are the mean accuracy over 3 seeds. The transformer uses a standard timm implementation with depth 6, embedding dimension 128, and 2 heads (1.7M parameters).
Why This Matters
Impact on research. The paper shows that a well-designed probing adapter can match or exceed a distilled multi-teacher model and approach joint fine-tuning of several models, without ever backpropagating through a backbone. It also provides a quantitative handle on which frozen models actually help a given task — a question that becomes more pressing as the number of released foundation models grows.
Real-world applications (drawn from the VTAB-1k task domains the paper evaluates on):
- Medical and satellite imagery, where specialised tasks are common and labelled data is scarce.
- Fine-grained natural-image classification, such as flowers, pets, and scene categories.
- Scene-structure tasks like object counting and depth-related prediction, where ComBo showed its largest gains.
- Deployments where a practitioner must choose among several available pre-trained checkpoints under a fixed compute budget.
Industry relevance. Because ComBo trains no backbone gradients and uses reduced training FLOPs and memory relative to Adapter+ and Multi-Adapter+ (for example, 3.212 GB versus 13.246 GB when adapting four backbones), the approach is positioned as practical for teams without the resources to distil or jointly fine-tune several large models. Its fixed hyperparameters across datasets also reduce the engineering cost of applying it to new tasks.
Future Directions
- Extending ComBo beyond ViTs. The authors state the core recipe — extract multi-layer features, align them spatially, compress with a learned projection, process with a transformer — could likely be extended to other architectures such as CNNs, given appropriate spatial alignment and interpolation strategies.
- Handling models of varying quality. The paper notes a performance gap between combination methods and strong individual models like DINOv2, and points to that gap as evidence that ComBo's management of unevenly capable models could be improved, possibly through the task-relevance mechanism.
- Choosing how many backbones to use. Top-2 selection performed best with a fixed number of models, but the optimal number varied per task; discovering it automatically from the importance scores is left as an open question.
- What tuned models keep. The finding that probing only the final layer suffices for Adapter+-tuned models, while all layers are needed for untuned ones, raises the question of how tuning redistributes task-relevant information, and whether that can be exploited further for efficiency.
Target Audience
Researchers and practitioners working on transfer learning, foundation-model adaptation, and multi-model representation fusion, particularly those who need to combine several large pre-trained vision backbones under constrained compute. It is also relevant to readers interested in benchmarking methodology, since it reports full VTAB-1k comparisons against PEFT, probing, and distillation baselines.
Authors’ abstract
Foundation models (FMs) trained with different objectives and data learn diverse representations, making some more effective than others for specific downstream tasks. Existing adaptation strategies, such as parameter-efficient fine-tuning, focus on individual models and do not exploit the complementary strengths across models. Probing methods offer a promising alternative by extracting information from frozen models, but current techniques do not scale well with large feature sets and often rely on dataset-specific hyperparameter tuning. We propose Combined backBones (ComBo), a simple and scalable probing-based adapter that effectively integrates features from multiple models and layers. ComBo compresses activations from layers of one or more FMs into compact token-wise representations and processes them with a lightweight transformer for task-specific prediction. Crucially, ComBo does not require dataset-specific tuning or backpropagation through the backbone models. However, not all models are equally relevant for all tasks. To address this, we introduce a mechanism that leverages ComBo's joint multi-backbone probing to efficiently evaluate each backbone's task-relevance, enabling both practical model comparison and improved performance through selective adaptation. On the 19 tasks of the VTAB-1k benchmark, ComBo outperforms previous probing methods, matches or surpasses more expensive alternatives, such as distillation-based model merging, and enables efficient probing of tuned models. Our results demonstrate that ComBo offers a practical and general-purpose framework for combining diverse representations from multiple FMs.