Research
The Triangle of Similarity: A Multi-Faceted Framework for Comparing Neural Network Representations
Overview Research area: Interpretability and representational similarity analysis of deep neural networks — specifically methods for comparing what different trained models have learned internally. Te

- arXiv
- 2601.17093
- Published
- 2026-01-23
- Authors
- Olha Sirikova, Alvin Chan
AI summary
Overview
Research area: Interpretability and representational similarity analysis of deep neural networks — specifically methods for comparing what different trained models have learned internally.
Technical level: Intermediate. Readers need familiarity with neural network representations, similarity metrics such as CKA, and pruning, but the framework itself is conceptually simple.
Scope: The paper proposes and empirically validates a three-part framework (static, functional, and sparsity similarity) for comparing vision model representations across CNNs, Vision Transformers, and Vision-Language Models.
What This Paper Is About
Methods for comparing neural network representations usually capture only one aspect of a model pair — either the geometry of internal features or the function computed in weight/output space — and treat models as static objects. This paper argues that a single view gives an incomplete picture, and proposes combining static representational similarity (CKA/Procrustes), functional similarity (Linear Mode Connectivity for same-architecture pairs, Predictive Similarity via Jensen-Shannon Divergence for different-architecture pairs), and sparsity similarity (how relationships survive progressive pruning) into one framework. The goal is a more holistic "fingerprint" of whether two models have converged on similar internal mechanisms, useful for model selection and scientific validation.
Key Contributions
- The Triangle of Similarity framework. A three-panel analysis for any pair of deep models with one or more hidden layers, integrating static representational view, functional view, and sparsity view into a single comparison protocol.
- An adaptation rule for the functional view. Linear Mode Connectivity is used when the two models share an architecture; when they differ, the framework substitutes Predictive Similarity measured as Jensen-Shannon Divergence between average softmax prediction distributions, where a score near 0 indicates similar predictions.
- Pruning repurposed as an analytical probe. Rather than using pruning for compression, the authors use global magnitude pruning as systematic "stress" to test whether model similarity is a robust shared core or an artifact of over-parameterization.
- A cross-view statistical validation. A quantitative correlation between the static view and the sparsity view across 21 model pairs on the ImageNetV2 testbed.
Main Findings
- Architectural family is the primary predictor of representational similarity. CKA and Procrustes heatmaps on out-of-distribution CIFAR-10 images show distinct high-similarity blocks: Transformer-based models (ViT, DeiT, DINO, CLIP, BLIP, LLaVA) form a highly coherent cluster attributed to the self-attention mechanism, while CNNs form a separate, less tight cluster. Cross-family similarity (e.g., RN50 vs. DINOv2) is consistently lower than within-family similarity (e.g., DeiT-TP16 vs. ViT-TP16). This clustering remains robust on out-of-distribution data.
- Task accuracy is more sensitive to pruning than representational structure. CKA self-similarity (pruned model vs. its original) often stays above 0.8 even at 40% sparsity, while Top-1 accuracy drops much more steeply — most pronounced for the Vision Transformers (DeiT, ViT). The paper interprets this as the core representational structure being more resilient and distributed across a larger set of parameters than the weights needed for the final percentage points of accuracy.
- Functional stability is more brittle than representational structure. The LMC barrier (minimum accuracy drop along the path between the unpruned model and its pruned variant) rises at lower pruning levels and far more sharply than CKA self-similarity falls. For Vision Transformers, the LMC barrier exceeds a 40% accuracy drop at 40% sparsity while CKA remains above 0.7.
- Static similarity and robustness under sparsity are strongly correlated, with important exceptions. Across 21 model pairs, Pearson's r = 0.882 (p < 0.0001). However, 4 of 21 pairs (19.0%) showed significant divergence between CKA and Procrustes scores, which the paper says highlights the risk of relying on a single metric. The paper describes the remaining pairs as exhibiting "Standard" patterns where robustness was consistent with static similarity.
- Pruning can act as a regularizer for some model pairs. For certain pairs, cross-model similarity remains stable or slightly increases at moderate pruning levels before declining, which the authors read as pruning removing run-specific weights while preserving a shared architectural core. At extreme sparsity, similarity may rise again as both models are reduced to a minimal set of generic features.
Methodology in Plain English
The authors take a broad set of pretrained vision models — ResNet18 and ResNet50 (trained on ImageNet, CIFAR-10, and random initialization), ViT-Tiny-Patch16, DeiT-Tiny-Patch16, DINOv2-Base, CLIP-ViT-B/32, BLIP-ViT-B/16, and LLaVA-1.5-7B using only its vision tower — and compare them pairwise.
For each pair, they build a three-panel picture. First, they extract layer-wise activations and compute a full similarity matrix using CKA for relational geometry and Procrustes for geometric alignment, establishing what the representations look like at rest. Second, they assess whether the models compute the same function: if the pair shares an architecture, they evaluate accuracy along the linear interpolation path between the two sets of weights (Linear Mode Connectivity), where a low-error path suggests functionally equivalent solutions; if architectures differ, LMC is not applicable and they instead compute Jensen-Shannon Divergence between the models' average softmax prediction distributions. Third, they progressively prune both models with global magnitude pruning at various sparsity levels, tracking each model's individual accuracy and the cross-model similarity over layers, and additionally trace the LMC barrier and CKA self-similarity between the original and pruned model to see which property degrades first.
The evaluation uses two tiers. Static similarity is computed on 5,000 CIFAR-10 test images upscaled to model resolution — used both to test whether architectural clustering survives out-of-distribution data and to enable rapid prototyping, with the justification that CKA is relatively invariant to input distribution. Dynamic analyses where task accuracy matters use the ImageNetV2 subset of 2,000 images from classes 0–200, so performance metrics reflect the models' intended evaluation domain.
Why This Matters
Impact on research. The paper argues that single-metric comparisons of model representations can support incorrect conclusions — 4 of 21 pairs showed large CKA/Procrustes disagreement — and that combining views yields a more reliable account of whether models share internal mechanisms. It also provides an initial baseline and a practical toolkit for validating findings across datasets rather than on one testbed.
Real-world applications (as motivated by the paper):
- Scientific domains where deep networks are already used, such as medical imaging and particle physics, where understanding whether independently trained models learned the same concepts supports validation of black-box models.
- Model selection, since a pair with high static similarity but low sparsity robustness signals a "brittle core," meaning that although representations align at rest, one model may be significantly less transferable than the other.
- Transfer learning decisions, guided by whether a candidate source model shares a robust computational core with the target architecture.
- Model merging and ensembling, the applied domain of loss landscape geometry research from which the LMC diagnostic is borrowed.
Industry relevance. Teams that must choose among many pretrained checkpoints, or that compress models for deployment, can use the sparsity panel to check whether a relationship between two models survives parameter removal, rather than trusting a static similarity score alone. The paper is explicit that the framework is computationally demanding — CKA scales quadratically with samples (O(N²)) and LMC requires multiple evaluations along high-dimensional paths — which may be expensive for large foundation models.
Future Directions
- Alternative pruning schemes. The authors used global magnitude pruning and state that structured pruning, lottery ticket rewinding, or learned sparsity patterns may reveal different similarity dynamics and should be explored.
- Reducing computational cost. Because CKA is O(N²) in samples and LMC needs repeated evaluations along high-dimensional paths, scaling the framework to large foundation models remains an open practical problem.
- Investigating the disagreement cases. Understanding why 4 of 21 pairs diverged between CKA and Procrustes, and what that implies about the models, is left as an open question.
- Broadening beyond vision. The empirical validation covers only CNNs, Vision Transformers, and Vision-Language Models on CIFAR-10 and ImageNetV2; extending the framework and testing the architectural-clustering claim in other modalities is a natural next step.
Target Audience
Researchers and practitioners in interpretability, representation learning, and model analysis who need to decide whether two models learned the same thing — especially those working on model selection, transfer learning, model merging, or pruning. It is also useful for scientists applying deep networks in domains such as medical imaging or physics who need validation tools for black-box models, and for readers who want an accessible entry point to CKA, Procrustes, and Linear Mode Connectivity in one place.
Authors’ abstract
Comparing neural network representations is essential for understanding and validating models in scientific applications. Existing methods, however, often provide a limited view. We propose the Triangle of Similarity, a framework that combines three complementary perspectives: static representational similarity (CKA/Procrustes), functional similarity (Linear Mode Connectivity or Predictive Similarity), and sparsity similarity (robustness under pruning). Analyzing a range of CNNs, Vision Transformers, and Vision-Language Models using both in-distribution (ImageNetV2) and out-of-distribution (CIFAR-10) testbeds, our initial findings suggest that: (1) architectural family is a primary determinant of representational similarity, forming distinct clusters; (2) CKA self-similarity and task accuracy are strongly correlated during pruning, though accuracy often degrades more sharply; and (3) for some model pairs, pruning appears to regularize representations, exposing a shared computational core. This framework offers a more holistic approach for assessing whether models have converged on similar internal mechanisms, providing a useful tool for model selection and analysis in scientific research.