Research
Why Prototypes Collapse: Diagnosing and Preventing Partial Collapse in Prototypical Self-Supervised Learning
Overview Research area: Self-supervised learning (SSL) for computer vision, specifically prototypical joint-embedding methods (DINO, DINOv2, CARP, iBOT, CAPI, SWaV-style frameworks) and the training d

- arXiv
- 2510.20108
- Published
- 2025-10-23
- Authors
- Gabriel Y. Arteaga, Marius Aasan, Rwiddhi Chakraborty, Martine Hjelkrem-Tan, Thalles Silva, Michael Kampffmeyer, Adín Ramírez Rivera
AI summary
Overview
Research area: Self-supervised learning (SSL) for computer vision, specifically prototypical joint-embedding methods (DINO, DINOv2, CARP, iBOT, CAPI, SWaV-style frameworks) and the training dynamics of learnable prototype/cluster-anchor vectors.
Technical level: Advanced. The paper relies on prototype-collapse definitions with angular thresholds, online expectation-maximization for Gaussian mixtures, and terminology from the joint-embedding SSL literature.
Scope: The paper diagnoses why multiple prototypes in prototypical SSL converge to near-identical representations ("partial prototype collapse") across a broad set of frameworks, and proposes a fully decoupled training scheme that estimates prototypes as an online Gaussian mixture independently of the encoder's loss.
What This Paper Is About
Prototypical SSL methods learn a set of prototype vectors that act as cluster anchors guiding the encoder to distribute samples into semantically coherent regions. In practice, many of these prototypes end up nearly identical, so the target signal becomes redundant and far less informative than intended. The authors argue this happens because the encoder and the prototypes are optimized jointly under one shared consistency loss, which lets prototypes drift into redundant shortcuts early in training, and they replace that joint optimization with a decoupled procedure in which prototypes are estimated by an independent online Gaussian-mixture model.
Key Contributions
- A systematic, quantitative analysis across many prototypical SSL frameworks — not just the DINO family — showing that partial prototype collapse is widespread (DINO, CARP, DINOv2, iBOT variants), and also objective-specific within a single model (DINOv2's dense versus instance-level heads).
- Identification of the underlying mechanism: joint optimization of encoder and prototypes under a shared loss, which the authors characterize as a form of shortcut learning in which prototypes minimize the loss without enriching representation diversity.
- A fully decoupled training framework that estimates prototypes from latent features via an independent online GMM/EM-style objective, while the encoder is trained separately against fixed prototypes, requiring no explicit diversity regularizer.
- Empirical validation that decoupling eliminates partial prototype collapse across all tested thresholds, improves early and late training performance relative to CARP, gives a 1.8-percentage-point k-NN improvement on ImageNet-1k, and yields a 3.0-percentage-point overall gain on long-tailed iNaturalist 2018 (with gains across head, medium, and tail classes), while slightly reducing peak memory at large batch sizes and adding 0.2 hours of training time.
Main Findings
- Collapse is not confined to DINO. Using Definition 2.1 with ε = 0.025 (prototypes within 12.84° counted as collapsed), DINO retained 908 unique prototypes out of 60,000 initialized (1.5%), and CARP retained 7,052 out of 65,536 (10.8%).
- CAPI, iBOT and DINOv2 behave very differently. CAPI (dense objective) kept 16,383 of 16,384 prototypes unique (99.9%); iBOT kept 3,057 of 8,192 (37.3%); iBOT-vMF + KP kept 7,895 of 8,192 (96.4%); DINOv2 (262,144 initialized prototypes) kept 110,201 unique dense prototypes and 2,556 unique instance-level prototypes, with the paper describing the dense head as relatively diverse and the instance-level head as suffering near-total collapse.
- Collapse appears very early. In training-dynamics tracking over the first 100 epochs on ImageNet-1k, after only 10 epochs two-thirds of the prototypes had already collapsed.
- Explicit regularization helps but does not fully solve it. KoLeo-Prototype regularization (Govindarajan et al., 2024) maintains diversity, and DINO + KP ended slightly above vanilla DINO despite a modest early-accuracy dip — but it introduced a trade-off hyperparameter λ_KP and, as the authors note, did not fully solve the problem.
- Framework-level differences at training end. CARL retained only about 0.2% of its initialized prototypes as unique by the end of training, while CARP preserved 9%, corresponding to a margin of roughly 4 percentage points in linear accuracy — though the authors caution this cannot be attributed to prototype uniqueness alone.
- Partial decoupling helps, full decoupling removes collapse. CAPI's teacher-branch prototypes are updated by a separate clustering module; the authors attribute CAPI's robustness to this partial decoupling. It retained approximately 38% of its initialized prototypes at the stricter threshold ε = 0.5. Under full decoupling, no evidence of partial prototype collapse was observed at any tested ε, including ε = 0 (no restriction), ε = 0.025, and ε = 0.5 (which counts prototypes as unique only if separated by at least 60°).
- Diversity is not universally beneficial. Applying decoupling to DINO on iNaturalist 2018 caused a large drop (overall 36.2%, down 9.1 points), which the authors attribute to DINO's high sensitivity to excessive prototype diversity and to hyperparameters that did not transfer from ImageNet-1k tuning. CARP + KP also slightly decreased on iNaturalist (overall 45.3%, down 0.6), which they link to reusing a KP strength tuned for DINO.
- Long-tailed gains span all class frequencies. On iNaturalist 2018 (about 430K images, 8,142 classes), CARP + Decoupling reached 59.1 head (>100), 49.3 medium (>20 and ≤100), 45.9 tail (≤20), and 48.9 overall — gains of 3.1, 2.4, 3.4, and 3.0 points over CARP. DINO + KP reached 49.0 overall, a 3.7-point gain, with gains in every regime.
- k-NN improves, linear evaluation is mixed. Adding decoupling to CARP improved k-NN by 1.8 percentage points (e.g., 67.7 to 69.1 for ResNet-50 at 400 epochs; 73.6 to 74.1 for ViT-S/16 at 300 epochs; 75.3 at ViT-S/16 with 800 epochs). Linear accuracy stayed essentially flat or slightly lower (75.3 to 75.3 for ResNet-50; 76.3 to 76.2 for ViT-S/16 at 300 epochs), and the authors emphasize k-NN as the more robust, tuning-free metric.
- Transfer learning does not substantially change. Across ResNet-50 and ViT-Small/16, decoupled variants matched CARP within small variations (mean over eight datasets: 79.18 versus 79.06 for ResNet-50; 79.66 versus 79.46 for ViT-Small/16).
- Memory and time costs are small. Peak memory per GPU for CARP + Decoupling was lower than CARP at every tested batch size (5.3G vs 5.4G at 128; 60.9G vs 62.4G at 2048, measured on AMD Instinct MI250X GPUs), with savings growing at larger batch sizes; training time increased by only 0.2 hours (37.9h to 38.1h for 100 epochs, batch size 1024, 12 crops).
Methodology in Plain English
The authors first check whether collapse is a DINO-only artifact by inspecting publicly released prototype weights from several frameworks and counting how many prototypes remain unique under Definition 2.1 at ε = 0.025, and additionally at ε = 0 and a stricter ε = 0.5. They train baselines themselves (except in the released-weights analysis) to obtain intermediate checkpoints and watch prototype uniqueness alongside ImageNet-1k linear accuracy over the first 100 epochs.
They then formalize the standard setup — minimize a consistency loss jointly over the encoder parameters and the prototype set — and argue this joint objective encourages prototypes to drift into redundant regions early on. Their fix is to alternate two separate problems: first estimate the prototypes from current latent features using an objective that does not involve the encoder's loss, then train the encoder with prototypes held fixed. Because clustering with K-Means every iteration is too noisy and too expensive, and clustering once per epoch leaves prototypes stale, they model prototypes as the means of an online Gaussian mixture updated incrementally by an EM-style procedure. To handle high-dimensional features and large prototype counts, they adapt the update with responsibility-weighted forgetting and deterministic annealing, so rarely used components stay stable while heavily used ones remain responsive. The result is a plug-in change: the encoder's loss and architecture are untouched, only the prototype update path is separated. Full implementation details are placed in Appendix A, and a theoretical motivation in Appendix C.1.
Why This Matters
Research impact. The paper reframes partial prototype collapse from a DINO-family quirk into a general consequence of joint optimization, gives a simple intervention that removes the need for ad-hoc diversity regularizers and their trade-off hyperparameters, and offers a diagnostic tool (unique-prototype counts tracked through training) that other researchers can apply to their own frameworks. It also shows that removing collapse is not automatically good — DINO degraded badly on imbalanced data under decoupling — which tempers a common assumption in the literature.
Potential real-world applications (derived from properties the paper reports, not applications tested in the paper):
- Foundation-model pretraining on web-scale, uncurated image collections, where long-tailed class distributions are the norm and the paper reports a 3.0-percentage-point overall improvement from decoupling on iNaturalist 2018.
- Biodiversity and ecological monitoring, where the tail-class behavior the paper measures (tail accuracy 45.9 versus 42.5) matters most because rare species are exactly the ones with few training images.
- Memory-constrained large-batch training pipelines, since decoupling reduced peak per-GPU memory at every tested batch size, with savings growing up to 1.5G at batch size 2048.
- Pipelines that cannot afford extra hyperparameter tuning, because decoupling requires no explicit diversity regularizer (and no λ_KP-style balancing term), and the authors report CARP maintains robust performance across a wide range of hyperparameters.
Industry relevance. Prototypical SSL backbones underpin modern vision foundation models, and prototype count is a major knob in those systems — the paper notes that practitioners over-parameterize prototype sets to compensate for collapse. Decoupling changes that cost calculus: it keeps prototypes diverse without adding regularizer tuning, reduces memory at scale, and adds only 0.2 hours of training time, while leaving the encoder architecture and consistency loss unchanged.
Future Directions
- Ensure prototype updates track the encoder's rate of change. The authors note the current implementation relies only on the forgetting factor η, which cannot adapt to shifts in the online mixture estimate.
- Reduce the assumptions and hyperparameters of the online GMM, which adds new hyperparameters and assumes incoming representations are Gaussian-distributed.
- Investigate patch-level objectives. Because patch-level objectives have substantially higher memory requirements, the authors expect memory savings from decoupling to be even larger for methods such as iBOT or DINOv2, and leave this to future work — along with isolating the effect of masked image modeling objectives on prototype diversity.
- Reconcile full diversity with framework-specific stability. The authors state that complete removal of collapse is not universally beneficial across frameworks, leaving open the question of how much prototype diversity each framework can absorb before optimization becomes unstable.
Target Audience
SSL researchers and graduate students working on joint-embedding and clustering-based representation learning; practitioners pretraining vision backbones who face decisions about prototype counts, diversity regularizers, and long-tailed data; and engineers evaluating memory and compute trade-offs for large-batch, large-prototype training runs. Readers without background in prototypical SSL or mixture-model EM will need to consult the appendices and cited frameworks (DINO, DINOv2, iBOT, CARP, CAPI, SwAV), since the main text assumes familiarity with the standard prototypical formulation.
Authors’ abstract
Prototypical self-supervised learning methods consistently suffer from partial prototype collapse, where multiple prototypes converge to nearly identical representations. This undermines their central purpose -- providing diverse and informative targets to guide encoders toward rich representations -- and has led practitioners to over-parameterize prototype sets or add ad-hoc regularizers, which mitigate symptoms rather than address the root cause. We empirically trace the collapse to the joint optimization of encoders and prototypes, which encourages a type of shortcut learning: early in training prototypes drift toward redundant representations that minimize loss without necessarily enhancing representation diversity. To break the joint optimization, we introduce a fully decoupled training strategy that learns prototypes and encoders under separate objectives. Concretely, we model prototypes as a Gaussian mixture updated with an online EM-style procedure, independent of the encoder's loss. This simple yet principled decoupling eliminates prototype collapse without explicit regularization and yields consistently diverse prototypes and stronger downstream performance.