Research
Diffusion Representations for Fine-Grained Image Classification: A Marine Plankton Case Study
Overview Research area: Computer vision / self-supervised representation learning applied to fine-grained image classification, with a domain focus on marine plankton imagery from Imaging FlowCytobot

- arXiv
- 2601.13416
- Published
- 2026-01-19
- Authors
- A. Nieto Juscafresa, Á. Mazcuñán Herreros, J. Sullivan
AI summary
Overview
Research area: Computer vision / self-supervised representation learning applied to fine-grained image classification, with a domain focus on marine plankton imagery from Imaging FlowCytobot (IFCB) instruments.
Technical level: Intermediate. The core idea (freeze a diffusion model, extract features, train a linear classifier) is simple, but the paper includes a full mathematical treatment of diffusion forward/reverse processes, noise schedules, and loss weighting.
Scope: The paper asks whether intermediate activations from a label-free, from-scratch-trained diffusion U-Net can serve as frozen features for fine-grained plankton recognition, and how layer choice, noise level, and training objective relate to classification accuracy.
What This Paper Is About
Diffusion models are trained to denoise and generate images, but the paper asks whether those same denoising networks incidentally learn features good enough to classify images they were never trained to label. The authors test this on plankton, a genuinely hard fine-grained problem with severe class imbalance, low contrast, and species that differ only by small spines, flagella, or faint internal structure. Their goal is to find which decoder layer and which noise timestep produce the most linearly separable features, and to understand how generation quality relates to classification quality.
Key Contributions
-
A formalized layer–timestep selection problem. The authors define a search over 12 U-Net decoder readout locations (
ℓ ∈ {1,…,12}) and a set of noise timesteps, fit a linear probe at every pair, and select the pair(t*, ℓ*)that maximizes validation accuracy. They report a stable low-depth, moderate-noise region as most useful. -
Evaluation across three plankton regimes. Curated balanced, realistic long-tailed, and out-of-distribution (OOD) evaluations, with compute amortized by reusing precomputed descriptors across related datasets.
-
A link between generation quality and recognition. The paper analyzes how the training objective and its per-timestep loss schedule shape learned features, and diagnoses when overfitting during diffusion training breaks the relationship between FID and discriminative accuracy.
Main Findings
-
Best readout pair is
(t*, ℓ*) = (25, 3). The highest linear-probe accuracy comes from timestep 25 at decoder readout location 3, which is RB3 at 16² resolution, under the paper's flattening indexℓ = 3(r−1)+b. That location combines self-attention with a compact 16×16 spatial grid and 512 channels. -
Accuracy drops deeper in the decoder. Probe accuracy decreases for later decoder blocks (higher spatial resolution, lower channel count), suggesting subsequent upsampling prioritizes synthesis over linear separability.
-
Competitive with supervised baselines on the balanced dataset. Frozen diffusion features reach 0.9240 accuracy and 0.9114 Macro F1, against EfficientNet-B0 at 0.9490 / 0.9455, EfficientNet-B3 at 0.9508 / 0.9447, ViT-B/16 at 0.9209 / 0.9128, and ResNet-50 at 0.9281 / 0.9229, with 57M parameters versus 87M for ViT-B/16.
-
Best Macro F1 on the unbalanced dataset. The diffusion representation reaches 0.9145 accuracy and 0.9064 Macro F1 on the 120-class unbalanced dataset, ahead of EfficientNet-V2-S (0.9100 / 0.9050) and ResNet-50 (0.9759 / 0.8864). Note accuracy and Macro F1 disagree here: ResNet-50 has higher accuracy but lower Macro F1.
-
Clear margin over self-supervised baselines. DINOv3 (ViT-B/16) reaches 0.8031 accuracy / 0.6145 Macro F1 balanced and 0.7909 / 0.3274 unbalanced; MAE (ViT-B/16) reaches 0.8467 / 0.8214 balanced and 0.8207 / 0.8001 unbalanced.
-
Strong transfer under distribution shift. On the temporal-shift Baltic OOD dataset, features trained on the balanced source reach 0.9133 accuracy / 0.7356 Macro F1; trained on the unbalanced source, 0.9112 / 0.7564. On the geographically and taxonomically shifted WHOI-22 dataset, 0.8411 / 0.8392 (balanced source) and 0.8272 / 0.8260 (unbalanced source). The larger Macro F1 decrease appears under the WHOI-22 shift.
-
MinSNR weighting outperforms uniform MSE. With
w_MinSNR(t) = min(SNR(t), γ)/SNR(t)and γ = 5, features maintain classification accuracy as SNR decreases and improve slightly in the high-SNR regime, whereas uniform MSE weighting degrades steadily on the IFCB experiments. -
FID and downstream accuracy decouple. After a certain point, validation loss shows overfitting while FID keeps improving. A checkpoint selected for better FID gave 0.8812 accuracy versus 0.9240 for the minimum-validation-loss checkpoint, roughly a four-point absolute decline.
-
Features appear organized by morphology. PCA on decoder feature tokens, with the first three components mapped to RGB, produced consistent colors along each organism across noise levels, suggesting organization by plankton morphology rather than raw pixel values.
Methodology in Plain English
The researchers train a U-Net diffusion model from scratch on plankton training images, with no labels used during denoiser training. It is unconditional, runs for T = 1000 diffusion steps with a cosine noise schedule, and is optimized with AdamW (learning rate 5·10⁻⁴, β₁ = 0.9, β₂ = 0.999, weight decay 10⁻⁴), cosine learning-rate decay with 5% warmup, gradient clipping at 1.0, mixed precision, and an EMA with decay 0.999, training for 250 epochs at batch size 256, selecting the checkpoint with lowest validation loss at epoch 100.
To get features, they take an image, corrupt it to a chosen noise level t to form x_t, and feed both x_t and t through the frozen denoiser. They read activations after each of the 12 residual blocks in the decoder, spanning four stages at {16², 32², 64², 128²} with three residual blocks each. Each activation tensor is reduced to a vector by global average pooling. For every (t, ℓ) pair they fit a separate linear softmax classifier on the training split, then pick the pair with the highest validation accuracy. The final feature for downstream tasks is the pooled activation at that selected pair. The experiments sweep t ∈ {1, 10, 25, 50, 75, 100, 200, 400, 600}.
Data comes from IFCB images in the SMHI and SYKE Baltic programs. After preprocessing, the balanced dataset has 70 classes with 500 single-channel grayscale images per class resized to 128×128 pixels; the unbalanced dataset has 120 classes with varying image counts. FID is computed by replicating grayscale ROIs to three channels and using a standard Inception-V3 feature extractor. Baselines include supervised ResNet-50, EfficientNet-B0/B3/V2-S, ViT-B/16, and an MLP on hand-crafted SIFT and edge features, plus self-supervised MAE and DINOv3 (ViT-B/16) that are fine-tuned on target training images with their own label-free objectives and then frozen for linear probing.
Why This Matters
Research impact. The paper extends work on diffusion representations (including DDAE) into the fine-grained, low-contrast regime, where prior diffusion-timestep choices like t = 45, t = 50, or t = 261 were established in other domains. It argues that generative fidelity and discriminative strength should be evaluated separately, and that FID should not be the optimization target when training diffusion models for representation learning.
Real-world applications.
- Marine ecosystem monitoring, where plankton communities control key ocean processes and serve as markers of climate-driven change including warming, acidification, and shifting nutrient regimes.
- Automated classification of the large, low-contrast image streams produced by in-situ instruments, supporting abundance estimates and trend analysis where labels are scarce.
- Cross-instrument and cross-region deployment, since the same frozen backbone transfers to temporally shifted Baltic data and geographically and taxonomically shifted WHOI-22 data with only a linear probe refit.
- Guidance for practitioners on where to read features from a diffusion U-Net, potentially reducing per-dataset adaptation cost.
Industry relevance. Ocean observing programs and operators of imaging flow cytometers can reuse one trained backbone across datasets and retrain only a lightweight linear probe. The reported cost is 15 hours on 4× NVIDIA A100-40GB GPUs for the initial diffusion training, with peak memory around 38 GB per GPU, 24 CPU threads, and approximately 11 GB of disk space for checkpoints. The paper notes diffusion training is costlier than CNNs, but that the cost is amortized through backbone reuse.
Future Directions
- Extending beyond plankton. The feature extractor makes no instrument- or dataset-specific assumptions, so the paper leaves open whether the same layer–timestep selection generalizes to other fine-grained or environmental microscopy domains.
- Reconciling FID with representation quality. The paper shows FID and downstream accuracy decouple under overfitting but does not propose a replacement selection criterion for diffusion training in representation-learning settings.
- Reducing diffusion training cost. The paper notes the costlier training relative to CNNs and cites complementary work on lightweight unsupervised fine-tuning of diffusion backbones (CleanDIFT) as an alternative to the noised-input paradigm; whether that approach changes the optimal
(t*, ℓ*)is not reported here. - Characterizing the attention contribution. Only one readout location includes a self-attention module, at the 16² decoder stage, and the authors attribute the strongest separability partly to it; a controlled comparison against purely convolutional variants is not reported.
Target Audience
Researchers and practitioners working on self-supervised representation learning and diffusion models who want evidence on using frozen diffusion features without fine-tuning. It is also directly relevant to marine ecologists, ocean-observing engineers, and machine learning engineers building taxon classifiers from IFCB or similar imaging flow cytometry data, particularly those facing long-tailed class distributions and domain shift. Readers seeking a detailed mathematical treatment of diffusion forward and reverse processes will also find a full background section.
Authors’ abstract
Diffusion models have emerged as state-of-the-art generative methods for image synthesis, yet their potential as general-purpose feature encoders remains underexplored. Trained for denoising and generation without labels, they can be interpreted as self-supervised learners that capture both low- and high-level structure. We show that a frozen diffusion backbone enables strong fine-grained recognition by probing intermediate denoising features across layers and timesteps and training a linear classifier for each pair. We evaluate this in a real-world plankton-monitoring setting with practical impact, using controlled and comparable training setups against established supervised and self-supervised baselines. Frozen diffusion features are competitive with supervised baselines and outperform other self-supervised methods in both balanced and naturally long-tailed settings. Out-of-distribution evaluations on temporally and geographically shifted plankton datasets further show that frozen diffusion features maintain strong accuracy and Macro F1 under substantial distribution shift.