Research
Perch 2.0 transfers 'whale' to underwater tasks
Overview Research area: Bioacoustics machine learning / transfer learning for marine mammal acoustic classification. Technical level: Intermediate. The evaluation protocol (linear probing on frozen em

- arXiv
- 2512.03219
- Published
- 2025-12-02
- Authors
- Andrea Burns, Lauren Harrell, Bart van Merriënboer, Vincent Dumoulin, Jenny Hamer, Tom Denton
AI summary
Overview
Research area: Bioacoustics machine learning / transfer learning for marine mammal acoustic classification.
Technical level: Intermediate. The evaluation protocol (linear probing on frozen embeddings with few-shot logistic regression) is conceptually simple, but the paper assumes familiarity with embeddings, ROC-AUC, and passive acoustic monitoring.
Scope: A comparative evaluation of seven pretrained bioacoustics embedding models—centered on Perch 2.0—on three marine and underwater audio datasets, using few-shot linear probing rather than the models' own classification heads.
What This Paper Is About
Marine bioacoustics has far less labeled data than terrestrial bioacoustics, because underwater recording requires specialized ceramic pressure sensors, water-sealed electronics, mooring buoys or divers, and visual confirmation of the vocalizing species is often impossible. Perch 2.0 is a strong terrestrial bioacoustics foundation model trained on 14,597 species (mostly birds, plus insects, mammals, and amphibians) with almost no marine mammal audio in its training data, so the authors ask whether its embeddings nevertheless transfer to underwater tasks. The goal is to give practitioners guidance on which pretrained embedding model to pick when building new linear classifiers for marine mammal sounds from only a few labeled examples.
Key Contributions
- A few-shot transfer-learning benchmark for Perch 2.0 on marine audio, evaluated on three validation sets: DCLDE 2026, NOAA PIPAN, and ReefSet, in addition to the previously reported Watkins Marine Mammal Sound Database (WMMSD) result from the BEANS benchmark.
- A head-to-head comparison of seven pretrained embedding models — Perch 2.0, Perch 1.0, SurfPerch, Google Multispecies Whale Model (GMWM), Birdnet V2.3, AVES-bio, and BirdAVES (large) — restricted to models supported in the Perch Hoplite GitHub repository.
- A decomposition of DCLDE 2026 into three distinct tasks (Species, Ecotype, Known species), including a five-way killer whale ecotype classification (NRKW, OKW, SAR, SRKW, TKW).
- An analysis of why terrestrial models transfer to whales, combining neural scaling arguments, the "bittern lesson" about birds as difficult supervision, and the shared myoelastic-aerodynamic sound production mechanism, supported by tSNE visualizations of ecotype separability.
Main Findings
- Perch 2.0 leads on most cetacean tasks. Under the non-contaminated comparison, Perch 2.0 outperforms the published comparison models on ReefSet and on all cetacean species tasks (DCLDE 2026 Species, DCLDE 2026 Ecotype, and NOAA PIPAN whales).
- One exception. On DCLDE 2026 Known Bio Species, Perch 2.0 comes in second to BirdNet V2.3, which reaches 0.990 at k=8 and 0.991 at k=16 versus Perch 2.0's 0.983 and 0.989.
- Headline AUC-ROC numbers for Perch 2.0 (k=8, k=16): DCLDE 2026 Species 0.970 / 0.977; DCLDE 2026 Ecotype 0.917 / 0.945; DCLDE 2026 Known Bio Species 0.983 / 0.989; NOAA PIPAN 0.863 / 0.924; ReefSet 0.975 / 0.981.
- Data contamination hurts transfer. SurfPerch, trained on ReefSet, scores 0.982* / 0.986* there but drops on other tasks; GMWM, trained on a large portion of NOAA PIPAN labeled audio, scores 0.868* / 0.917* there but is weaker elsewhere.
- Off-the-shelf classification heads underperform embeddings. GMWM's pretrained classification scores for supported classes on DCLDE yield AUC-ROC 0.612, but using its embeddings for few-shot learning jumps to 0.954. The authors suggest the poor classification-head result may be due to overfitting to a particular microphone or other training-data characteristics.
- BirdNet V2.3 is a strong runner-up. It trails Perch 2.0 closely on most tasks (e.g., NOAA PIPAN 0.855 / 0.924, DCLDE 2026 Species 0.942 / 0.959).
- AVES variants lag on small sample counts. AVES-bio and BirdAVES generally show lower scores at low k, though the authors note the pooling method could potentially be modified for improvement.
- Embeddings are discriminative at fine granularity. A tSNE plot of Perch 2.0 embeddings (PCA-compressed to 32 dimensions first) shows strong separability between killer whale ecotypes within the same species.
- Qualitative tSNE comparisons. GMWM embeddings show weak linear separability; AVES-bio and BirdAVES show more class entanglement, especially for the SRKW ecotype; AVES models do not separate TKW and NRKW well; Perch 1.0 separates SRKW less well than Perch 2.0, SurfPerch, or BirdNet; Perch 2.0 appears to have the best boundary between NRKW and TKW.
- Class dropping in NOAA PIPAN. Class "Bm" is dropped for k=16, and classes "Bm" and "Be" are dropped for k=32, because fewer than k+1 recording-level embeddings were available.
- Model scale differs widely. Among the compared models, parameters range from 4.1M (GMWM) to 315.4M (BirdAVES large), with Perch 2.0 at 101.8M and a 1536-dimensional embedding—the largest embedding dimension in the comparison.
Methodology in Plain English
The authors do not use the models' own predictions. Instead, they treat each model purely as an embedding extractor and train a simple classifier on top:
- Chunk and embed. Each recording is split into fixed windows (5 seconds or 3 seconds, depending on what the model expects), with the hop size equal to the window size, and every window is passed through the model.
- Pool to recording level. If a labeled recording is longer than the window, the window embeddings are averaged and normalized into a single recording-level vector. Mean pooling is applied for all models on NOAA PIPAN and for DCLDE examples longer than 3.0 or 5.0 seconds.
- Few-shot linear probing. For each class, k recording-level embeddings are sampled at random, where k ∈ {4, 8, 16, 32}, and a logistic regression classifier is trained. The remaining embeddings are held out to compute a one-vs-all ROC-AUC.
- Repeat and average. The whole process is repeated 5 times with independently sampled training sets, and the average ROC-AUC is reported.
- Score only clean comparisons. The highest performance among non-contaminated models is bolded, and asterisks mark datasets that were in an embedding model's training data. Perch 2.0's training data includes around a dozen cetacean recordings from iNaturalist, but these were mostly above-water phone recordings, not underwater hydrophone recordings.
- Note on implementation. The authors state their scores may differ slightly from previously published numbers because they use scikit-learn's
LogisticRegression(L-BFGS optimizer with default weight decay), whereas SurfPerch's earlier ReefSet evaluation used mini-batches and the Adam optimizer.
Why This Matters
Impact on research: The results suggest that large terrestrial bioacoustics foundation models are viable embedding backbones for underwater work, which matters because marine labeled data is expensive to collect and verify. The embedded-in-a-vector-database plus linear-probe workflow supports agile modeling: an active learning loop where new sounds can be found by vector search and classified within minutes rather than requiring a new training run.
Real-world applications:
- Passive acoustic monitoring for conservation and ecology, which the paper calls a critical tool, including detecting and annotating large hydrophone archives at scale.
- Baleen whale species identification, covering common minke, humpback, sei, blue, fin, and Bryde's whales, plus unknown whale sounds and anthropomorphic noise in the NOAA PIPAN class set.
- Killer whale ecotype discrimination (NRKW, OKW, SAR, SRKW, TKW), which is relevant to differentiating locally adapted sub-populations.
- Rapid triage of newly discovered or evolving sounds, such as mystery underwater sounds later attributed to minkes and Bryde's whales and constantly evolving humpback songs that move across populations.
Industry relevance: Any organization running hydrophone deployments—fisheries science centers, environmental consultancies, marine energy developers, shipping and naval monitoring programs—can adopt pretrained embeddings instead of training models from scratch. The evaluation is deliberately limited to models available in the open-source Hoplite repository, so the recommendations are directly actionable with existing tooling.
Future Directions
- Improve embedding pooling for AVES-family models. The authors explicitly note that lower performance at small k may be due to the mean-pooling approach and could potentially be modified for better results.
- Add real marine mammal data to pretraining. Perch 2.0's only cetacean exposure is roughly a dozen mostly above-water iNaturalist phone recordings; increasing underwater representation is the obvious lever.
- Test broader taxonomic and geographic coverage. The current evaluation covers three validation sets plus WMMSD; whether these findings hold for other species and recording conditions is not reported.
- Investigate why bird supervision transfers. The proposed explanations—scaling laws, the "bittern lesson" about bird classification as a difficult supervision task, and the shared myoelastic-aerodynamic sound production mechanism—remain hypotheses that are not directly tested here.
Target Audience
Practitioners and researchers building classifiers for marine mammal and underwater audio, especially those working with passive acoustic monitoring and limited labeled data. Also relevant to bioacoustics ML researchers studying cross-taxon transfer, and to engineers choosing a pretrained embedding model for an agile modeling pipeline. Readers who only want the practical takeaway can read the abstract, Table 2, and the discussion; readers interested in the negative results should look at the off-the-shelf GMWM comparison and the appendix tSNE plots.
Authors’ abstract
Perch 2.0 is a supervised bioacoustics foundation model pretrained on 14,597 species, including birds, mammals, amphibians, and insects, and has state-of-the-art performance on multiple benchmarks. Given that Perch 2.0 includes almost no marine mammal audio or classes in the training data, we evaluate Perch 2.0 performance on marine mammal and underwater audio tasks through few-shot transfer learning. We perform linear probing with the embeddings generated from this foundation model and compare performance to other pretrained bioacoustics models. In particular, we compare Perch 2.0 with previous multispecies whale, Perch 1.0, SurfPerch, AVES-bio, BirdAVES, and Birdnet V2.3 models, which have open-source tools for transfer-learning and agile modeling. We show that the embeddings from the Perch 2.0 model have consistently high performance for few-shot transfer learning, generally outperforming alternative embedding models on the majority of tasks, and thus is recommended when developing new linear classifiers for marine mammal classification with few labeled examples.