Research
SpidR: Learning Fast and Stable Linguistic Units for Spoken Language Models Without Supervision
Overview Research area: Self-supervised speech representation learning and textless (pure) spoken language modeling. Technical level: Intermediate. The core idea is intuitive (students predicting a te

- arXiv
- 2512.20308
- Published
- 2025-12-23
- Authors
- Maxime Poli, Mahi Luthra, Youssef Benchekroun, Yosuke Higuchi, Martin Gleize, Jiayi Shen, Robin Algayres, Yu-An Chung, Mido Assran, Juan Pino, Emmanuel Dupoux
AI summary
Overview
- Research area: Self-supervised speech representation learning and textless (pure) spoken language modeling.
- Technical level: Intermediate. The core idea is intuitive (students predicting a teacher's own clustering labels), but the paper assumes familiarity with masked prediction, exponential moving averages, and vector quantization.
- Scope: This paper introduces SpidR, a single-pass self-supervised speech encoder whose intermediate layers predict codebook assignments from the matching teacher layers, and shows that this objective stabilizes online clustering, improves zero-shot spoken language modeling, and cuts pretraining cost to 23 hours on 16 A100 GPUs.
What This Paper Is About
Training a language model directly from speech, without any text, requires a speech encoder that turns audio into discrete units carrying linguistic rather than speaker or acoustic information. Existing encoders such as HuBERT are expensive to train (multiple passes alternating between clustering and model training), while the single-pass alternative DinoSR tends to suffer codebook collapse. The goal of this work is to build a fast, stable, single-pass speech encoder whose units are good enough to make textless spoken language models work well.
Key Contributions
- SpidR, a new self-supervised speech encoder. Trained on raw waveforms with a masked prediction objective combined with self-distillation and online clustering. Its distinguishing change from DinoSR is that the student's layer k representation predicts the pseudo-labels derived from the teacher's layer k, instead of routing all K prediction heads through the student's final layer. This layer-aligned objective is reported to make online clustering more resistant to codebook collapse.
- Systematic validation of unit-quality metrics as proxies. The paper evaluates across models and layers the correlation between speech unit quality (ABX, PNMI) and downstream zero-shot spoken language modeling at lexical, grammatical, and semantic levels, concluding that these metrics reliably predict downstream scores.
- A minimal PyTorch codebase and a large training speedup. Full pretraining on LibriSpeech is reported as one day on 16 GPUs instead of a week, with 23 hours and 369 GPU hours for SpidR's own configuration, enabled by the training method and an efficient implementation with model code based on HuBERT from torchaudio.
- Open release. Training code and model checkpoints are released at https://github.com/facebookresearch/spidr.
Main Findings
- SpidR reaches the best phonetic discriminability among the compared encoders. At layer 6, SpidR records ABX within-speaker of 3.32 (dev-clean) and 3.74 (dev-other), and ABX across-speaker of 3.66 (dev-clean) and 4.95 (dev-other), versus HuBERT at layer 11 (3.38, 4.26, 4.01, 6.49), DinoSR at layer 5 (4.05, 5.11, 4.72, 7.29), wav2vec 2.0 at layer 6 (4.47, 5.63, 5.25, 7.82), and data2vec at layer 4 (4.41, 5.49, 5.07, 7.40). Chance level is 50%.
- Word-level retrieval (MAP over words). SpidR scores 66.50 (dev-clean) and 55.26 (dev-other), ahead of HuBERT (46.07, 33.37), DinoSR (63.02, 45.86), wav2vec 2.0 (44.81, 31.92), and data2vec (39.34, 27.90). data2vec 2.0 scores 69.38 on dev-clean and 53.49 on dev-other.
- SpidR improves zero-shot spoken language modeling. With K-means units (layer 6), SpidR reaches 71.89 sWUGGY all / 82.46 in-vocab, 59.48 sBLIMP, and 70.46 tSC, compared with HuBERT at layer 11 (65.50, 73.67, 55.60, 68.75) and wav2vec 2.0 at layer 6 (62.29, 68.50, 53.34, 65.97). With codebook units it reaches 69.78, 79.98, 58.10, and 70.14, and the paper states this is enough to surpass WavLM Base (69.74, 79.88, 56.60, 70.35).
- Codebook predictions are competitive with K-means, unlike for DinoSR. DinoSR's codebook units score 60.10, 64.56, 57.04, and 69.44, while its K-means units drop to 56.69, 59.42, 54.32, and 65.76. The paper attributes the gap to SpidR's aligned prediction scheme.
- Higher codebook diversity during training. Codebook and prediction perplexities tracked on LibriSpeech dev-clean with K = 8 codebooks show that DinoSR collapses especially in the last layers, while SpidR does not, which the authors link to reduced distribution shift between embeddings and codebooks.
- Pretraining is far cheaper than prior work. SpidR and the authors' DinoSR reimplementation use 16 A100 GPUs, 400k steps, a 63-minute batch, 23 hours, and 369 GPU hours, against 2880 GPU hours for the original DinoSR (16 V100s, 180 hr), 1984 for HuBERT (32 A100s, 62 hr), 1920 for Academic HuBERT, 688 for data2vec 2.0, 513 for k2SSL Zipformer, and 300 for MelHuBERT.
- The advantage holds across vocabulary sizes and data scales. In the scaling analysis over 600h, 6k, and 60k hours of Libri-Light, SpidR consistently beats HuBERT in all conditions but with similar scaling slopes, described as a constant advantage rather than a steeper curve. Text LMs trained under matching conditions achieve better performance and superior scaling, particularly on tSC.
- Results across numbers of units. With vocabulary sizes of 50, 100, 200, and 500, SpidR (K-means) scores sWUGGY all of 68.51, 71.27, 70.98, and 70.94, and tSC of 73.13, 70.78, 70.35, and 70.03. The paper notes SpidR with standard K-means matches HuBERT with DC-Spin units, with DC-Spin ahead on sBLIMP and SpidR ahead on the other metrics.
Methodology in Plain English
The model takes a speech waveform, cuts it into 20 ms frames using a stack of convolutional layers, and randomly masks some of those frames. A student Transformer with 12 layers sees the masked input; a teacher Transformer with the same architecture sees the unmasked input. For each of the top 8 layers, a small prediction head on the student tries to guess which codebook entry the teacher's corresponding layer would have been assigned to at the masked positions, trained with cross-entropy.
The labels come from online clustering: each layer has its own codebook of 256 codewords, and the teacher's frame embedding is assigned to its nearest codeword. The teacher is not trained by gradients; it is an exponential moving average of the student, with a smoothly increasing decay schedule. Codewords are updated with an exponential moving average of the teacher embeddings assigned to them, with non-activated codewords left in place.
The key design choice is layer alignment: layer k of the student predicts labels from layer k of the teacher, rather than all heads reading the student's last layer as in DinoSR. The authors also remove biases from the attention Q, K, V projections to avoid weight-norm growth that caused loss spikes, and freeze the feature extractor after 200k steps.
For evaluation, they pick the layer with the lowest average ABX, quantize it either with K-means (trained on LibriSpeech train-clean-100) or by taking the codebook prediction, deduplicate tokens, and train OPT-125M language models on the 6k-hour Libri-Light subset for 25k steps with a context length of 2048 and up to 81920 tokens per batch. Zero-shot scores are then computed on sWUGGY (lexical), sBLIMP (syntactic), and tSC (semantic) by comparing sequence log-likelihoods normalized by token count.
Why This Matters
Impact on research. The paper shows that a single-pass training objective, not just architecture, can fix codebook collapse, and it offers a cheap pretraining recipe plus an open codebase, which lowers the entry cost for experimenting with speech self-supervised models. It also gives evidence that ABX and PNMI are usable proxies for downstream spoken language modeling, so researchers can screen models without training full language models each time.
Real-world applications the work points toward:
- Textless NLP for languages that lack enough written resources for standard text-based or ASR-based pipelines, which the paper explicitly cites as motivation.
- Discrete speech units as inputs for speech generation, described in the paper as usable by GAN-based synthesis systems that can be conditioned on speaker, pitch, or style tokens.
- Rapid iteration on speech encoders in academic or small-lab settings, since pretraining falls from roughly a week to a day.
- Faster exploration of the design space of speech tokenizers, including larger or syllable/word-like units built on top of SSL models.
Industry relevance. The compute reduction is the headline: 369 GPU hours against 2880 for the original DinoSR and 1984 for HuBERT. The PyTorch-native implementation removes dependence on the unmaintained fairseq library, and the released checkpoints let teams use pretrained units without training from scratch.
Future Directions
- Scaling the language model and data. The paper notes that better performance on larger datasets could potentially be achieved with larger models than the OPT-125M used here, and that text LMs still scale better than SLMs.
- Closing the gap to text models. SpidR changes the intercept, not the slope, of the scaling curves, so the question of what would change the scaling behaviour itself remains open.
- Better units beyond frame-level clustering. The paper situates SpidR alongside work that merges units into larger units closer to syllables or words, and alongside methods like Spin and DC-Spin; combining SpidR with those adaptations is left as an extension.
- Non-linguistic and paralinguistic evaluation. The paper acknowledges that its metrics capture linguistic knowledge only, and points to complementary metrics for speaker, style, and other information.
- Synthesis and hybrid systems. No vocoder is trained in this work; the paper points to prior GAN-based synthesis from discrete units and to hybrid speech-text models as adjacent directions.
Target Audience
Researchers and engineers working on self-supervised speech representation learning, speech tokenization, and textless spoken language modeling, who are already comfortable with masked prediction, teacher-student distillation, and discrete codebooks. It is also relevant to practitioners who care mainly about the practical side: teams that want a fast, stable, open PyTorch pipeline for pretraining speech encoders on limited compute, and anyone evaluating whether ABX and PNMI are trustworthy proxies before committing to downstream language model training.
Authors’ abstract
The parallel advances in language modeling and speech representation learning have raised the prospect of learning language directly from speech without textual intermediates. This requires extracting semantic representations directly from speech. Our contributions are threefold. First, we introduce SpidR, a self-supervised speech representation model that efficiently learns representations with highly accessible phonetic information, which makes it particularly suited for textless spoken language modeling. It is trained on raw waveforms using a masked prediction objective combined with self-distillation and online clustering. The intermediate layers of the student model learn to predict assignments derived from the teacher's intermediate layers. This learning objective stabilizes the online clustering procedure compared to previous approaches, resulting in higher quality codebooks. SpidR outperforms wav2vec 2.0, HuBERT, WavLM, and DinoSR on downstream language modeling benchmarks (sWUGGY, sBLIMP, tSC). Second, we systematically evaluate across models and layers the correlation between speech unit quality (ABX, PNMI) and language modeling performance, validating these metrics as reliable proxies. Finally, SpidR significantly reduces pretraining time compared to HuBERT, requiring only one day of pretraining on 16 GPUs, instead of a week. This speedup is enabled by the pretraining method and an efficient codebase, which allows faster iteration and easier experimentation. We open-source the training code and model checkpoints at https://github.com/facebookresearch/spidr.