Research
O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization
Overview Research area: Online speaker diarization — determining who spoke when in a recording, in a streaming (causal) setting. Technical level: Advanced. The paper assumes familiarity with end-to-en

- arXiv
- 2512.15229
- Published
- 2025-12-17
- Authors
- Elio Gruttadauria, Mathieu Fontaine, Jonathan Le Roux, Slim Essid
AI summary
Overview
Research area: Online speaker diarization — determining who spoke when in a recording, in a streaming (causal) setting.
Technical level: Advanced. The paper assumes familiarity with end-to-end neural diarization, encoder-decoder attractor models, transformer decoders, gated recurrent units, permutation-invariant training, and diarization error rate.
Scope in one sentence: The paper introduces O-EENC-SD, an online speaker diarization system built on EEND-EDA with a GRU-based neural clustering "stitching" stage and a new centroid refinement decoder, evaluated on two-speaker conversational telephone speech from CallHome.
What This Paper Is About
Most strong speaker diarization systems are offline and need the whole recording, while online systems typically adapt offline models by splitting audio into chunks — which creates the problem of matching speaker identities across chunks because end-to-end models are trained under permutation-invariant training and give speakers in arbitrary order. Existing online solutions either rely on unsupervised clustering (flexible but hyperparameter-heavy) or on large speaker-tracing buffers (accurate but computationally costly, often using buffers of 100 s or more).
The goal of O-EENC-SD is to solve the cross-chunk speaker permutation problem with a lightweight, fully differentiable, hyperparameter-free neural clustering mechanism, so that online diarization can run efficiently — even on non-overlapping chunks — while remaining competitive with the state of the art in the two-speaker conversational telephone speech domain.
Key Contributions
- An end-to-end online diarization system (O-EENC-SD) based on EEND-EDA, with a novel RNN-based stitching mechanism that performs online neural clustering over speaker attractors across consecutive chunks.
- A new centroid refinement decoder, a transformer decoder that uses cross-attention between stored speaker centroids and the current chunk's attractors (concatenated with a trainable "ghost speaker" embedding) to contextually update centroids before matching.
- A rigorous ablation study over architecture components and over different latency/computation settings, showing the individual and joint benefit of the attractor refinement decoder and the centroid refinement decoder.
- An efficiency analysis demonstrating that O-EENC-SD can operate with a buffer equal to the latency and on independent, non-overlapping chunks, with constant-time per-chunk computation, while remaining competitive on the CallHome test set. Code is released at https://github.com/egruttadauria98/O-EENC-SD.
Main Findings
- Competitive DER at higher latency: Using a 100-s FIFO buffer, O-EENC-SD reaches 9.53% DER at 5-s latency and 9.50% DER at 10-s latency. The best configuration reaches 9.33% DER at 5-s latency when all past frames are included in the buffer.
- High-latency training helps low-latency inference: Tested at 1-s latency with a 100-s buffer, the model trained at 5-s latency reaches 11.96% DER, better than the 12.47% DER of the model trained at 1-s latency. With a much smaller 25-s buffer, the 5-s-latency-trained model reaches 12.14% DER at 1-s latency, still better than the 1-s-latency-trained model with a 100-s buffer.
- Chunk-level quality drives online quality: Average EEND-EDA performance on a single chunk is 11.84% DER when trained at 1-s latency versus 8.89% when trained at 5-s latency, which helps explain the training-latency effect.
- Decoder ablation at 1-s latency: Starting from a base model at 15.38% DER (100-s buffer), adding the attractors decoder gives 15.07%, adding the centroid decoder gives 12.69%, and using both gives 12.47% DER. Joint use raises clustering accuracy from 90% to 93% and improves the per-chunk EEND-EDA DER from 12.68% to 11.84%.
- Low-computation regime: With latency set equal to the buffer, O-EENC-SD at a 10-s buffer reaches 10.75% DER with 97.17% clustering accuracy, outperforming BW-EDA-EEND with L = ∞ (11.82% DER, 10-s buffer). With a 5-s buffer it reaches 13.20% DER and 96.16% accuracy; its base variant reaches 13.94% DER and 93.14% accuracy at 5 s, and 12.93% DER and 94.24% accuracy at 10 s. All variants beat BW-EDA-EEND with L = 1 (16.18% DER, 10-s buffer).
- Small buffers suit the clustering component: Clustering accuracy exceeds 97% for the 10-s buffer model, while models trained with a 50-s buffer average around 93% accuracy.
- Comparison context: In the top section of Table I, EEND-EDA+FW-STB improves from 12.70% DER to 9.08% DER with variable chunk-size training (VCT) and improved speaker-tracing buffer sampling; EEND-GLA-Small+BW-STB reports 9.01% and EEND-GLA-Large+BW-STB reports 9.20%, all at 100–101 s buffers and 1-s latency.
- Small-buffer limitation: At 1-s latency, O-EENC-SD trained with a 5-s latency and 50-s buffer degrades to 19.99% DER with a 5-s buffer and 14.5% with a 10-s buffer, while the same training setup gives 12.54% with a 50-s buffer.
Methodology in Plain English
Audio is cut into chunks processed in sequence. Each chunk goes through an EEND-EDA model, which produces two things: frame-level speaker activity predictions and "attractors" — vector representations of the speakers found in that chunk. Because the model is trained with permutation-invariant training, the order of speakers in those predictions is arbitrary, so the system must decide which speaker in the new chunk corresponds to which speaker seen in earlier chunks.
That matching is done by a recurrent neural network that maintains a "centroid" (hidden state, implemented with gated recurrent units) for each speaker seen so far, plus a common trainable embedding h⁰ that acts as an average speaker and lets the problem be treated as a closed-set, fully differentiable classification task. Each attractor in the new chunk is classified against the known centroids or h⁰ using a cross-entropy loss, and active speakers' centroids are updated with their matched attractors via teacher forcing; non-active speakers' states are left unchanged. This is what makes the approach hyperparameter-free compared with unsupervised clustering.
Two transformer decoders refine the latent speaker representations before matching. The attractor refinement decoder refines the chunk's attractors using only that chunk's frame embeddings (to keep computation light). The new centroid refinement decoder contextually updates the centroids by cross-attending to the chunk's refined attractors concatenated with a trainable "ghost speaker" embedding, giving non-active speakers an extra option to attend to rather than the current chunk's attractors.
Training uses four loss terms: global and chunk-level EEND-EDA losses, the clustering cross-entropy loss, and a diarization loss applied to the stitched output. The total objective is the global EEND-EDA loss plus ten times the chunk-level EEND-EDA loss plus the clustering cross-entropy loss plus the stitching diarization loss. Models are first trained on 10-s segments split into 10 non-overlapping 1-s chunks, then pretrained on simulated telephone conversations built from licence-free CallHome and CallFriend stereo data on TalkBank, and finally fine-tuned on the CallHome development set, with learning rates of 10⁻⁴ and 10⁻⁵ respectively. Latency is controlled by modifying the attention mask of the EEND-EDA encoder, and inference uses first-in-first-out buffering.
Why This Matters
Impact on research: The work positions online neural clustering as a compact alternative to storing raw acoustic context in large buffers or to unsupervised clustering with extensive tuning. Storing one evolving centroid per seen speaker is far cheaper than retaining up to 100 s of frame features, and the paper shows this compact representation can be competitive on CallHome. The centroid refinement decoder and the downstream diarization loss applied to stitched output are concrete, reusable mechanisms for anyone adapting permutation-invariant end-to-end models to streaming use.
Real-world applications:
- Live call-center and telephony analytics, where a 1-s to 10-s latency budget and constant per-chunk cost matter.
- Real-time meeting transcription and note-taking that must attribute speech to participants as the meeting unfolds.
- Edge and on-device diarization, such as smart speakers or wearables, where the paper's O(1) per-chunk setting with a buffer equal to the latency is directly relevant.
- Streaming monitoring of recorded or broadcast telephone speech, including emergency-call and broadcast-media indexing pipelines.
Industry relevance: Deployments rarely have the compute budget for a 100-s buffer with a 1-s hop, which the paper explicitly notes is not feasible for many resource-limited applications. A system whose per-chunk computation does not depend on sequence length, and that can operate on non-overlapping chunks, lowers the barrier to shipping diarization into latency-sensitive and energy-constrained products. The released code further reduces the cost of adoption and reproduction.
Future Directions
- Bridging the gap between high-latency training and low-latency inference, since training at higher latency improves both chunk-level EEND-EDA quality and final DER at test-time latencies as low as 1 s, yet introduces a domain shift.
- Strengthening online neural clustering at low-latency requirements, where the paper's 5-s-buffer and 10-s-buffer results at 1-s latency (19.99% and 14.5% DER) remain far from the larger-buffer results.
- Extending evaluation beyond the two-speaker conversational telephone speech setting on CallHome, including recordings with more speakers, given that unsupervised clustering has been used to handle speaker counts beyond what an end-to-end model permits in one chunk.
- Pursuing real-time speaker diarization on edge devices, the direction the paper explicitly opens by analyzing performance when neither data nor computation is shared between subsequent chunks.
Target Audience
Researchers and practitioners in speech processing and machine learning who work on speaker diarization, streaming or causal speech systems, and end-to-end neural architectures. It is also relevant to engineers building production diarization for telephony, meeting transcription, or edge devices who need to weigh diarization error rate against latency and compute budget. Readers without a background in EEND-EDA, permutation-invariant training, or transformer attention will find the methodology sections demanding despite the clear structure.
Authors’ abstract
We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system provides key advantages over existing methods: a hyperparameter-free solution compared to unsupervised clustering approaches, and a more efficient alternative to current online end-to-end methods, which are computationally costly. We demonstrate that O-EENC-SD is competitive with the state of the art in the two-speaker conversational telephone speech domain, as tested on the CallHome dataset. Our results show that O-EENC-SD provides a great trade-off between DER and complexity, even when working on independent chunks with no overlap, making the system extremely efficient.