Research
Make It Hard to Hear, Easy to Learn: Long-Form Bengali ASR and Speaker Diarization via Extreme Augmentation and Perfect Alignment
Overview Research area: Automatic Speech Recognition (ASR) and speaker diarization for Bengali, with a focus on long-form, multi-speaker audio in a low-resource setting (cs.SD). Technical level: Inter

- arXiv
- 2602.23070
- Published
- 2026-02-26
- Authors
- Sanjid Hasan, Risalat Labib, A H M Fuad, Bayazid Hasan
AI summary
Overview
Research area: Automatic Speech Recognition (ASR) and speaker diarization for Bengali, with a focus on long-form, multi-speaker audio in a low-resource setting (cs.SD).
Technical level: Intermediate. The paper is a competition system description, so it assumes familiarity with model names, VAD-based chunking, and error metrics (WER, DER, RTF), but its central arguments are stated plainly and can be followed without deep mathematical background.
Scope: A full account of Team Villagers' dual ASR and diarization pipeline submitted to the DL Sprint 4.0 competition, the failures that shaped it, and the 882-hour Bengali dataset (Lipi-Ghor-882) built to support it.
What This Paper Is About
Bengali speech technology has advanced for short clips, but long-duration audio and reliable speaker diarization (deciding who spoke when) remain unsolved, largely because there is no large, temporally aligned multi-speaker Bengali dataset. The authors set out to build such a resource and to find a pipeline that works on it, using their entry in the DL Sprint 4.0 (BUET CSE FEST 2026) competition, where models were graded on a 22-hour hidden test set of long-form audio, as the proving ground. Their headline lesson is that engineering effort paid off in opposite ways for the two tasks: corrupting the training audio helped recognition, while algorithmic post-processing of model output, not retraining, helped diarization.
Key Contributions
- Lipi-Ghor-882: An 882-hour multi-speaker Bengali dataset assembled from YouTube with
yt-dlp, spanning 1,019 videos, 596 unique channels, more than 150 categorized domains, with roughly 856 annotated hours and speaker boundaries derived from the Pyannote API, stored in SSTT annotation format. - A training recipe for long-form Bengali ASR: Targeted fine-tuning of Whisper-Medium on a small, perfectly aligned subset in which 20 percent of the audio was artificially corrupted with noise and reverberation, rather than scaling raw data or ensembling.
- A diarization result that contradicts the usual assumption: Global state-of-the-art models and retraining failed; performance came from strict heuristic post-processing of Pyannote outputs, formalized as Algorithm 1 with fixed thresholds (merge 3.79 s, gap 0.17 s, segment 0.75 s, speaker 9.0 s).
- An efficiency-optimized dual pipeline: CTranslate2 plus
faster-whisperon dual T4 GPUs cut Whisper's inference on the 22-hour set from 4 hours to 26 minutes, an RTF of roughly 0.019.
Main Findings
- Raw scaling and ensembling did not work for ASR. ROVER ensembling of the best predictions produced a public WER of 0.33770 and private WER of 0.35762, worse than the standalone fine-tuned faster-whisper model at 0.30842 and 0.31070.
- Corrupted training data was the single most effective ASR lever. Training on a perfectly aligned subset with 20 percent of audio artificially degraded by noise and reverberation forced reliance on phonetic features instead of acoustic memorization. The gain plateaued at a ceiling set by annotation quality in the broader dataset.
- Inference engineering delivered the biggest speed win. Moving Whisper-Medium to CTranslate2 and
faster-whisperwith parallel processing on dual T4s cut runtime from 4 hours to 26 minutes (RTF about 0.019).faster-whisperbeatwhisperX, and default VAD parameters had to be adjusted manually because they were truncating the first words of speech chunks. Silero VAD gave more reliable boundaries than Pyannote VAD. - Base model trade-offs. Moonshine ran in 5 minutes but transcribed poorly; Hishab-Titu-Bn-Conformer-Large ran in 17 minutes with competitive accuracy; Wav2Vec2 took 2 hours with chunking at moderate accuracy; Whisper-Medium (Tugstugi) took 4 hours but had the highest fidelity (public WER 0.44443, private 0.45064).
- Contextual biasing backfired over time. It improved scores early but degraded catastrophically past a training threshold and failed on out-of-distribution data (public WER 0.35578, private 0.38231).
- Frozen-decoder PEFT gave only small gains. The adapter variant reached public WER 0.31073 and private 0.32145, an improvement over the base model but not the final strategy.
- Demucs is a trade-off, not a free win. Removing music and noise lowered ASR error (public WER 0.30354, private 0.30612) but raised the RTF from 0.019 to 0.068 and inflated inference time beyond acceptable limits, so it was dropped. In diarization, Demucs actually made scores worse.
- Global diarization SOTA underperformed here. Diarizen, described as a leading open-source SOTA, did poorly on both leaderboards (public DER 0.28077, private 0.27893). Adding a Bengali-trained WavLM to Diarizen, or using VBx with a Pyannote pipeline, changed nothing.
- Retraining the diarization model was a dead end. Training the segmentation model on muffled audio for 20 epochs produced negligible DER improvement.
- Leaderboard behavior differed between splits. Pyannote 3.1 topped the public leaderboard (DER 0.20283) but was weaker privately (0.28145), whereas Pyannote Community-1 was more robust privately (0.26640, public 0.24522). ECAPA-TDNN reached public DER 0.25322 and private 0.31640.
- Rigid post-processing was the diarization driver. The custom algorithm forces inter-speaker gaps, renames speakers serially by first appearance, resolves overlaps, merges same-speaker micro-segments, filters short segments, and prunes speakers below a total speaking duration.
- Compute was a hard constraint. Training ran on a single L40S GPU with 48 GB VRAM, limited to 5 free hours per month from Lightning AI; all inference was evaluated on 2x T4 GPUs.
Methodology in Plain English
The team treated the two tasks as separate pipelines rather than one joint model, which let them tune each side independently.
For recognition, they started by racing pre-trained Bengali models against each other on a 22-hour validation set, measuring both accuracy and speed. Whisper-Medium was the most accurate but far too slow, so they converted it to a faster runtime format, ran it on two GPUs in parallel, and added a voice-activity detector to split long audio into chunks. Then they tried a series of training ideas, most of which failed: freezing the decoder, adding contextual hints, stripping background music with a source-separation model, and combining outputs with ROVER. The one that worked was deliberately degrading the training audio — adding noise and reverberation to 20 percent of a small, cleanly aligned set — so the model had to learn the sound of speech rather than the sound of a particular recording.
For diarization, they tested several off-the-shelf systems and found the well-regarded ones disappointing on this data. Fine-tuning did not help either. Their solution was therefore procedural: take the output of Pyannote Community-1 and rewrite the timeline with a fixed set of rules — enforce a minimum gap between different speakers, merge fragments from the same speaker that are close together, delete very short segments, and drop speakers who barely talk. The specific thresholds were chosen values of 3.79 s for merging, 0.17 s for the inter-speaker gap, 0.75 s for minimum segment length, and 9.0 s for minimum total speaker duration.
Finally, they built Lipi-Ghor-882 to fill the data gap they kept running into, harvesting audio and transcripts with yt-dlp and using the Pyannote API to mark speaker boundaries.
Why This Matters
Impact on research: The paper is a counterexample to two common defaults in low-resource speech work — that more data and more ensembling help, and that a bigger or newer pretrained model will beat a well-tuned pipeline. It also contributes a large joint ASR-and-diarization resource for a language where such data has been scarce, and it reports negative results (contextual biasing decay, ROVER regression, failed diarization retraining) that are rarely published.
Real-world applications:
- Transcribing and indexing Bengali broadcast media, interviews, and panel discussions, where knowing who spoke when is as important as the words.
- Subtitle and caption generation for long-form Bengali video, where the reported RTF of about 0.019 makes batch processing on modest hardware feasible.
- Meeting and call analytics for Bengali-speaking organizations that need speaker-attributed transcripts.
- Archival and legal transcription of multi-speaker Bengali recordings, where automatic segmentation with enforced speaker boundaries reduces manual labeling effort.
Industry relevance: The paper's engineering findings — CTranslate2 conversion, faster-whisper over whisperX, manual VAD parameter tuning, and the observation that Demucs cleanup can cost more in latency than it returns in accuracy — translate directly into production deployment decisions. The demonstration that heuristic post-processing can outperform model work also suggests where teams should allocate effort when annotation budgets are tight.
Future Directions
- Raising the annotation-quality ceiling. The authors state that ASR training gains stopped at a ceiling caused by poor annotation quality in the broader dataset, so improving or filtering the roughly 856 annotated hours of Lipi-Ghor-882 is the obvious next lever.
- Better speaker separation under overlap. The diarization results suggest that long-form multi-speaker audio remains difficult for current architectures; overlap-heavy conversational speech is the hardest unresolved case mentioned.
- End-to-end joint modeling. The team deliberately kept ASR and diarization separate for computational efficiency; whether a joint model could beat the split pipeline on this dataset is untested here.
- Larger-scale tuning under realistic compute budgets. The reported constraint of 5 GPU hours per month on one L40S left "extensive experimentation and hyperparameter tuning" undone, and the authors believe more compute hours could meaningfully improve performance.
Target Audience
Speech researchers and engineers working on low-resource languages, especially those dealing with long-form or multi-speaker audio; competition participants and practitioners building ASR-plus-diarization pipelines under tight latency and hardware budgets; and dataset curators looking for a documented example of assembling a large Bengali speech corpus with yt-dlp and the Pyannote API. Readers seeking new model architectures will find less here than readers looking for practical engineering judgment and honestly reported negative results.
Authors’ abstract
Although Automatic Speech Recognition (ASR) in Bengali has seen significant progress, processing long-duration audio and performing robust speaker diarization remain critical research gaps. To address the severe scarcity of joint ASR and diarization resources for this language, we introduce Lipi-Ghor-882, a comprehensive 882-hour multi-speaker Bengali dataset. In this paper, detailing our submission to the DL Sprint 4.0 competition, we systematically evaluate various architectures and approaches for long-form Bengali speech. For ASR, we demonstrate that raw data scaling is ineffective; instead, targeted fine-tuning utilizing perfectly aligned annotations paired with synthetic acoustic degradation (noise and reverberation) emerges as the singular most effective approach. Conversely, for speaker diarization, we observed that global open-source state-of-the-art models (such as Diarizen) performed surprisingly poorly on this complex dataset. Extensive model retraining yielded negligible improvements; instead, strategic, heuristic post-processing of baseline model outputs proved to be the primary driver for increasing accuracy. Ultimately, this work outlines a highly optimized dual pipeline achieving a $\sim$0.019 Real-Time Factor (RTF), establishing a practical, empirically backed benchmark for low-resource, long-form speech processing.