Research
Bridging the Language Gap: Synthetic Voice Diversity via Latent Mixup for Equitable Speech Recognition
Overview Research area: Speech recognition (automatic speech recognition, ASR) and multilingual fairness, specifically data augmentation for low-resource languages using latent-space interpolation in

- arXiv
- 2511.20534
- Published
- 2025-11-25
- Authors
- Wesley Bian, Xiaofeng Lin, Guang Cheng
AI summary
Overview
Research area: Speech recognition (automatic speech recognition, ASR) and multilingual fairness, specifically data augmentation for low-resource languages using latent-space interpolation in voice conversion models.
Technical level: Intermediate. The core idea is intuitive (blend two speaker "voices" numerically), but the pipeline relies on a diffusion-based voice conversion model (Diff-HierVC) and a 255-dimensional style encoder, and assumes familiarity with word error rate (WER) and standard ASR training setups.
Scope: The paper proposes and evaluates LatentVoiceMix, a method that performs mixup inside the speaker-timbre latent space of a voice conversion model to synthesize more speaker-diverse training audio, and measures whether this reduces the ASR performance gap between low-resource Wolof and high-resource English.
What This Paper Is About
ASR systems work far better for English than for the world's other 7,000+ languages, largely because English has vastly more transcribed speech data. Collecting more data for low-resource languages is expensive and slow, so the authors instead generate additional synthetic speech that keeps the original words and transcript but sounds like a new, blended speaker. The goal is to improve recognition accuracy on low-resource languages and narrow the performance gap without collecting any new data.
Key Contributions
- A new augmentation method, LatentVoiceMix, that applies mixup inside the latent style-encoder space of a voice conversion model rather than in waveform, spectrogram, or encoder-activation space. The authors state that no prior work had explored mixup within the latent code space of style encoders.
- A complete, reusable augmentation pipeline built on Diff-HierVC: denoising with the
noisereducepackage, extraction and storage of 255-dimensional speaker timbre vectors, random target and mixup-speaker selection, convex combination of timbres, voice conversion, post-denoising, and transcript inheritance from the source audio. - Controlled empirical comparisons against waveform augmentation, SpecAugment-style spectrogram augmentation, and conventional voice conversion augmentation, holding the amount of generated synthetic audio constant (33% more data on AN4; a tripled dataset on Wolof).
- Evidence that latent convexity matters: experiments showing mixup reduces the Wolof–English WER gap, plus a PCA analysis showing synthetic mixup timbres stay closer to the real speaker timbre distribution than waveform-augmented timbres do. The authors describe it as the first fairness-oriented synthetic data generator at the style layer.
Main Findings
- AN4 (low-resource English, NVIDIA NeMo trained from scratch for 50 epochs, +33% data): WER was 0.785 with no augmentation, 0.436 with waveform augmentation, 0.424 with voice conversion augmentation, and 0.339 with mixup — the best of the four.
- Wolof vs. English bias (NeMo, 50 epochs): Training on 8 hours of Wolof plus 24 hours of English produced WER of 0.796 on Wolof versus 0.562 on English, a gap of 0.234. After adding 16 hours of synthetic mixup data to the Wolof side (24h Wolof augmented + 24h English), Wolof WER fell to 0.725 and English to 0.550, narrowing the gap to 0.175.
- Whisper finetuning on Wolof (
whisper-tiny, 8 hours original data, 4 epochs, tripled dataset): WER was 0.283 (none), 0.242 (spectrogram), 0.217 (waveform), 0.215 (voice conversion), and 0.202 (mixup) — lowest overall. - SpeechMOS nuance: On the same Whisper experiment, SpeechMOS (a 1–5 predicted human-perceived quality score) was 2.661 for no augmentation, 2.117 for waveform augmentation, 2.710 for voice conversion augmentation, and 2.243 for mixup; it was not computed for spectrogram augmentation because that method does not generate waveforms. Mixup achieved the best WER but did not have the highest perceived quality score.
- Ablation study (8 hours original Wolof, Whisper finetuning): Full proposed mixup with 16 hours reached WER 0.202. Removing post-denoising gave 0.214; using three speaker timbres gave 0.221; setting source equal to target gave 0.235 at 8 hours and 0.221 at 16 hours. Post-denoising and careful mixup design both mattered.
- Timbre distribution analysis (PCA on the 14 Wolof speakers): Synthetic timbres from mixup sat closer to the distribution of real speaker timbres, while waveform-augmented timbres showed greater variance and tended to fall outside the real distribution — a possible explanation for mixup's better WER.
Methodology in Plain English
The team starts from an existing voice conversion model (Diff-HierVC) that splits audio into two parts: what was said (linguistic content) and who said it (speaker timbre, a 255-number fingerprint of a voice).
- Every audio file in the dataset is denoised with the
noisereducepackage. - The style encoder extracts a 255-dimensional timbre vector for each file, and these vectors are saved to disk for reuse.
- For each training example, the original file supplies the words (the source).
- A different file supplies a target timbre, and a third, separate speaker supplies a mixup timbre.
- The two timbres are blended with the formula t_mixed = λ·t_target + (1 − λ)·t_mixup, where λ is drawn from a Beta(0.5, 0.5) distribution — so the blend is usually close to one of the two voices rather than a 50/50 average.
- The model speaks the original words in this new blended voice.
- The output audio is denoised again to remove artifacts.
- The synthetic clip inherits the transcript of the original source file, since the words never changed.
This produces extra training audio with new-sounding speakers but correct labels, and the authors compare it against time-stretching/pitch-shifting (waveform), masking frequency and time bands (spectrogram, SpecAugment), and ordinary voice conversion, always matching the amount of synthetic data added.
Why This Matters
Impact on research: The paper argues that the location of interpolation matters — doing mixup in a style encoder's latent space, rather than in waveforms or spectrograms, keeps synthetic speakers inside the real speaker-timbre distribution and yields better recognition accuracy. It connects latent interpolation theory (Mixup, Manifold Mixup, MixRep, Latent Filling) with codec-level voice conversion models and frames augmentation as a fairness intervention rather than only a robustness trick.
Real-world applications:
- Speech interfaces and dictation tools for languages with little transcribed audio, such as Wolof and other West African languages.
- Voice assistants, call-center transcription, and accessibility tools (e.g., live captioning) deployed in multilingual regions where users currently get worse service than English speakers.
- Any organization that wants to expand an ASR system to a new language or accent group but cannot afford a large new recording campaign.
- Rapid prototyping and benchmarking of ASR pipelines in resource-constrained settings, such as the small an4 dataset used here.
Industry relevance: The Whisper finetuning experiment targets the common industrial workflow of adapting a large pretrained model to a new language with limited data. Because the method only requires an existing speech corpus and a voice conversion model, it offers a cheaper alternative to data collection for companies and public institutions trying to serve underrepresented linguistic communities.
Future Directions
- Raising synthetic audio quality. Mixup had the lowest WER but a SpeechMOS of 2.243, below no-augmentation (2.661) and voice conversion augmentation (2.710). Closing that perceptual quality gap is an open problem.
- Tuning the mixup distribution. The experiments fix α = 0.5 and β = 0.5 and use two timbres; the ablation shows three timbres performed worse (0.221 vs. 0.202), so the right number of speakers and shape of the Beta distribution remain open questions.
- Extending beyond the three tested corpora. Results cover Wolof (16 hours, 14 speakers), VCTK English (~44 hours, 109 speakers), and an4; whether the gains transfer to other low-resource languages, larger speaker pools, or heavily accented data is not reported.
- Evaluating fairness metrics beyond the WER gap. The paper reports the Wolof–English WER gap (0.234 to 0.175) but does not report per-speaker or per-accent fairness breakdowns, or test on ASR architectures beyond Whisper and NVIDIA NeMo.
Target Audience
Researchers and practitioners working on multilingual or low-resource ASR, speech data augmentation, and voice conversion; machine learning fairness researchers interested in language equity; and engineers who need to adapt pretrained speech models to under-resourced languages without new data collection. Readers with a basic grounding in ASR metrics and neural audio models will get the most from it.
Authors’ abstract
Modern machine learning models for audio tasks often exhibit superior performance on English and other well-resourced languages, primarily due to the abundance of available training data. This disparity leads to an unfair performance gap for low-resource languages, where data collection is both challenging and costly. In this work, we introduce a novel data augmentation technique for speech corpora designed to mitigate this gap. Through comprehensive experiments, we demonstrate that our method significantly improves the performance of automatic speech recognition systems on low-resource languages. Furthermore, we show that our approach outperforms existing augmentation strategies, offering a practical solution for enhancing speech technology in underrepresented linguistic communities.