Skip to content
AI.info

The Pulse

Humyn Labs Puts 23 ASR Models Against Overlapping Speech

Humyn Labs' BRIDGE ASR 2.0 benchmark evaluates 23 commercial speech-recognition models on spontaneous, overlapping conversations across 21 languages and six metrics.

Humyn Labs Puts 23 ASR Models Against Overlapping Speech

AI.info Team ·

Humyn Labs has released BRIDGE ASR 2.0, a benchmark that puts 23 commercial speech-recognition models through the kind of audio that clean laboratory tests often avoid: spontaneous two-person conversations, background noise, code-switching, long pauses and people speaking over one another.

The benchmark's citation identifies it as a September 2026 report. It evaluates models on real conversational recordings rather than single-speaker dictation. Its leaderboard ranks systems across six measures, including word error rate, semantic similarity, code-switching and word information loss.

The benchmark is designed to test whether speech systems can preserve useful information when people do not speak in neat, isolated turns.

BRIDGE Uses Conversations Instead of Dictation

Every recording in the benchmark features two speakers and real cross-talk. Humyn Labs says the conversations run for 10 to 15 minutes and were captured at 44–48 kHz across everyday acoustic environments. The company presents the dataset as a test of streaming recognition and speaker separation, two tasks that can fail when speech overlaps.

The sample covers 18 Indic languages alongside Latin American Spanish, Brazilian Portuguese and Vietnamese. Humyn Labs tags the recordings across seven cohort dimensions, allowing users to compare results by language, region, silence, conversational behavior and other conditions rather than relying on one overall average.

The six-metric stack combines Word Error Rate, Character Error Rate, Semantic Similarity, Code-Switch F1, Phoneme-Informed Error Rate and Word Information Lost. The company also separates substitutions, deletions and insertions. That distinction matters: a model can mishear a word, drop an entire segment or add speech that was never recorded, while producing similar headline error rates.

ElevenLabs Leads the Indic Results

ElevenLabs Scribe v2 takes the top position in the report’s Indic-language leaderboard, with a 10.99% loanword-adjusted word error rate. Soniox stt-async-v4 follows at 16.32%, a gap of 5.3 percentage points. Humyn Labs says Scribe v2 ranks first in 14 of the 18 Indic languages tested.

The spread is far wider lower down the table. The six weakest Indic-language models record loanword-adjusted error rates between 70.77% and 88.03%. OpenAI’s GPT-4o and GPT-4o mini score near the bottom of the displayed ranking, while AssemblyAI Universal, Universal 2 and Universal 3 Pro sit below them.

Results are closer in the Latin-script group. Scribe v2 records a 4.83% WER across Spanish, Portuguese and Vietnamese, compared with 8.14% for Microsoft MAI Transcribe 1.5. The full range runs from 4.83% to 22.10%, which leaves latency, pricing and language coverage as meaningful deployment factors when several systems produce usable transcripts.

Long Silence Causes More Damage Than Fast Turn-Taking

One of BRIDGE’s clearest findings concerns silence. After controlling for language and model, calls with gaps longer than 20 seconds show an average 5.05-point loss in loanword-adjusted accuracy. Rapid turn-taking produces a smaller average penalty of 1.33 points.

The pattern appears across most of the viable models tested. Google Chirp 3 loses 13.22 points with long silences compared with 6.69 points from rapid speech. Gemini 2.5 Pro loses 10.87 points from long gaps and 2.16 points from speed, while Scribe v2 loses 7.50 points and 0.79 points respectively.

Humyn Labs attributes the likely failure to segmentation rather than basic recognition. Long pauses are concentrated in Marwari, which is already the hardest language in the Indic set. Without controlling for language, the apparent silence penalty rises to 16.33 points, overstating the effect of silence alone.

Code-Switching Exposes a Second Quality Gap

BRIDGE also measures whether systems preserve English words embedded in Indic-language speech. Humyn Labs argues that ordinary WER can hide an important difference: a model may retain the English term but write it in a local script, or remove the term entirely.

ElevenLabs Scribe v2 leads the code-switching measure with a CS F1 score of 0.897, followed by Gemini 3 Pro at 0.891 and Gemini 3 Flash at 0.886. AssemblyAI Universal records 0.157, placing it in the report’s zero-code-switching category. Humyn Labs says 15 models score at least 0.7, while seven fall between 0.2 and 0.7.

The benchmark separates script mismatch from missing words. Sarvam saarika v2.5, Sarvam saaras v3, AWS Transcribe, Google Chirp 3 and Soniox show script penalties while maintaining roughly 90–95% code-switch precision. Humyn Labs describes that problem as potentially recoverable through transliteration-aware matching; a word that disappears altogether cannot be recovered in the same way.

A Benchmark Built for Routing Decisions

BRIDGE’s results challenge the usefulness of a single universal ranking. Humyn Labs reports that file differences explain 36.1% of the variation in loanword-adjusted error among viable Indic models, compared with 25.0% for language and 14.6% for the model itself. The median pairwise correlation between model scores is 0.42, suggesting that systems often struggle with different recordings.

The evaluation set, including audio files, reference transcripts, speaker metadata, cohort labels and scoring scripts, is available through the benchmark page. Humyn Labs says additional languages and an overlap-focused corpus are in preparation. For developers building call transcription, voice interfaces or speech systems for physical devices, the immediate value is not simply identifying a winner; it is seeing whether a model drops speech, invents it, mishandles silence or loses the English terms that downstream software needs.

Source

Humyn Labs

Explore

More articles