Research
VākQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question Answering
Overview Research area: Spoken question answering (SQA) and low-resource language NLP, focused on Telugu factoid QA and the reliability of automatic evaluation. Technical level: Intermediate. The pape

- arXiv
- 2609.19879
- Published
- 2026-09-17
- Authors
- Bhavana Akkiraju, Ravi Sastry Kolluru, Sri Charan D, Srihari Bandarupalli, Santosh Kesiraju, Anil Vuppala
AI summary
Overview
Research area: Spoken question answering (SQA) and low-resource language NLP, focused on Telugu factoid QA and the reliability of automatic evaluation.
Technical level: Intermediate. The paper combines benchmark construction, speech processing pipelines (VAD, ASR, MT), and LLM evaluation methodology; readers should be comfortable with metrics such as WER, BLEU, EM/F1, and correlation coefficients.
Scope in one sentence: The paper releases VākQA, a 2,001-question Telugu spoken factoid QA benchmark with 2.53 hours of audio and bilingual transcriptions, and uses it to measure how input modality, input language, model tier, and cascaded ASR→MT pipelines affect answer correctness, while validating how well automatic judges agree with human ratings.
What This Paper Is About
Spoken question answering has advanced mainly for high-resource languages, and no spoken QA benchmark existed for Telugu. Beyond building the dataset, the authors ask whether automatic evaluation can even be trusted in this setting, since exact-match metrics break down when ASR errors and bilingual answers produce correct responses that differ in surface form from the reference. The goal is to create the first Telugu spoken factoid QA benchmark and use it to quantify both model behavior and judge reliability.
Key Contributions
- The benchmark itself: VākQA is released publicly (https://hf.co/datasets/Bhavanaakkiraju/VakQA) and described as the first benchmark for spoken QA in Telugu across six domains, containing spoken audio and bilingual (Telugu and English) transcriptions.
- Evaluation reliability study: The authors analyze LLM-based automatic evaluation for Telugu QA and show reliability depends strongly on the judge model, with Gemini-as-a-judge exhibiting non-uniform strictness.
- Systematic benchmark study: They isolate the effects of input language, input modality, model size in parameters, and proprietary versus open-weights tier, and additionally analyze cascaded ASR→MT error compounding and domain-wise performance.
- Human-verified data construction pipeline: A semi-automatic pipeline spanning VAD and chunking, ASR and merging, QA extraction and audio–text alignment, followed by human verification and manual translation into English.
Main Findings
- Gemini is the most reliable judge but is still imperfect: Against average human judgment, Gemini-as-a-judge reaches Spearman ρ = 0.86 and Kendall τ = 0.77, the highest of all methods, with ME = -0.28 and MAE = 0.46. It is slightly stricter on average with the narrowest limits of agreement (LoA: -1.28 to 0.72), but its behavior is non-uniform: more lenient for low-quality answers and stricter for high-quality ones.
- Open-weight judges systematically penalize correct Telugu answers: Gemma-3-12B reaches ρ = 0.81 and τ = 0.71 with ME = 0.34 and LoA -1.15 to 1.83; Gemma-3-27B reaches ρ = 0.80 and τ = 0.70 with ME = -0.07 and LoA -1.57 to 1.43; Gemma-3-4B is weakest (ρ = 0.57, τ = 0.50) with the widest spread (LoA: -2.31 to 2.93). Switching the judge from Gemini to Gemma-12B made 46.23% of candidate answers score worse, 21.14% better, and left 32.63% unchanged.
- Lexical and embedding metrics correlate poorly with humans: EM (ρ = 0.37, τ = 0.33), Avg. F1 (ρ = 0.49, τ = 0.43), and BLASER-2.0 (ρ = 0.36, τ = 0.27) all fall below 0.5 in Spearman correlation. In example E1, a more detailed but correct answer receives 0 from EM and F1, while BLASER-2.0 yields 2.43.
- Proprietary models lead consistently: With oracle Telugu text, Gemini scores 3.63 (1.71) on the 1–5 scale, versus Gemma-27B 2.55 (1.79), Gemma-12B 2.01 (1.60), Sarvam-m 1.98 (1.60), and Gemma-4B 1.43 (1.11); Llama-3.1-8B (1.41), Hex-1 (1.33), and Qwen-3-4B (1.17) score near or below 1.5.
- Telugu input preserves scope better than English translation: Oracle Telugu text (O1) gives Gemini 3.63 (1.71) versus 3.52 (1.74) for oracle English (O2). Pairwise, English degrades 18.8% of questions, improves 16.5%, and leaves 64.7% unchanged. The example given is a Telugu possessive pronoun "mana" (our) that implicitly scopes the question to India, becoming ambiguous in English and causing "Sputnik" instead of the correct "Aryabhata."
- Speech input introduces phonetic confusions: Speech (S1) scores 3.28 (1.84) versus 3.63 (1.71) for oracle text. Pairwise, speech degrades 21.2% of questions, improves 13.1%, and leaves 65.7% unchanged. In example E4, the model confuses "paṁdu" (fruit) with the phonetically closer "paṁḍuga" (festival), answering Bathukamma instead of mango.
- ASR→MT errors compound progressively: ASR alone (A1, Seamless FT) degrades 11.9% of questions versus oracle text and improves only 4.8%. Cascaded pipelines score lower still: O1→Seamless MT (O3) gives Gemini 2.74 (1.83) versus 3.63 for O1, and full ASR+MT cascades (M1–M4) range from 2.44 to 2.80 for Gemini.
- Component-level ASR and MT quality: Seamless FT ASR achieves 30.25 WER and 10.12 CER; IndicWhisper ASR achieves 35.23 WER and 22.38 CER. Translation BLEU/ChrF++ drops from 37.22/60.12 (oracle text → Seamless MT) and 36.91/61.55 (oracle text → Indic MT) to 27.95/52.03 for IndicWhisper → Seamless MT.
- Domain matters, and Culture is most fragile: With oracle Telugu text, Gemini scores 3.54–3.76 across domains, with Culture strongest (3.76) and Science and Politics slightly weaker (3.54). Under translation, Culture drops the most (3.76 → 2.42), while Geography becomes strongest in O3 (2.94). For open-weight models, Science is generally easiest and Culture consistently hardest; Gemma-27B ranges 2.31–2.80 in O1 and Gemma-12B 1.63–2.27.
- Human annotation was reliable: Five native Telugu speakers rated 400 (question, reference answer, candidate answer) triplets from 100 questions and four QA models on a 1–5 scale; after excluding 20 outlier items with the full 1–5 rating spread, Krippendorff's α reached 0.836.
Methodology in Plain English
The team gathered Telugu YouTube videos featuring quiz-style and multiple-choice question content, chosen so each spoken question and answer sat in separable audio segments. Sources were curated across six domains: General Knowledge, Science, Geography, History, Politics, and Culture.
To turn raw video into data, they detected non-silent regions with Pyannote VAD and cut them into 7-second chunks with 2-second overlap, keeping original timestamps. Each chunk was transcribed with a fine-tuned Seamless-large-v2 Telugu ASR model trained on roughly 900 hours of data drawn from IndicVoices, Kathbath, Google FLEURS, SyspIn TTS, and IndicTTS. The chunk transcripts were merged into a single passage, and Gemini was used to extract question–answer pairs verbatim. For finer alignment, word-level timestamps were computed separately using Whisper-timestamped with IndicWhisper, and the extracted QA text was matched to this word-level output via fuzzy string matching at a Levenshtein ratio of at least 85%, allowing precise adjustment of QA audio boundaries. Five annotators then checked that transcripts matched their audio and manually translated everything into English.
For evaluation, they gathered human ratings as the gold standard, then compared those ratings against Exact Match, token-level F1, BLASER-2.0, and four LLM judges (Gemini and Gemma-3 4B/12B/27B) using Spearman's ρ, Kendall's τ, mean error, and mean absolute error. They then benchmarked QA models under controlled input conditions: direct speech, Telugu ASR text, cascaded ASR→MT English (two ASR systems and two MT systems, giving four configurations), and oracle ground-truth text in both languages. Testing the same model on the same questions in both languages separates a model that lacks knowledge from one that has knowledge but cannot access it in a given language. QA models tested were Gemini-2.5-Flash plus the open-weight Gemma-3 family (4B, 12B, 27B), Llama-3.1, Hex-1, Sarvam-m, and Qwen-3-4B.
Why This Matters
Impact on research: This is described as the first spoken QA benchmark for Telugu, and it is unusual in benchmarking input language, modality, cascaded ASR→MT errors, and judge reliability together. It also challenges a common practice: most Indic QA datasets rely on automatic metrics such as EM and F1, and the paper shows those metrics correlate with human judgment below 0.5 on this data. The finding that smaller open-weight judges fail to recognize semantic equivalence in Telugu is a direct warning for evaluation practice in low-resource settings.
Real-world applications:
- Voice-based information access and voice assistants for Telugu speakers, where users ask spoken factoid questions and need direct answers.
- Government, health, and agricultural voice helplines serving Telugu-speaking populations, where query phrasing and domain terminology vary widely.
- Educational and quiz-style content delivery, matching the YouTube quiz format the dataset was built from.
- Speech-to-text and translation product development for Indic languages, since the paper quantifies where ASR and MT errors break downstream tasks.
Industry relevance: The results give a concrete ordering of practical design choices: matching input language to the user's language rather than translating, avoiding cascaded ASR→MT pipelines where a speech-capable model is available, and treating judge model selection as a substantive engineering decision rather than a formality. The absence of any open-weight Telugu speech LLM at the time of writing, noted explicitly by the authors, marks a clear product and research gap.
Future Directions
- Improve evaluation for Telugu specifically: The authors note that model-based metrics such as Bert Matching, BERTScore, BLEURT, and ORCA cannot be directly applied to Telugu without language-specific fine-tuning, which they left outside the scope of this work.
- Reduce judge non-uniformity: Gemini-as-a-judge's non-uniform strictness, being more lenient at low scores and stricter at high scores, limits fine-grained comparisons; the authors treat this as a limitation to be addressed.
- Handle translation scope ambiguity: YouTube-sourced audio requires faithful transcript-based translation rather than clarified rewrites, which introduces scope ambiguity, as when a Telugu possessive pronoun becomes ambiguous in English. Richer translation approaches or source-language-native models could be explored.
- Assess transliterated reference answers: Some Science and Geography reference answers use English transliterations, and the paper does not assess whether judges score these consistently against native Telugu equivalents.
- Build open-weight Telugu speech LLMs: The authors state that, to the best of their knowledge, no open-weight Telugu speech LLM is currently available, which is why open-weight models were not evaluated in the direct speech setting.
Target Audience
Researchers and engineers working on spoken question answering, Indic-language NLP, and low-resource speech processing will get the most from this paper, particularly those building or evaluating datasets in languages without existing speech QA resources. Evaluation researchers interested in LLM-as-a-judge reliability across languages and tasks will find the reliability comparison directly useful. Product teams building voice interfaces, ASR, or machine translation for Telugu and related Indic languages can use the pipeline error analysis to prioritize engineering effort. Speech and language technology students will find the semi-automatic, human-verified construction pipeline a practical template for building similar benchmarks.
Authors’ abstract
Question answering has advanced rapidly with large language models, but predominantly for high-resource languages, in both text and spoken settings. Spoken question answering (SQA) benchmark for Telugu remains unexplored, and the reliability of automatic evaluation in this setting remains unquantified. We introduce VākQA, a Telugu SQA benchmark of 2,001 factoid question-answer pairs across six domains, with 2.53 hours of speech audio, bilingual transcriptions, and human-verified reference answers. We first validate evaluation methods against human judgements: Gemini-as-a-judge best approximates human ratings but is non-uniformly strict, while open-weight judges systematically penalize correct Telugu answers that differ in surface form from the reference. Using this validated setup, we benchmark proprietary and open-weight models across input modality, language, and domain. We observe that Telugu phrasing retains cultural specificity that is lost in translation, speech input introduces phonetic confusions that alter question meaning, and cascaded ASR-MT errors compound progressively. VākQA is publicly released.