Skip to content
AI.info

Research

SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling

Overview Research area: Spoken language understanding (SLU), specifically slot filling with speech-based large language models (speechLLMs), evaluated on call-center data. Technical level: Intermediat

SpeechLLMs for Large-scale Contextualized Zero-shot Slot Filling
arXiv
2510.15851
Published
2025-10-17
Authors
Kadri Hacioglu, Manjunath K E, Andreas Stolcke

AI summary

Overview

Research area: Spoken language understanding (SLU), specifically slot filling with speech-based large language models (speechLLMs), evaluated on call-center data.

Technical level: Intermediate. The paper assumes familiarity with ASR/NLU pipelines, parameter-efficient fine-tuning (LoRA/QLoRA), and multimodal model architecture, but the experimental narrative is readable.

Scope: The paper builds a large-scale call-center slot-filling dataset and instruction-based training recipe, then systematically studies modality adapters and training strategies to close the gap between a composite speechLLM and a text-based upper bound, while quantifying out-of-domain and zero-shot weaknesses.

What This Paper Is About

Slot filling asks a system to extract structured information — slot types and their values — from what a user says, and it is traditionally done by a cascade of speech recognition followed by one or more text NLU components. The authors ask whether a single speechLLM, which couples a speech encoder to a text LLM through a modality adapter, can do this at industrial scale while generalizing zero-shot to slot labels it has never seen in training. Their goal is to establish an empirical upper bound for the task, measure the gaps in performance, robustness, and generalization, and then close those gaps through better data, architecture, and training.

Key Contributions

  1. A new industrial-scale slot-filling dataset and annotation recipe. CallCenter-A contains approximately 31K scripted call-center calls, almost 1M turns, and 2.1K hours of speech across banking, telecommunication, insurance, and retail. It was annotated with GPT-4o, deliberately without priming the LLM for any fixed slot label set, using the complete call as context and annotating turn-by-turn for real-world entities, events, dates, times, and numerals while excluding abstract notions.

  2. An instruction-based training data construction strategy. Training triplets of (audio, instruction, response) are built with turn-by-turn slot filling with and without context, where context size T is randomized in the range 0 ≤ T ≤ 3; prompts optionally query specific slots with a randomized number of distractors S in the range 1 ≤ S ≤ 5; and each case samples from a set of 10 prompts.

  3. A systematic comparison of modality adapters and training strategies. Four adapters of increasing parametric complexity (CNN, linear, MLP with SwiGLU, 2-layer transformer with 8 attention heads), each designed for 8x subsampling, are compared with and without LoRA adaptation of the LLM, followed by a comparison of single-stage joint training against three multistage strategies (Multistage-A, -B, -C) and multitask data expansion.

  4. An explicit distinction between composite and foundational speechLLMs. The authors separate their task-specific "composite" model (two unimodal foundation models fine-tuned bimodally) from Qwen2-Audio, which they treat as a foundational speechLLM trained at scale on diverse data with pretraining, instruction-supervised training, and DPO, and they fine-tune both.

Main Findings

  • The text-based fine-tuned LLM is the empirical upper bound. On CallCenter-A, the fine-tuned LLM (FT-LLM, Llama 3.2 1B fine-tuned on the textual part of CallCenter-A) reaches precision/recall/F1 of 0.8523/0.9901/0.9160. A vanilla LLM without fine-tuning reaches 0.0436/0.7259/0.0823, illustrating its inability to perform this structured task zero-shot. The cascaded ASR-plus-NLU system (Whisper | FT-LLM, with the fine-tuned Whisper operating at 14.20% word error rate) reaches 0.7006/0.8982/0.7872, and the baseline speechLLM with a CNN adapter and frozen encoder and LLM reaches only 0.3892/0.4151/0.4017.

  • Adapter complexity matters, and so does adapter design, not just size. With frozen foundation models, F1 rises from 0.4017 for the CNN adapter (3.41M parameters), to 0.6581 for the linear adapter (8.39M), to 0.6970 for the MLP adapter (20.98M), to 0.7420 for the transformer adapter (67.16M). A widened CNN variants, CNN-XL (68.17M parameters), reached only 0.6994 F1 despite matching parameter scale, showing that modeling capability, not parameter count alone, drives the result.

  • Fine-tuning the LLM, rather than freezing it, is critical. Enabling LoRA for the LLM raised F1 for the CNN adapter by +0.2632 (to 0.6649), for the linear adapter by +0.0490 (to 0.7071), and for the MLP adapter by +0.0707 (to 0.7677). The transformer adapter diverged at default settings, dropping to 0.5504 F1 (Δ −0.1916); reducing the learning rate brought it to 0.7242 F1 (Δ −0.0178) and increasing the warm-up period brought it to 0.7479 F1 (Δ +0.0059). The authors continued subsequent LoRA experiments with the MLP adapter.

  • Multistage training improves over single-stage joint training. On CallCenter-A, single-stage joint training gives 0.6553/0.9268/0.7677; Multistage-A gives 0.6722/0.9297/0.7803; Multistage-B gives 0.6653/0.9448/0.7808; and Multistage-C gives 0.6842/0.9391/0.7916. Multistage-C also showed the fastest convergence and lowest loss in the learning curves.

  • Multitask data expansion helps both in-domain and out-of-domain. Adding AST, SIT, and SQIT auxiliary tasks ("SpeechLLM++") improved CallCenter-A F1 from 0.7677 to 0.7797 (precision/recall/F1 0.6632/0.9458/0.7797), CallCenter-B ID F1 from 0.5140 to 0.5496 (0.4240/0.7810/0.5496), and CallCenter-B OOD F1 from 0.3119 to 0.3561 (0.2645/0.5448/0.3561).

  • The cascaded system still generalizes better than the composite speechLLM. On CallCenter-B, the FT-LLM reaches 0.8360/0.9530/0.8907 on in-domain slots and 0.5439/0.8474/0.6626 on out-of-domain slots; Whisper | FT-LLM reaches 0.5909/0.7067/0.6436 and 0.4280/0.6278/0.5090; SpeechLLM reaches only 0.3964/0.7306/0.5140 and 0.2386/0.4500/0.3119. The authors conclude that modality alignment and LLM fine-tuning may not be sufficient for the joint model to beat the sequential performance of unimodal models without training at scale (more tokens and parameters). CallCenter-B contains labels that overlap only 48% with CallCenter-A, leaving 52% entirely new labels.

  • A foundational speechLLM fine-tuned on the task outperforms the composite model but remains below the text upper bound. Fine-tuned Qwen2-Audio reaches 0.7710/0.9799/0.8630 on CallCenter-A, 0.6708/0.8697/0.7574 on CallCenter-B ID, and 0.4507/0.7619/0.5664 on CallCenter-B OOD. The same model with prompt engineering only (PE-Qwen2-Audio) collapses to 0.1757/0.6334/0.2751, 0.1059/0.6178/0.1807, and 0.0116/0.6034/0.0227, illustrating that fine-tuning is necessary for the specific task.

  • Training instabilities increase with adapter complexity. Gradient explosion and divergent learning curves, known from prior work when training adapters and LoRA from scratch, first emerged in this study with the transformer adapter. Standard remedies (warm-up, gradient clipping, initialization, layer-norm placement, smaller architecture) were already in place, but the defaults were suboptimal for that adapter.

  • Reported numbers are representative, not statistically characterized. The paper states that a single performance estimate per configuration is reported due to compute constraints, that performance fluctuates by ±2 percentage points depending on checkpoint selection, and that the saved best-loss model may not correspond to the best target-metric model.

Methodology in Plain English

The authors adopt a speechLLM architecture modeled closely on SpeechVerse with three parts: a pretrained speech encoder, a modality adapter, and a pretrained decoder-only LLM. Audio features and text embeddings are combined along the time/position dimension and fed to the LLM, which generates a text response conditioned on both.

They start with a CNN adapter and frozen encoder and LLM, trained only on the slot-filling task, then progressively swap in more expressive adapters (linear projector, two-layer MLP with SwiGLU, 2-layer transformer) while keeping the same 8x subsampling and projection to the LLM input dimension. Next they enable LoRA on the LLM to align input representations with output generation, not just align the modalities to each other. They then test three multistage recipes: Multistage-A fine-tunes the Whisper encoder on (speech, text) pairs and the LLM on (instruction, response) pairs before joint adapter-plus-LLM training; Multistage-B fine-tunes the adapter on a "continuation" task built by prompting the LLM to complete transcripts, then converts those to (audio, continuation) pairs to align spoken and textual inputs; Multistage-C fine-tunes the adapter on automatic speech transcription before joint adaptation. Finally, they expand the task-specific data with auxiliary tasks (AST, SIT, SQIT) to improve alignment and reduce overfitting, and they compare their composite model against a fine-tuned foundational speechLLM, Qwen2-Audio.

All training used four NVIDIA A10G GPUs with 24GB each, the Hugging Face ecosystem, PEFT with QLoRA 4-bit quantization, Accelerate and DeepSpeed, LoRA rank 32 with α = 128 and dropout 0.05, effective batch size 128 (batch 4 per GPU, accumulation 8), cosine scheduling with maximum learning rate 2×10⁻⁴ and 20% warm-up, over 10 to 15 epochs with AdamW defaults and gradient clipping at 1.0. Inference used greedy search with temperature 0 and a maximum of 512 new tokens.

Why This Matters

The paper provides an unusually candid empirical account of where speechLLMs stand relative to well-established cascaded ASR-plus-NLU pipelines for a commercially important task, showing that a composite speechLLM can surpass the cascade in-domain but still trails it out-of-domain and on unseen slot labels. It also shows that freezing the LLM to preserve its skills is the wrong strategy when the model lacks the target capability, and that data curation choices (whole-call context, randomized context length, distractor slots, diverse prompt templates) are as consequential as architecture.

Real-world applications:

  • Call-center agents that extract structured information from customer conversations in banking, telecommunication, insurance, and retail without retraining for each new slot type.
  • Faster onboarding of new intents and slot labels by specifying them as text at inference time rather than collecting annotated training data.
  • Assistive and analytics tools that summarize spoken customer interactions into structured fields for downstream processing.
  • Any voice-driven enterprise workflow where adaptable extraction of entities, dates, times, and numerals from speech is required.

Industry relevance: The work targets practitioners operating under real compute constraints (four 24GB GPUs) and documents a practical recipe — adapter choice, staged training, multitask expansion — along with the trade-offs between composite and foundational models. The authors also note that their dataset cannot be distributed and that no public benchmark exists at the scale they considered, so reproducibility depends on the documented configurations and prompt templates in the appendices.

Future Directions

  1. Scaling the bimodal model and its training data. The authors attribute the out-of-domain gap to insufficient scale in model size and training data, and call for continued scaling and refined training strategies.

  2. Closing the zero-shot generalization gap. Calls containing entirely unseen label types (the 52% of CallCenter-B labels not overlapping CallCenter-A) remain the hardest setting, and the paper reports no solution that closes this gap.

  3. Making training stable at higher adapter complexity. The transformer adapter required deviation from default learning rate and warm-up settings; how to train larger adapters reliably remains open.

  4. Quantifying statistical significance. The authors report only single point estimates per configuration with ±2 percentage point fluctuation and no variance estimates, and they did not use comprehensive human annotation validation, leaving both the reliability of small differences and annotation quality as open questions.

Target Audience

Industry practitioners and applied researchers building multimodal speech understanding systems under real compute constraints will benefit most, particularly those deciding between cascaded pipelines and speechLLMs. It is also relevant to researchers working on parameter-efficient fine-tuning, modality adapters, and zero-shot slot filling, and to data-curation teams designing LLM-generated annotations for spoken conversational data. Readers looking for public benchmark comparisons or reproducible leaderboard numbers should note that the paper reports none at this scale.

Authors’ abstract

Slot filling is a crucial subtask in spoken language understanding (SLU), traditionally implemented as a cascade of speech recognition followed by one or more natural language understanding (NLU) components. The recent advent of speech-based large language models (speechLLMs), which integrate speech and textual foundation models, has opened new avenues for achieving speech understanding tasks in a more unified, generative, and instruction-following manner while promising data and compute efficiency with zero-shot abilities, generalizing to unseen slot labels. We address the slot-filling task by creating an empirical upper bound for the task, identifying performance, robustness, and generalization gaps, and proposing improvements to the training data, architecture, and training strategies to narrow the gap with the upper bound result. We show that each of these measures improve performance substantially, while highlighting practical challenges and providing empirical guidance and insights for harnessing these emerging models.

Read the original paper