Skip to content
AI.info

Research

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models

Overview Research area: Spoken conversational memory and evaluation of large audio language models (LALMs) — a benchmark paper in audio/speech processing (eess.AS). Technical level: Intermediate. It a

VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models
arXiv
2609.32607
Published
2026-09-26
Authors
Yang Xiao, Vidhyasaharan Sethu, Eun-Jung Holden, Ting Dang

AI summary

Overview

  • Research area: Spoken conversational memory and evaluation of large audio language models (LALMs) — a benchmark paper in audio/speech processing (eess.AS).
  • Technical level: Intermediate. It assumes familiarity with context windows, token budgets, and multimodal LLM evaluation, but the core argument is conceptual rather than mathematical.
  • Scope in one sentence: VoxMem is a benchmark that jointly tests what acoustic information a model must remember from multi-session speech and how that information must be used, across 3,196 instances built from 34,743 spoken sessions (177 hours) at four controlled context lengths.

What This Paper Is About

Spoken assistants are expected to remember earlier conversations, but memory in speech involves more than the words that were said — it also includes who spoke, how they spoke, and what sounds were audible around them, information that a transcript cannot recover. The authors argue that existing spoken-memory benchmarks cover only lexical content, use a narrow and ad hoc set of memory operations, and treat memory as a single-session problem. VoxMem builds a two-dimensional taxonomy of acoustic evidence and memory operations into a quality-controlled, multi-session benchmark, then uses it to measure how 15 LALMs actually perform as history grows from 8K to 64K tokens.

Key Contributions

  1. A two-dimensional taxonomy of spoken conversational memory that jointly characterizes the acoustic evidence to be remembered (speech semantics, speaker identity, paralinguistic cues, environmental sound) and the operations applied to it (information extraction, multi-session reasoning, temporal evolution tracking, answer refusal), yielding 15 valid combinations out of the 4 × 4 grid.
  2. The VoxMem benchmark: 3,196 quality-controlled evaluation instances over 34,743 spoken sessions (177 hours), built on multi-session histories with evidence, haystack, and filler sessions, and stratified across four context budgets from 8K to 64K tokens so the same question and evidence are held fixed while history length grows.
  3. A systematic evaluation of 15 LALMs (ten open-weight, five proprietary), showing that no model exceeds 40% overall accuracy at 32K and that memory performance separates sharply by evidence type, operation, and context length.
  4. A fine-grained error analysis attributing failures with an LLM judge, showing that error modes differ qualitatively across question types — speaker-identity errors are mostly binding failures, while paralinguistic and environmental errors are mostly localization failures — pointing to distinct underlying deficiencies rather than one shared bottleneck.

Main Findings

  • No model is close to reliable: At the 32K reference budget, no evaluated model exceeds 40% overall accuracy, and the strongest model reaches 38.5%. The five proprietary models average 33.0% at this budget while the ten open-weight models average 21.9%.
  • What was said is remembered far better than who, how, or what was audible: At 32K, proprietary models average 55.6% on Speech Semantics versus 32.7% on Speaker Identity, 20.0% on Paralinguistic Cues, and 21.9% on Environmental Sound. Open-weight means show the same order at lower levels: 31.6%, 26.7%, 14.5%, and 15.5%.
  • The retained audio-native questions genuinely require the audio: On all 669 answerable questions at 8K, switching from full audio to transcript-only input drops Gemini-3.7-Flash accuracy from 46.2% to 4.4% on audio-native questions, but only from 75.9% to 71.0% on speech semantics.
  • Operation and evidence type interact, and difficulty is not uniform: Multi-Session Reasoning is comparatively robust — 41.8% on Speech Semantics and 43.1% on Speaker Identity, still reaching 26.7% on Paralinguistic and 25.2% on Environmental evidence, above the corresponding Information Extraction scores of 8.8% and 13.5%. Temporal Evolution Tracking tracks speech semantics well (44.5%) but falls to 20.0% for Speaker Identity and nearly collapses for Paralinguistic (3.4%) and Environmental (1.2%) evidence.
  • Refusal has an inverted profile: Models abstain more successfully on Paralinguistic (34.3%) and Environmental (34.2%) questions than on Speech Semantics (19.9%) and Speaker Identity (16.2%) — the opposite of the answerable-question pattern, suggesting they abstain because they cannot use the acoustic evidence, not because they recognize it is missing.
  • Longer histories degrade access to evidence that is still present: From 8K to 32K, mean accuracy drops from 40.1% to 33.0% for proprietary models and from 26.6% to 21.9% for open-weight models, and the decline continues to 64K.
  • Decay rate is not the same as level of difficulty: At 64K, models retain 70.3% of their 8K accuracy on Speech Semantics and 69.8% on Paralinguistic Cues, compared with 66.5% for Speaker Identity and 63.2% for Environmental Sound. Paralinguistic memory is weak at every length but decays at a rate comparable to speech semantics, whereas speaker and environmental memory erode specifically as history grows.
  • Failure modes are qualitatively distinct: Information Extraction errors are dominated by binding and association failures (46%); Multi-Session Reasoning errors split between evidence localization (48%) and operation execution (32%); Temporal Evolution Tracking errors are overwhelmingly localization failures (55%) with operation-execution errors rare (5%). By evidence type, speaker errors are most often binding failures (48%), paralinguistic errors are dominated by localization failures (63%), and environmental errors split between localization (41%) and unsupported answers (39%).

Methodology in Plain English

The authors first define two axes. On one axis, they ask what kind of information has to be remembered from the audio, giving four evidence types: the words spoken, the identity of the speaker, the way something was said (hesitation, surprise, emphasis, speaking rate, laughter), and non-speech environmental sound. On the other axis, they ask what the model must do with that information, giving four operations: pull a fact from one session, combine evidence across sessions, track how a state changes over time, or refuse to answer when the history does not support a unique answer.

Items are then planned before any audio exists: over 30,000 plans specify the question, gold answer, and required evidence, with natural-language questions rendered via Gemini-3.7-Flash and GPT-5.6-Luna. Evidence and haystack sessions are written as user–assistant dialogues, with the two sides authored independently so the answer cannot leak into the wording. User turns are synthesized with Higgs-TTS-3 using fixed VCTK voices across sessions; paralinguistic cues are added through style controls and environmental sounds from ESC-50 are mixed in at 10dB SNR. Filler sessions come from the InstructS2S corpus. Assistant turns are supplied as text by design, since they carry no answer-critical acoustic evidence, which also lets the authors evaluate models that accept audio but do not generate speech.

Each history mixes evidence sessions (which contain the answer), haystack sessions (same topic or acoustic context but a different or absent answer), and filler sessions, and is assembled at 8K, 16K, 32K, and 64K Whisper-encoder tokens — roughly 2.5 to 20 minutes of audio. Longer histories strictly extend shorter ones, and evidence and haystack sessions are spread uniformly so position does not signal the answer. Two levels of quality control check each session and each question family, including whether audio-native questions remain answerable from a transcript or from the paired rendition with the cue removed. Evaluation scores open-ended short answers with an LLM judge (Gemini-3.7-Flash, not one of the evaluated models), with re-judging by GPT-5.6-Luna yielding κ = 0.95 agreement.

Why This Matters

Impact on research. Existing spoken benchmarks mostly evaluate memory for what was said in a single recording or continuous dialogue, and benchmarks that vary history length often vary the question at the same time, confounding length with difficulty. VoxMem supplies a shared taxonomy, a multi-session history structure, and a controlled-length design in which the question and its evidence are held fixed — so a drop in accuracy can be attributed to context growth rather than to a different question. The error attribution further argues that the failures seen are not one bottleneck but several distinct ones.

Real-world applications.

  • Persistent voice assistants that must remember a user's earlier statements, plans, and preferences and attribute them to the right person in a multi-user household.
  • Customer-service and call-centre agents that need to carry state across separate calls rather than within one conversation.
  • Assistive and care-oriented spoken systems where tone of voice or an audible alarm in the background changes what the system should do.
  • Meeting and archival tools that must locate a fact in a long history of separate recordings while avoiding plausible but wrong sessions on the same topic.

Industry relevance. The result that larger context windows do not guarantee stable access to information already inside them is directly relevant to anyone scaling long-context audio models. The finding that abstention on acoustic questions reflects inability rather than genuine uncertainty also matters for safety and trust: a system that refuses because it cannot hear is not the same as a system that knows it does not know.

Future Directions

  • Preserving non-lexical acoustic information. The paper concludes that future systems must retain paralinguistic and environmental evidence and maintain its associations with speakers, events, and sessions over time; how to build such memory is left open.
  • Fixing localization versus binding failures separately. Since paralinguistic errors are dominated by localization failures (63%) and speaker errors by binding failures (48%), the paper's error analysis implies different remedies are needed for different evidence types.
  • Closing the state-tracking gap. Temporal Evolution Tracking nearly collapses for paralinguistic (3.4%) and environmental (1.2%) evidence, which the authors suggest may reflect limited stateful modeling in current LALMs.
  • Extending the taxonomy and the setting. The benchmark covers 20 everyday topics and 799 questions across 15 valid combinations, excluding information extraction over speech semantics by design; whether the taxonomy and the multi-session design extend to other scenarios, languages, or evidence types is not reported.

Target Audience

Researchers and engineers working on large audio language models, long-context multimodal systems, and spoken dialogue memory; benchmark designers who need a principled way to separate evidence type from memory operation; and practitioners building multi-session voice assistants or call-handling systems who need to know where current models actually fail.

Authors’ abstract

Spoken conversational systems must recover information from prior interactions (i.e., memory), yet relevant information in speech extends beyond what was said to who said it, how it was spoken, and what was audible, information that exists only in the audio signal and cannot be recovered from a transcript. Beyond what to remember, memory also demands diverse operations: retrieving a single fact, integrating evidence across turns, tracking an evolving state. Real interactions further unfold across sessions, meaning information accumulates across distinct episodes rather than a single continuous recording. Existing benchmarks fall short on all three dimensions: they focus primarily on lexical content, adopt limited and ad hoc memory operations, and treat memory as a single-session problem. We argue that principled memory evaluation requires jointly characterizing the acoustic evidence to be retained and the operations applied to it, and introduce a taxonomy along these two axes. Building on this taxonomy, we present VoxMem: 3,196 evaluation instances over 34,743 spoken sessions (177 hours) crossing four acoustic evidence types (speech semantics, speaker identity, paralinguistic cues, environmental sound) with four memory operations (information extraction, multi-session reasoning, temporal tracking, and answer refusal), grounded in multi-session histories and stratified across context budgets from 8K to 64K tokens. Evaluating 15 LALMs, no model exceeds 40% at 32K. Models retain what was said far better than who said it, how, or what was audible, a gap that widens for complex operations, grows with history length, and manifests as qualitatively distinct failure modes across evidence types. VoxMem aims to provide a foundation to measure and drive progress on the full scope of spoken conversational memory.

Read the original paper