Skip to content
AI.info

Research

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations Overview Research area: Long-term agent memory for multi-party spoken con

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations
arXiv
2609.32522
Published
2026-09-26
Authors
Wenxu Jia, Xize Cheng, Zihan Zhang, Dongjie Fu, Linjun Li, Wenshi Chen, Yangyang Wu, Tao Jin

AI summary

Beyond Dyadic Memory: Interaction-Aware Multimodal Memory with Adaptive Agentic Retrieval for Multi-Party Spoken Conversations

Overview

  • Research area: Long-term agent memory for multi-party spoken conversations — combining speaker identification, multimodal memory construction, and reinforcement-learned retrieval.
  • Technical level: Advanced.
  • Scope: The paper proposes VoxPolyMem, an interaction-aware multimodal memory framework with a trainable retrieval agent, and VoxPolyBench, a benchmark of 1,527 QA pairs over 176 sessions and 18.9 hours of synthesized multi-party speech.

What This Paper Is About

Existing agent memory research mostly targets dyadic (two-person) text or image-text conversations, so systems rarely have to remember who spoke to whom across sessions when many people share a conversation. The authors argue that multi-party spoken conversations require three extra things: recovering participant identities from voice across sessions, storing interaction roles (who said what to whom), and retrieving evidence that may be spread across sessions, memory layers, and modalities. Their goal is a memory framework and a benchmark that jointly evaluate these abilities.

Key Contributions

  1. VoxPolyBench, a benchmark for long-term memory in multi-party spoken conversations covering recurring participant identification, memory evolution, personalized answering, retrieval and reasoning, and interaction attribution — built from 18 scenarios, 176 sessions, 18.9 hours of synthesized speech, and 1,527 QA pairs.
  2. VoxPolyMem, a framework combining incremental (online) speaker identification with a three-layer memory hierarchy: interaction memory, fact memory, and participant profiles.
  3. An adaptive agentic retrieval formulation in which a trainable policy rewrites queries and selects memory layers, tools, and participant filters based on accumulated evidence and a remaining budget.
  4. Evidence-Gain GRPO (EG-GRPO), a GRPO adaptation that assigns round-wise credit for the coverage and ranking of newly retained supporting evidence, explicitly excluding evidence already retained earlier in the branch.

Main Findings

  • Strong overall result: VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points, and leads all four task subsets (retrieval and reasoning 86.4, evolution and conflict 92.1, persona 76.1, interaction attribution 87.5).
  • Generalizes to public image-text benchmarks: It scores 89.6 on Mem-Gallery and 74.4 on H2HMem-Multi, exceeding the strongest external baselines by 11.8 and 8.4 points respectively; the paper's weighted average is 86.6 versus 68.1 for the strongest external baseline.
  • Structure alone goes a long way: The no-RL variant already exceeds all external baselines on every benchmark with a weighted average of 84.4, indicating the memory hierarchy and agentic retrieval loop are effective before policy optimization.
  • Memory abstraction matters most on H2HMem-Multi: Removing fact memory and participant profiles drops the H2HMem-Multi score from 72.1 to 63.6.
  • Interaction structure matters most on VoxPolyBench: Removing interaction edges and relation-based filtering while keeping speaker identities lowers the VoxPolyBench score from 84.0 to 78.6.
  • Iterative retrieval beats query rewriting as an ingredient: Single-round retrieval causes larger declines (e.g., VoxPolyBench 84.0 to 76.1) than keeping multiple rounds with the original query (84.0 to 83.2), showing sensitivity to follow-up searches.
  • EG-GRPO outperforms other retrieval-policy optimizations: It leads all three benchmarks in answer score and both annotated benchmarks in recall, beating MoT-GRPO on Mem-Gallery (89.6 versus 87.1 score; 93.6 versus 91.5 recall), while Search-R1 lowers answer scores across all three benchmarks relative to no RL.
  • Fewer retrieval rounds: EG-GRPO uses an average of 1.2–1.3 rounds, compared with 1.4–1.7 for no RL and 1.4–1.8 for MoT-GRPO.
  • Terminal-coverage reward is insufficient: Terminal-coverage GRPO improved recall on both annotated benchmarks without consistent answer improvements, suggesting higher evidence coverage does not automatically yield better answers.
  • Speaker tracking transfers beyond the synthetic data: The tracker reaches attribution accuracies of 95.0% on IEMOCAP and 87.9% on AMI with reference utterance boundaries.
  • Benchmark transcription quality: Corpus-level WER of 2.14% on all 9,599 dialogue audio segments using Whisper large-v3-turbo.
  • Task composition: VoxPolyBench contains 443 retrieval-and-reasoning instances (29.0%), 324 evolving-memory instances (21.2%), 326 interaction-attribution instances (21.3%), and 434 personalized-memory instances (28.4%), across nine fine-grained task types.

Methodology in Plain English

The authors built a pipeline in three parts.

Build the benchmark first. They defined 18 scenarios (telemarketing, meetings, in-car assistance, household assistance) with recurring participants, then planned "event anchors" — small records of what happens, when, involving whom, and who addressed whom. Human reviewers checked anchors before dialogue generation. Dialogs were generated with GPT-4.1 from those anchors, and speech was synthesized with a fixed reference voice per participant across sessions. Quality control screened synthesis duration and used ASR to flag transcription mismatches, identifier errors, and repetitions, with targeted resynthesis and manual spot checks. QA pairs were then written from the anchors and dialogs, with supporting dialogue turns annotated and checked by humans and an LLM.

Build the memory in layers. Each audio turn is encoded into a voiceprint with an ECAPA-TDNN encoder and compared by cosine similarity against a set of stored voiceprints. A sufficiently close match assigns the turn to an existing anonymous speaker ID; otherwise a new ID is created. Matched voiceprints are updated by exponential moving average only when similarity exceeds a higher update threshold, and after each session provisional IDs may be conservatively merged. Conversations are stored as a directed interaction graph recording who speaks each message and whom it addresses. An LLM then processes the current turn plus up to three preceding turns to link anonymous speaker IDs to names and aliases while jointly extracting facts, participant attributes, and addressees. Facts store content, speaker, addressees, and time; profiles store name, aliases, associated voiceprints, and attributes.

Learn how to retrieve. Retrieval is a sequential decision problem: at each round the agent sees the query, the asker identity (matched from a spoken query against profile voiceprints when possible), previously retained evidence, action history, and remaining budget, and chooses a memory layer, a tool (vector or BM25 text search, text-to-image via image descriptions, image-to-image via image embeddings), a rewritten query, and optional speaker/addressee filters. Results are fused (reciprocal rank fusion), deduplicated, and capped at k = 15 records; the agent either gathers more evidence or answers, with a maximum of three rounds. Training uses EG-GRPO: at each state, 8 candidates are sampled from the same pre-action context, each candidate's newly retained supporting IDs are rewarded by coverage and by their rank in the selected context (discounted logarithmically), previously retained evidence earns no credit, and advantages are normalized within the group. The retrieval policy is an 8B Qwen-VL model trained for one epoch on roughly 900 QA instances (300 from Mem-Gallery, 400 from LoCoMo, 200 from VoxPolyBench). GPT-4.1-mini handles memory extraction and serves as the answer model for all methods; GPT-4.1-mini and GPT-4.1 independently judge answers, and gold annotations are training-only.

Why This Matters

  • Research impact: The paper shifts memory evaluation from dyadic text or image-text settings toward multi-party spoken histories, where acoustic identity recovery and interaction roles must be modeled alongside content. It also offers a credit-assignment scheme (round-wise evidence gain) as an alternative to terminal answer rewards for retrieval agents.
  • Real-world applications:
    • In-car assistants that must track which passenger requested or was told what across multiple trips.
    • Meeting assistants that attribute decisions, commitments, and follow-ups to the right speaker and addressee over many sessions.
    • Household assistants serving multiple recurring residents with distinct preferences and relationships.
    • Telemarketing or customer-service agents that need cross-session continuity about specific participants.
  • Industry relevance: The framework targets persistent, personalized assistance where a fixed context window is insufficient, and the paper reports gains comparable to or above public multimodal memory baselines while using fewer retrieval rounds — relevant to cost and latency budgets for deployed memory systems. The authors also flag voiceprint and profile data as sensitive, recommending informed consent, restricted access, limited retention, and correction/deletion mechanisms, and cautioning against use for surveillance or high-stakes identity decisions.

Future Directions

  • Extending evaluation to real-world spoken interactions with greater acoustic and conversational diversity, since the current benchmark uses generated dialogues and synthesized speech and does not establish robustness across real speakers, accents, or recording conditions.
  • Determining whether the memory hierarchy and EG-GRPO transfer to domains beyond the 18 synthesized scenarios and to the public benchmarks used, which lack the paper's interaction-attribution annotations.
  • Understanding why higher evidence coverage does not always translate into better answers, given the terminal-coverage GRPO result.
  • Testing whether joint optimization of layer selection, tool choice, and query rewriting continues to help as budgets, memory sizes, and numbers of participants grow.

Target Audience

Researchers and engineers working on agent memory, long-term conversational systems, multi-party speech understanding, speaker diarization and identification, and reinforcement learning for retrieval and tool use. It is also relevant to product teams building persistent assistants for meetings, vehicles, households, or customer service, and to benchmark designers interested in how memory evaluation can be extended to multi-party spoken settings.

Authors’ abstract

Long-term memory enables agents to accumulate information and reason across sessions, yet existing research primarily focuses on dyadic text or image-text conversations, leaving long-term memory for multi-party spoken conversations underexplored. This setting requires preserving conversational content, identifying participants across sessions, and retaining who speaks to whom. To this end, we propose VoxPolyMem, an interaction-aware multimodal memory framework combining incremental speaker identification with a memory hierarchy comprising interaction memory, fact memory, and participant profiles. We formulate retrieval as sequential decision-making, where an agent rewrites queries and selects retrieval tools and memory layers based on accumulated evidence to address information gaps. We further introduce Evidence-Gain GRPO (EG-GRPO), which uses round-wise credit assignment to encourage complementary evidence acquisition. We also construct VoxPolyBench to evaluate memory evolution, personalized answering, memory retrieval and reasoning, and interaction reasoning and attribution in multi-party spoken conversations. VoxPolyMem achieves an overall score of 85.0 on VoxPolyBench, surpassing the strongest evaluated baseline by 23.6 points. On Mem-Gallery and H2HMem-Multi, it scores 89.6 and 74.4, respectively, exceeding the strongest evaluated public memory baselines by over 8 points each. These results highlight its potential for persistent, personalized assistance in multi-party multimodal interactions. Code and datasets are available at https://voxpolymem.github.io/VoxPolyBench/demo/

Read the original paper