Skip to content
AI.info

Research

Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies

Overview Research area: Long-term memory for Large Language Model (LLM) agents, specifically retrieval-augmented generation (RAG)-based memory systems for dialogue. Technical level: Intermediate. The

Memory Type Varies: Empowering LLM Agents for Long-Term Memory with Diverse Strategies
arXiv
2610.11573
Published
2026-10-08
Authors
Yi Wen, Derong Xu, Pengyue Jia, Yichao Wang, Yingyi Zhang, Maolin Wang, Junyi Li, Wenlin Zhang, Xiaopeng Li, Yong Liu, Xiangyu Zhao

AI summary

Overview

  • Research area: Long-term memory for Large Language Model (LLM) agents, specifically retrieval-augmented generation (RAG)-based memory systems for dialogue.
  • Technical level: Intermediate. The paper mixes a concrete engineering framework with a formal theoretical bound, but its central idea (memories are not all alike) is explained in accessible terms.
  • Scope: This paper argues that treating all dialogue memories with one retrieval strategy is suboptimal, and proposes a three-way memory taxonomy, a labelled dataset (TriMEM) and a routing framework (MemoType) that applies a different retrieval strategy to each memory type.

What This Paper Is About

Most RAG-based memory systems store past conversations and retrieve them by a single similarity-matching procedure, applying the same treatment to every memory. The authors observe that memories differ fundamentally: an event recollection ("we met at the café on Tuesday") and a stable fact ("I work as a nurse") carry different kinds of information, and one retrieval strategy cannot serve both well. The paper builds a labelled benchmark of memory types and a framework that classifies each memory and query, then retrieves with a strategy tailored to the type.

Key Contributions

  1. A memory categorisation principle and the TriMEM dataset. Drawing on Tulving's distinction theory from cognitive psychology, the authors split dialogue memory into Episodic Memory, Personal Semantic Memory, and General Semantic Memory, and release TriMEM, a memory multi-class dataset of 6,000 samples with fine-grained type annotations across user-user and user-assistant scenarios.

  2. The MemoType framework. MemoType uses a learned router to recognise memory and query types adaptively, retrieves within the predicted memory categories, and applies a different retrieval strategy per type (decompose-then-align for episodic, memory-side query expansion for personal semantic, hypothetical memory for general semantic). A query-conditioned memory pruning module then removes query-irrelevant content before generation.

  3. A theoretical upper bound for single-strategy retrieval. The authors prove that in a corpus with C categories separated by angular distance δ, any single shared embedding encoder has expected query-document alignment bounded above by 1 − (C−1)δ/(2C), so no single strategy can excel across all categories simultaneously.

  4. Empirical validation across four long-term memory benchmarks. MemoType is compared against nine baselines on LongMemEval-S, LoCoMo, LongMemEval-M and PerLTQA, reporting up to 16.18% improvement in Recall@1.

Main Findings

  • Episodic memory is rare in real dialogue. In LongMemEval-S, episodic memory accounts for only 12.35% of instances, meaning semantic memory dominates and is poorly served by the coarse episodic/semantic dichotomy.

  • Semantic memory splits visibly. t-SNE visualisations of LongMemEval reveal a clear semantic gap between personal content and general knowledge, which motivated the finer three-way classification.

  • Retrieval gains on LongMemEval-S. MemoType achieves Recall@1 of 59.15%, outperforming the strongest baseline LightMem by 2.16% and the structural method HippoRAG2 by 8.51%.

  • Largest retrieval gain on LongMemEval-M. MemoType reaches Recall@1 of 48.09%, surpassing SeCom by over 16 points, which the authors attribute to the advantage of type-aware retrieval over uniform segmentation.

  • Gains on user-user dialogue and on PerLTQA. On LoCoMo, MemoType beats A-Mem by 2.27% in Recall@1; on PerLTQA it reaches 71.92% Recall@1, exceeding the second-best method MemoryOS by 14.73%.

  • Generation improvements. On LongMemEval-S, MemoType attains a GPT4Judge score of 56.80 and BLEU of 4.66, surpassing LightMem by over 13 points and 2.38 points respectively. On LoCoMo it achieves F1 of 22.22 and GPT4Judge of 48.79. On LongMemEval-M it reports GPT4Judge 50.20 and F1 18.77; on PerLTQA GPT4Judge 62.83, F1 47.26, BLEU 20.65, Rouge1 51.20, Rouge2 33.02, RougeL 44.78, RougeLsum 44.79 and BERTScore 91.61 — the best reported for every method on that dataset.

  • Type-aware routing beats every single strategy. Against three uniform strategies (Key Expansion, Hypothetical Query, Hypothetical Memory), the type-aware strategy reaches Recall@1 of 59.15 on LongMemEval-S (5.24 over the best single strategy), 31.07 on LoCoMo (5.14 over the best), and 71.92 on PerLTQA (5.68 over the best).

  • Pruning helps generation. Adding the query-conditioned memory pruning module raises F1 on LongMemEval-S from 19.55 to 20.31, showing that retrieved segments carry query-irrelevant noise.

  • Routing quality comes from data, not model size. The memory router is a fine-tuned lightweight Qwen3-1.7B model with LoRA, trained as a multi-label classifier, and the authors report that even lightweight encoders such as BERT achieve competitive routing performance — suggesting dataset quality rather than model scale drives accuracy.

  • Filtering shrinks the search space. Type-aware filtering eliminates approximately 50% of unrelated memories across datasets, with K=5 hypothetical memories used to balance coverage and robustness.

  • Some baselines could not be run on LongMemEval-M. A-Mem, HippoRAG2 and Raptor had memory construction latency exceeding one week on that dataset, so their results are not reported; Raptor also cannot be evaluated on retrieval because it generates new text.

Methodology in Plain English

The researchers start from a cognitive-science distinction between memories of events and memories of facts, then refine it. Because event-centred dialogue turned out to be rare (12.35% of LongMemEval-S) and because visualisation showed personal and general knowledge forming separate clusters, they settle on three types: episodic, personal semantic, and general semantic.

They build TriMEM, a 6,000-sample annotated dataset spanning both user-user and user-assistant conversations, and use it to train a small router that labels a memory with one or more of the three types, since real conversations often mix topics.

Routing queries is harder, because queries are short and underspecified while memories are rich. So instead of classifying the query directly, they have an LLM generate K=5 hypothetical memories the query might retrieve, classify those with the memory router, and take the most frequent category as the query type. Retrieval is then restricted to the memories in those predicted categories, cutting roughly half the corpus.

Each type gets its own retrieval treatment. Episodic memories are decomposed into typed elements (time, participants, location, event context) and matched element by element, because a single embedding would drown fine-grained cues in dominant event semantics. Personal semantic memories are handled by reversing expansion: the system generates hypothetical questions that each memory could answer and matches the query against those. General semantic memories are standardised enough that hypothetical memory generation works directly. Each strategy's scores are fused with the original similarity using Reciprocal Rank Fusion.

Finally, a pruning module makes a binary relevance judgement for each retrieved segment and drops the irrelevant ones before the LLM generates an answer.

The theory section formalises why one strategy cannot serve all: with C categories separated by at least δ in embedding direction, the expected alignment between queries and documents under any shared encoder is capped at 1 − (C−1)δ/(2C). Improving alignment for one category necessarily suppresses it for others.

Why This Matters

Impact on research: The paper challenges a default assumption in memory-augmented LLM systems — that a better encoder or a better index is the main lever. It argues that heterogeneity itself is the bottleneck, supplies a labelled resource and a formal bound that explains the ceiling, and offers evidence that data quality, not model scale, drives routing accuracy.

Real-world applications:

  • Personal assistants that must answer both "what did we discuss last Tuesday?" (episodic) and "what are my dietary restrictions?" (personal semantic) from the same conversation history.
  • Customer-service agents that need to recall a specific complaint alongside stable account facts.
  • Companion or tutoring agents running long multi-session dialogues where stale or irrelevant retrieved turns degrade responses.
  • Enterprise knowledge assistants mixing conversational context with general world knowledge.

Industry relevance: The framework is designed around a lightweight Qwen3-1.7B router and reports that even BERT-class encoders are competitive, which lowers deployment cost. The reported ~50% reduction in retrieval space speaks to serving efficiency, and the paper contrasts RAG-based memory with agent-memory approaches that need reinforcement learning and incur higher construction latency.

Future Directions

  • Extending beyond three types. The authors explicitly flag the question of whether different memory types at the corpus level can benefit from better strategies as a promising direction, implying the taxonomy may need to grow.
  • Balancing performance and latency. The paper notes that RL-trained memory agents achieve strong performance but carry higher construction latency, and that balancing this trade-off remains open for those methods.
  • Improving episodic retrieval. Episodic memory is described as the most challenging of the three categories, and its decompose-then-align strategy adds computational cost; reducing that overhead is a natural next step.
  • Robustness of query routing. The routing pipeline depends on LLM-generated hypothetical memories and a K=5 window whose coverage-versus-robustness trade-off the paper sets but does not exhaustively explore.

Target Audience

Researchers and engineers working on LLM agents, retrieval-augmented generation and long-term conversational memory, who want a taxonomy and labelled benchmark for memory types alongside a practical routing architecture. It also suits readers interested in formal limits on single-encoder retrieval over heterogeneous corpora, and practitioners building production memory layers who need evidence on which retrieval strategy to apply to which kind of memory.

Authors’ abstract

The memory capabilities of Large Language Models (LLMs) have garnered increasing attention recently. Despite great success achieved, existing retrieval-based memory approaches typically overlook the differences between memories and employ a unified strategy to process all memories, leading to suboptimal performance. Thus, an intuitive question arises: can we categorize memory into different types and select appropriate strategies? However, given the topic-rich, scenario-complex, and boundary-blurred nature of memory scenarios, achieving precise classification of memories is not easy. To address this challenge, we propose a memory multi-class dataset in this paper, termed TriMEM, which provides precise annotations for memory types across diverse scenarios. Building upon this foundation, we propose a novel memory framework, named MemoType, which can adaptively recognize each memory and query type with the learned router model. With the memory and query routing, MemoType can retrieve the memory with corresponding query types and design tailored retrieval strategies, thereby enhancing the retrieval performance. Moreover, we theoretically prove that any single retrieval strategy is subject to a fundamental upper bound on its expected retrieval precision in multi-class corpora, leading to systematic precision degradation. Extensive experiments on three datasets demonstrate that MemoType consistently outperforms existing methods, achieving up to 16.18% improvement in Recall@1.

Read the original paper