Research
From Retrieval to Reconstruction: Constructing Evolvable Cognitive Memory for Long-Term Dialogue
Overview Research area: Natural Language Processing — long-term memory architectures for LLM dialogue agents, combining retrieval-augmented generation, knowledge-graph schemas, and cognitive-science-i

- arXiv
- 2610.11314
- Published
- 2026-10-08
- Authors
- Zirui Liao, Zhengxian Wu, Zhuohong Chen, Yunyao Yu, Xiaoyu Liu, Yifan Xu, Haoqian Wang
AI summary
Overview
Research area: Natural Language Processing — long-term memory architectures for LLM dialogue agents, combining retrieval-augmented generation, knowledge-graph schemas, and cognitive-science-inspired memory design.
Technical level: Advanced. The paper assumes familiarity with RAG, graph retrieval, episodic vs. semantic memory, and benchmark evaluation (F1, BLEU-1, LLM-as-judge accuracy).
Scope: The paper proposes CogMem, a training-free PEC²F (Person-Event-Concept-Claim-Fact) cognitive graph architecture in which source-attributed claims are separated from event/fact records, episodic traces are consolidated into semantic facts, and a ReAct-style agent composes four deterministic graph operators to reconstruct answers — evaluated on LoCoMo and LongMemEval with GPT-4o-mini and Qwen2.5-14B-Instruct backbones.
What This Paper Is About
Existing memory systems for long-term dialogue treat memory as passive storage: text is vectorized into a flat index and retrieved by similarity. The authors argue this causes two failures — Semantic Collapse, where an attributed opinion (e.g., "Gina thinks Jon's job loss is a mistake") is returned as if it were an unattributed record of what happened, and a lack of intentionality, since static pipelines cannot adapt their search strategy to connect evidence across sessions. The goal is a memory architecture that stores the epistemic status of statements, evolves episodic experience into semantic knowledge, and actively reconstructs query-relevant evidence through graph operations rather than one-shot similarity matching.
Key Contributions
-
Epistemic Completeness through the PEC²F schema. A person-centric graph with five node types — Person, Event, Concept, Claim, Fact — where Claim nodes model subjective statements as two-hop structures (
Person → CLAIMS → Claim node → ABOUT → Target) so attributed beliefs are never flattened into unattributed attribute edges. The authors frame the Concept and Claim components (the "C²" components) as a structural fix for Semantic Collapse. -
Dynamic Stability through memory consolidation. An offline mechanism clusters episodic Event nodes and, when a cluster reaches a threshold, uses an LLM-parameterized abstraction function to synthesize a semantic Fact node with a computed temporal envelope, while Evidence Mounting adds
SUPPORTED_BYprovenance edges back to original episodes so retrieval can operate at two resolutions. -
Active Plasticity through agentic recall. A rule-based retrieval controller driven by LLM intent parsing composes four deterministic graph operators — anchoring, traversal, intersection, and evidence grounding — implemented inside a ReAct-based Cognitive Search Agent, replacing fixed k-hop pipelines with query-adaptive operator selection.
-
Provenance-preserving Claim reconciliation. Claims sharing a canonical source, target, and predicate are ordered temporally and marked CURRENT, SUPERSEDED, or CONFLICTING; claims from different speakers are never merged, so updates are tracked without erasing disagreement.
Main Findings
-
Semantic Collapse is measurable and largely reduced by the schema. On 150 synthetic fact–claim pairs modeled on LoCoMo conversational patterns, matched fact–claim pairs have a mean BGE-M3 cosine similarity of 0.8231. Flat RAG misattributes the speaker's opinion as the requested record in 64 of 150 cases (42.7%), while type-constrained retrieval over the PEC²F graph reduces this to 9 of 150 cases (6.0%).
-
Strong results on LoCoMo. With GPT-4o-mini, CogMem scores F1/BLEU-1 of 59.20/56.40 (Single Hop), 50.70/46.10 (Multi Hop), 63.80/57.90 (Temporal), and 56.10/51.20 (Open Domain). With Qwen2.5-14B-Instruct, it scores 60.40/55.80, 48.42/44.65, 56.20/49.10, and 55.80/50.10 respectively.
-
Largest gains in multi-hop and temporal reasoning. With Qwen2.5-14B, CogMem reaches 48.42% F1 on Multi-Hop, exceeding the best memory baseline (GAM, 42.96%), the learned traversal baseline (RoG, 41.80%), and the best heuristic graph baseline (HippoRAG, 34.78%).
-
Best overall LongMemEval accuracy with both backbones. CogMem reaches 68.40% with GPT-4o-mini and 66.50% with Qwen2.5-14B, compared with e.g. LightMem (64.29% and 61.95%), GAM (63.82% and 58.71%), and A-MEM (62.60% and 65.20%).
-
Ablations isolate three complementary components (Qwen2.5-14B backbone). Removing the agentic loop causes the largest drop, with Multi-Hop F1 falling 25.61 points (48.42% → 22.81%). Removing Claim nodes reduces Knowledge Update accuracy by 16.41 points (73.08% → 56.67%). Removing consolidation reduces Multi-Hop F1 to 43.56% and Knowledge Update accuracy to 66.67%; the flat variant scores 41.20% multi-hop and 48.10% with a fixed pipeline.
-
Consolidation improves both quality and efficiency. It improves Multi-Hop F1 by 4.86 percentage points and reduces average reasoning steps from 3.42 to 2.67. Per category, F1/steps before→after: Single Hop 58.06→60.40 F1 and 2.84→2.35 steps; Multi Hop 43.56→48.42 and 3.42→2.67; Temporal 56.20→56.20 and 2.78→2.44; Open Domain 54.50→55.80 and 2.86→2.55.
-
Human preference favors CogMem. In a blind study on 50 randomly sampled LoCoMo Multi-Hop and Temporal queries, CogMem wins 68% of comparisons, ties 20%, and loses 12%, judged by two graduate students with a third adjudicating disagreements and no monetary compensation.
-
Zero-shot agentic search beats zero-shot learned traversal. RoG performs competitively on Single-Hop but is lower on Open Domain and Temporal questions; the authors note this measures out-of-the-box transfer since RoG's pre-trained checkpoint is used without task-specific fine-tuning.
-
Residual failure modes. Appendices analyze construction cost and graph growth, operator sensitivity, and 100 manually categorized LoCoMo failures, attributing remaining errors primarily to embedding drift, incomplete search scope, and temporal normalization.
Methodology in Plain English
CogMem processes dialogue incrementally. After each turn, an Atomic Extraction Engine reads the utterance along with speaker, session ID, and timestamp, and writes graph records: time-bounded actions become Event nodes, stable self-reported attributes become Fact nodes, and opinions about another person or concept become Claim nodes carrying explicit source and target fields. Person and Concept strings are normalized and matched against existing anchors before new nodes are created, and every node keeps its source-turn identifier and timestamp.
Separately, an offline consolidation pass groups Event nodes connecting the same person and concept within a time window; once a cluster is large enough, an LLM turns the pattern into a higher-level Fact node with a temporal envelope, and SUPPORTED_BY edges mount the new fact onto its original episodes.
At query time, a controller first performs a mandatory anchoring call using hybrid lexical-plus-embedding scoring to pick entry nodes. Then, in a loop of at most five post-anchoring steps, an LLM parses the query intent and the controller dispatches one of four deterministic operators: intersection when the query asks for comparison or commonality, traversal over Claim relations when the query asks what someone believes, traversal with temporal gating when the query contains time expressions, and evidence grounding when the query asks for verification or detail. The loop stops when evidence is judged sufficient or the step budget is exhausted, and the final answer is generated from the accumulated subgraph.
Evaluation used GPT-4o-mini and Qwen2.5-14B-Instruct as backbones with BGE-M3 (1024-d) embeddings, temperature fixed at 0.0, and the stated hyperparameters: α = 0.3, K_anchor = 15, K_traverse = 15, D_trav = 2, K_intersect = 10, D_step = 5, τ = 3 events, Δt_max = 30 days. Experiments ran on a node with 4× NVIDIA A800 GPUs (80GB VRAM each), a 48-core CPU, and 720GB of RAM. LoCoMo contains ten long multi-session conversations and roughly 1.5K question–answer pairs; LongMemEval contains 500 questions over long cross-session histories. LoCoMo is scored with F1 and BLEU-1, LongMemEval with LLM-as-judge accuracy.
Why This Matters
Impact on research: The paper argues that reliable long-term memory is not just about storing more context but about preserving provenance and reconstructing evidence through task-appropriate paths. It operationalizes Theory of Mind and Complementary Learning Systems ideas within a unified, training-free retrieval system, positioning architectural and representational design — not a new graph primitive — as the contribution. It also names and quantifies a specific failure mode (Semantic Collapse) with a reproducible probe.
Real-world applications:
- Personal assistants that must track how a user's stated preferences and circumstances change across months without treating an outdated statement as current.
- Multi-speaker settings such as care coordination or mediation, where keeping straight who believes what about whom is essential and disagreement must not be silently resolved.
- Customer-support agents that need to distinguish a customer's reported problem (a record) from an agent's or customer's characterization of it (an attributed claim).
- Auditable memory systems in regulated domains, where each answer must be traceable back to the original utterance via evidence-grounding.
Industry relevance: The architecture is training-free and built on top of standard components (BGE-M3 embeddings, an LLM for extraction and intent parsing, a graph store), which lowers adoption cost. The measured trade-off — consolidation increases graph storage while reducing online reasoning steps from 3.42 to 2.67 — is directly relevant to latency and cost planning for deployed memory services. Performance is reported on both a closed-source (GPT-4o-mini) and an open-source (Qwen2.5-14B-Instruct) backbone.
Future Directions
- Incremental, real-time consolidation. The current consolidation mechanism runs periodically and offline; the authors propose updating semantic facts during the conversation stream to reduce the delay between an event and its availability as generalized knowledge.
- Neural tool learning. Replacing the fixed set of hand-crafted, orthogonal operators with operators that are composed or learned from feedback on failed queries — for example, a learned composite operator for counterfactual search over conflicting claims.
- Schema expansion. Extending PEC²F with dynamic Goal nodes that track evolving user intent, or Emotion Vectors that modulate retrieval weight by affective intensity.
- Broader and stronger evaluation. Testing multilingual dialogue, noisy real-world conversations, and continuously deployed assistants, plus larger independent human evaluations beyond the 50-example study with two primary annotators and one adjudicator, and addressing LLM-judge bias in LongMemEval. Hybrid retrieval and more adaptive operator selection are also named as open directions.
Target Audience
Researchers and engineers working on long-term memory, retrieval-augmented generation, or agentic dialogue systems, particularly those designing knowledge-graph-backed memory stores or evaluating on LoCoMo and LongMemEval. It is also relevant to practitioners building persistent personal assistants who need provenance-aware, auditable recall, and to cognitive-science-adjacent researchers interested in how Theory of Mind and Complementary Learning Systems concepts are translated into concrete graph structures and retrieval operators.
Authors’ abstract
Large Language Models (LLMs) serving as long-term dialogue agents require memory systems that support reliable reasoning over extended interactions. However, existing Retrieval-Augmented Generation (RAG) frameworks typically treat memory as passive storage, making it difficult to distinguish source-attributed beliefs from unattributed event/fact records and to connect evidence dispersed across sessions. We introduce CogMem, a cognitive memory architecture based on the PEC$^2$F (Person-Event-Concept-Claim-Fact) graph schema. Dedicated Claim nodes preserve the source and target of subjective statements, while Fact and Event nodes represent semantic and episodic knowledge. Dialogue turns are incrementally converted into provenance-aware graph records, consolidated into higher-level facts, and reconciled into temporally scoped Claim views when the same source provides conflicting updates. For retrieval, a rule-based controller driven by LLM intent parsing composes four deterministic graph operators---anchoring, traversal, intersection, and evidence grounding---to reconstruct query-relevant context. Experiments on LoCoMo and LongMemEval show strong performance, especially on multi-hop, temporal, and knowledge-update tasks. Ablations and a semantic-collapse probe support complementary contributions from epistemic separation, consolidation, and agentic retrieval. Code: https://github.com/Silent-Rain02/CogMem.