Skip to content
AI.info

Research

Capability-Driven Self-Evolution of Agent Memory

Overview Research area: LLM agent memory — specifically self-evolving memory programs that store and retrieve information from past interactions, with this paper focusing on the optimization/search pr

Capability-Driven Self-Evolution of Agent Memory
arXiv
2610.06361
Published
2026-10-05
Authors
Yaoqi Chen, Yuru Feng, Qianxi Zhang, Baotong Lu, Jianan Lu, Zhirui Wang, Shusen Xu, Zewen Jin, Zengzhong Li, Cheng Li, Qi Chen

AI summary

Overview

  • Research area: LLM agent memory — specifically self-evolving memory programs that store and retrieve information from past interactions, with this paper focusing on the optimization/search procedure used to improve them.
  • Technical level: Intermediate. The core idea (capability-level instead of overall-score guidance) is easy to grasp, but the paper's scheduling formula, interface decomposition, and evolution stages are detailed enough to require familiarity with agent memory systems and LLM-driven code search.
  • Scope in one sentence: The paper introduces capability-driven evolution and the PrisMem system, which guides memory-program self-evolution along five individual memory capabilities rather than a single overall score, and reports gains over seven baselines on BEAM-1M and LongMemEval-M.

What This Paper Is About

Agent memory systems are executable programs that decide how an LLM agent stores and retrieves information from past conversations. Existing self-evolution methods improve these programs using one aggregate performance number, which the authors argue both blurs the direction of the next revision and hides improvements in one capability that are cancelled out by regressions in another. The paper's goal is to lift evolution guidance from a single overall-performance dimension to multiple capability-level dimensions, so that useful revisions survive long enough to be refined and later merged into one stronger memory program.

Key Contributions

  1. Capability-driven evolution as a new paradigm. The authors extend search guidance from overall performance to individual capability dimensions, so that capability-specific feedback supplies explicit revision directions and capability-level evaluation preserves revisions that improve at least one capability even when the overall score does not improve.
  2. The PrisMem system. PrisMem combines three mechanisms: dependency-aware capability selection (prioritizing capabilities that are weak themselves or act as bottlenecks for others), history-guided diagnosis (selecting severe, historically actionable, and diverse failure cases), and trace-guided integration (merging capability specialists using paired differential cases).
  3. A fine-grained memory interface. The memory program is decomposed into five components — Extraction, Indexing, Planning, Retrieval, and Answer — giving the coding model more localized modification boundaries than the usual "add" and "retrieve" abstraction.
  4. Empirical gains on million-token histories. Evolving on shorter splits (BEAM-100K and LongMemEval-S) and testing on longer ones (BEAM-1M and LongMemEval-M) yields improvements of 7.83–10.54 percentage points over the best-performing self-evolving baselines.

Main Findings

  • Holistic evolution plateaus. M⋆ improves over the strongest static baseline by 0.99 pp on BEAM, and EvolveMem by 0.34 pp on LongMemEval. The authors attribute this to mixed feedback obscuring optimization directions and aggregate scores masking capability gains offset by regressions.
  • Hidden gains are common. In the evolution traces of M⋆ and EvolveMem on BEAM, 80.5% of revisions without an overall gain still improve at least one capability.
  • PrisMem leads on both benchmarks. It improves overall score over the strongest static and self-evolving baselines by 8.17–11.53 pp and 7.83–10.54 pp respectively, achieving best performance in nearly all capability comparisons. On BEAM-1M its overall score is 63.39 ± 1.02 versus 52.85 ± 2.49 for M⋆ and 51.86 ± 0.87 for HippoRAG2; on LongMemEval-M it is 75.50 ± 1.80 versus 67.67 ± 1.53 for EvolveMem and 67.33 ± 0.76 for A-MEM.
  • Static methods have uneven capability strengths. HippoRAG2 leads in factual retrieval (65.24 on BEAM-1M), while A-MEM is competitive in preference extraction (65.39 on BEAM-1M).
  • Generalization across splits and models. Programs evolved on the shorter splits retain their advantage on the longer splits. With GPT-5.5, PrisMem outperforms baselines by 6.17–20.41 pp on BEAM-1M. The paper also reports cross-model and cross-dataset transfer experiments in Appendix C.2.
  • Evolution cost is competitive. Under the same 20-round budget, PrisMem uses 73.7 ± 6.4 M tokens on BEAM and 143.3 ± 9.6 M on LongMemEval, a 4%–30% reduction over M⋆ (76.5 ± 7.1 M and 202.3 ± 10.3 M). EvolveMem is cheaper (69.2 ± 7.3 M and 102.7 ± 8.4 M) because it evolves only retrieval and answer configurations, whereas M⋆ and PrisMem evolve the full program and pay additional extraction costs. Task token usage for PrisMem is 100.59 ± 5.02 K per question on BEAM-1M and 2232.65 ± 32.72 K on LongMemEval-M.
  • Ablations attribute the gain. On BEAM-1M with Qwen3.8-27B, removing capability evolution costs 10.15 pp (53.24 ± 2.38), removing trace-guided integration costs 3.16 pp (60.23 ± 1.86), a fixed round-robin schedule costs 2.55 pp (60.84 ± 1.13), and random case selection costs 5.26 pp (58.13 ± 1.25). Without trace-guided integration the variant still outperforms the baselines.
  • Integration recovers the peak. Capability scores improve intermittently across Stage II, and the final integrated program closely matches the peak capability-wise performance reached during evolution, exceeding the boundary of holistic evolution.
  • Question-only capability annotation is unreliable. Tested with GPT-5.5 on 700 BEAM-1M questions, the predicted primary capability matched the benchmark-type mapping for only 337 questions (48.14%), motivating annotation that also uses reference answer and evidence.

Methodology in Plain English

The starting point is a deliberately simple seed memory program and a training set of tasks. Each task is annotated offline by an LLM with one primary capability and up to two secondary capabilities drawn from five: factual retrieval, temporal tracking, preference extraction, multi-session synthesis, and adversarial. These labels organize feedback but are never shown to the memory program during evaluation, so comparisons with baselines stay fair.

Evolution runs in three stages. Cold start runs for 5 rounds, using a metric table that aggregates diagnostics by capability and component — comparing reference evidence against extracted memory units, retrieved content, and content retrieved for the actual question — to fix broad weaknesses and produce a balanced base program. Capability refinement then picks one target capability at a time. Selection scores each capability by its own error plus its strongest measured pressure on other capabilities (an excess-over-base-rate computation that catches cases where one capability's errors are disproportionately concentrated on tasks that also need it), divided by one plus the number of times it has already been chosen. Refinement starts from that capability's best-performing specialist rather than the global best program. Diagnosis selects 5 cases per round: tasks with the target as primary capability, scoring at or below 0.8, ranked by a weight that favors failures past programs handled well (regression witnesses) while discounting cases and secondary-capability combinations already covered often. Trace-guided integration then merges specialists into a unified program. For each capability it picks the training case with the largest score gap between the base program and the specialist being integrated — a "paired differential case" — and uses the contrasting execution traces to plan the merge. Integration runs for up to 3 rounds, and if it fails to improve overall score or drops any capability beyond a threshold, the failed implementation and its traces are handed back to the coding model to retry.

Cost is reduced two ways: diagnostic contexts are compacted (cutting diagnostic-case token cost by 44% on average) and, because the memory interface is fine-grained, outputs from components a revision did not touch are reused instead of re-executed. All evolution methods train for 20 iterations, on Qwen3.8-27B with Qwen3-Embedding-8B, on NVIDIA A100 GPUs with vLLM 0.20.0 under Python 3.12; each experiment is run three times and reported as mean ± standard deviation.

Why This Matters

Impact on research. The paper makes the case that the objective function of memory self-evolution, not just the memory architecture, is a design lever. It provides evidence that aggregate scores systematically discard useful revisions, and it offers a reusable recipe — capability vocabularies, dependency-aware scheduling, capability-level retention, and differential-case integration — for anyone building evolutionary search over LLM-written programs.

Real-world applications:

  • Long-running personal assistants that must simultaneously track facts, preferences, changing states, and cross-session context without letting one capability degrade another.
  • Customer-support and CRM agents that need to abstain or flag contradictions rather than invent answers (the paper's adversarial capability) while still retrieving preferences and order history.
  • Enterprise knowledge agents over million-token histories, where the paper's compact diagnostic contexts and reuse of unaffected pipeline components directly reduce the cost of iterating on a memory design.
  • Agent platforms that want to tune memory to a deployment's own question mix instead of adopting a single fixed configuration.

Industry relevance. The reported 4%–30% evolution-token reduction over M⋆ and the 44% diagnostic-token saving matter for the economics of repeatedly evaluating candidate memory programs. The finding that programs evolved on short splits transfer to million-token splits, and that gains hold with a different underlying model (GPT-5.5), speaks to practical deployment where re-evolution per context length is otherwise expensive.

Future Directions

  • Broadening the capability vocabulary. The paper presents its five capabilities as one concrete instantiation drawn from a review of benchmarks and existing systems; other domains and benchmarks would likely require different or additional capability sets.
  • Whether capability-first refinement is the decisive factor or the specific scheduler is. The ablation shows that removing capability-driven evolution entirely costs 10.15 pp, while swapping only the scheduling policy costs 2.55 pp and swapping case selection costs 5.26 pp, leaving open how much further scheduling or case-selection design could add.
  • Extending beyond the five-component interface. The Extraction, Indexing, Planning, Retrieval, and Answer decomposition sets the boundary of what a revision can localize, and reuse for efficiency depends on it; finer or differently factored interfaces remain untested.
  • Behavior when a capability genuinely cannot be improved without harming another. Integration aborts and retries when a capability drops beyond the threshold, but the paper does not report how often such conflicts are irreconcilable versus resolvable.

Target Audience

Researchers and engineers working on LLM agent memory, long-context agents, and automated program search or self-improvement loops. It will also interest practitioners who must choose or tune a memory system for a specific deployment, and readers studying evaluation methodology, since the paper's central argument concerns what optimization signal should be measured rather than only which architecture performs best.

Authors’ abstract

Memory self-evolution uses task feedback to iteratively improve executable memory programs that store and retrieve information from past interactions. Existing approaches typically adopt holistic evolution, deriving revision directions from mixed feedback and judging progress by overall performance. This can obscure optimization directions and hide capability-specific gains offset by regressions elsewhere, leaving promising directions underexplored. We introduce capability-driven evolution, which extends search guidance from overall performance to individual capability dimensions, preserving promising revisions and expanding exploration beyond the boundaries of holistic evolution. We propose PrisMem, which uses dependency-aware capability selection to prioritize targets with potential cross-capability benefits and history-guided diagnosis to refine capability specialists. Trace-guided integration compares evaluated programs on paired differential cases, using their behavioral differences to consolidate complementary gains into a unified memory program. Experiments show that PrisMem outperforms the strongest baselines by 10.54 and 7.83 percentage points on BEAM-1M and LongMemEval-M, respectively, demonstrating its effectiveness on million-token histories.

Read the original paper