Research
ExpVoyager: Direct Experience Navigation for Dynamic Agent Skill Synthesis
Overview Research area: Natural Language Processing / LLM agents — specifically experience-based skill synthesis and self-evolving agents. Technical level: Advanced (assumes familiarity with LLM agent

- arXiv
- 2609.32630
- Published
- 2026-09-26
- Authors
- Kwangwook Seo, Dongha Lee
AI summary
Overview
- Research area: Natural Language Processing / LLM agents — specifically experience-based skill synthesis and self-evolving agents.
- Technical level: Advanced (assumes familiarity with LLM agents, retrieval, trajectories, and agent harnesses).
- Scope: The paper reframes agent skill synthesis as an active navigation problem over raw past experience rather than a one-shot abstraction into fixed skills, and evaluates the resulting framework (ExpVoyager) on three interactive agent benchmarks.
What This Paper Is About
LLM agents accumulate trajectories from past task attempts, and existing methods try to turn that experience into reusable "skills" by abstracting it into fixed procedural knowledge before anyone knows what future tasks will require. The authors argue this risks throwing away knowledge that later turns out to be critical while keeping instance-specific details that never matter, and that similarity-based retrieval of top-k trajectories is a poor substitute because irrelevant knowledge hides in similar trajectories while critical cues can sit in superficially dissimilar ones. Their goal is to let a skill curator actively navigate accumulated raw trajectories — choosing what to inspect and at what resolution — and synthesize a task-specific skill on demand for a frozen executor.
Key Contributions
- The paper reframes agent skill synthesis as a dynamic navigation problem over past experience, enabling targeted and fine-grained access to task-relevant procedural knowledge at the moment it is needed.
- It introduces ExpVoyager, a framework built on two components: a Navigable Interface that exposes raw experience through multiple views (trajectory-level and step-level) and resolutions, and a Navigation State that continually connects what the curator has accessed with what it still needs to investigate.
- It reports extensive experiments on three agent benchmarks (ALFWorld, WebShop, ScienceWorld) showing consistent downstream task improvements, continual gains as the experience space scales, and compatibility with pre-existing skills under efficient experience access.
- It provides two preliminary analyses that quantify the failure modes of prior paradigms: pre-constructed skills lose the majority of oracle knowledge, and existing retrieval paradigms achieve low knowledge recall while precision drops rapidly as retrieval depth increases.
Main Findings
- Pre-constructed skills lose knowledge: The first preliminary analysis compared skills built in advance by existing methods against oracle knowledge from the corresponding source trajectories. Even the best-performing method failed to preserve the majority of the oracle knowledge contained in raw experience, indicating much of the later-critical knowledge is already lost during pre-construction. The analysis relied on 80 oracle skills verified through execution on ALFWorld test tasks and 300 sampled ALFWorld training source trajectories.
- Retrieval gives limited access: The second analysis found that existing methods achieve low knowledge recall, leaving a substantial portion of target-relevant knowledge inaccessible. Increasing retrieval depth k produced marginal recall gains that were increasingly outweighed by redundant and irrelevant knowledge, causing precision to drop rapidly.
- Improved task performance with Qwen3.5-9B as curator: With a Qwen3.5-9B executor, ExpVoyager reached 63.7 SR on ALFWorld (versus 41.4 for base ReAct), 28.7 SR on WebShop (versus 19.6), and 37.4 SR on ScienceWorld (versus 27.0). It outperformed AWM, RBank, Trace2Skill, and SkillTTA on all three benchmarks.
- Improved task performance with Gemma4-31B as executor: ExpVoyager reached 69.6 SR on ALFWorld (versus 53.9 base), 37.9 SR on WebShop (versus 24.7), and 55.6 SR on ScienceWorld (versus 48.2), again beating all listed baselines.
- Fewer execution steps: ExpVoyager also reduced the number of execution steps — for example, 25.9 steps on ALFWorld versus 34.1 for base ReAct with the Qwen3.5-9B executor — indicating the synthesized skills guide more efficient execution.
- Transferable across backbone configurations: With a fixed Gemma4-31B executor, ExpVoyager with curator Qwen3.5-9B scored 69.6/37.9/55.6 SR (ALFWorld/WebShop/SciWorld), with Gemma4-31B 71.3/38.4/60.8, and with GPT-5.4-mini 75.9/38.1/62.1. A smaller curator improved the performance of a much larger executor.
- Online self-evolution: In an online setting with no pre-collected experience corpus and three random task orderings, ExpVoyager showed limited gains and could even degrade performance early when experience was scarce, but continued improving as experience accumulated while baselines plateaued, progressively widening the gap.
- Scaling with experience: Under a controlled offline setup using nested subsets of 1,000 source trajectories (100%), baseline performance declined as more source experience became available, whereas ExpVoyager consistently improved and its advantage became more pronounced at larger scales.
- Synergy with pre-constructed skills: Combining ExpVoyager with pre-constructed skills significantly boosted the corresponding baselines, and combined variants consistently reduced task-time cost compared to standalone ExpVoyager while achieving higher task performance with only a single exception.
- Better use of access budget than agentic search: Against iterative top-k retrieval extensions of baselines under comparable experience-access budgets, ExpVoyager consistently outperformed them. Extra budget gave baselines only marginal gains and could hurt performance, while ExpVoyager benefited from additional budget up to 20 rounds before declining with further navigation.
- Ablation on the two components: On ALFWorld, full ExpVoyager scored 63.7 SR, removing the Navigable Interface dropped it to 55.3, and removing the Navigation State dropped it to 58.9. Removing the Navigable Interface lowered knowledge recall and increased irrelevant knowledge; removing the Navigation State slowed recall growth by repeatedly visiting duplicated knowledge.
- Navigation behavior: The curator selected actions according to information needs — trajectory-view for procedure overview, step-view for local execution records, and trajectory inspection to understand context across steps — and explored procedural knowledge distributed across multiple source trajectories.
Methodology in Plain English
The authors first ran two diagnostic studies to show existing approaches fall short. They built a set of oracle skills by repeatedly generating candidate skills from collected successful and failed executor trajectories on ALFWorld tasks and keeping only those that let the same frozen executor succeed, producing 80 verified oracle skills, and then annotated which of 300 sampled source trajectories contained knowledge contributing to each skill. Separately, they measured how well existing retrieval setups surfaced that needed knowledge at both the source level and the knowledge level, across different retrieval depths.
Building on those findings, they designed ExpVoyager. A skill curator is given an interface over the pool of past trajectories that exposes each trajectory in complementary views: trajectory-level views (ordered action sequence, execution outcome, environment metadata) and step-level views (observation, recorded reasoning, executed action, immediate result, plus surrounding context). The curator acts through operations — search_exp, which takes a view level, fields, a regular-expression pattern, and a result limit, returning complete records that retain a reference to their source trajectory, and inspect_traj, which expands a reference into the full chronological trajectory.
To keep navigation coherent, the curator maintains a Navigation State made of two parts: accumulated knowledge items curated for the current target, and open questions still to be resolved. After each round the state is updated: newly accessed observations are interpreted into target-relevant knowledge (distinguishing what stays source-specific, what transfers procedurally, and what still needs verification), earlier judgments can be revised, resolved questions are closed, and new gaps are added. The curator then picks a useful open question and the action that can help answer it, iterating for up to a set number of rounds. Finally it synthesizes a skill.md from the accumulated state, turning procedural knowledge into execution guidance, retaining unresolved conditions as execution-time checks and failure lessons as cautions or recovery guidance. The skill is supplied as extra context to a frozen executor, whose weights are never modified.
Evaluation used ALFWorld (household interaction; six task categories across 120 rooms; 1,000 source trajectories; all 140 official test instances; 50-step episodes), WebShop (online shopping; built from roughly 1.18 million Amazon products; 1,000 training instructions sampled with seed 42; all 500 held-out test instructions; 15-step episodes), and ScienceWorld (scientific reasoning — further dataset details are cut off in the provided content). Unless stated otherwise, Qwen3.5-9B served as both curator and frozen executor, the navigation budget was R = 20 rounds, and results were averaged over three random runs. Baselines covered base ReAct without skills, pre-constructed skill methods (AWM, ReasoningBank/RBank, Trace2Skill), and test-time skill synthesis (SkillTTA), with all methods sharing the same backbones and source pool within each configuration.
Why This Matters
-
Impact on research: The paper shifts the conversation from "how do we compress experience well in advance" to "how do we let agents decide, at task time, which parts of raw experience to inspect." Its preliminary analyses quantify two concrete failure modes of current practice — knowledge loss during pre-construction and precision collapse as retrieval depth grows — giving the field measurable targets. It also connects agent skill synthesis to agentic search and corpus interaction, arguing that execution experience has a structure (context, decision, consequence) that generic retrieval interfaces do not respect.
-
Real-world applications:
- Agent harness systems that must supply procedural guidance to a fixed executor at runtime without retraining model weights.
- Deployed agents that accumulate experience incrementally and need to improve online rather than from a pre-collected corpus.
- Enterprise or assistive agents that already maintain curated skill libraries but need those skills specialized on demand for particular tasks.
- Shopping, household-robotics-style interaction, and scientific reasoning assistants, the three domains covered by the benchmarks.
-
Industry relevance: Because ExpVoyager leaves the executor frozen and works as a complementary layer, it fits the harness-centric view of agent deployment, where the surrounding execution stack (interfaces, context, external resources) is treated as a first-class design object. The finding that a small curator can improve a much larger executor suggests a cost-effective deployment pattern, and the budget analysis showing benefits up to 20 navigation rounds versus the marginal returns of iterative retrieval speaks directly to serving cost and latency trade-offs.
Future Directions
- ScienceWorld and other domains: The provided content truncates the ScienceWorld dataset description and appears to cut off before the appendix; full dataset details and evaluation protocols for that benchmark remain to be examined.
- Navigation budget policy: ExpVoyager improved with additional budget up to 20 rounds and then declined, so how to decide when to stop navigating — or adapt the budget per task — remains open.
- The early-experience problem: In the online setting ExpVoyager showed limited gains and could degrade performance when accumulated experience was still small; understanding and mitigating this cold-start behavior is an open question.
- Deeper integration with existing skills and harnesses: The paper shows synergy with pre-constructed skills and reports that combined variants reduce task-time cost; how far that combination can go, and how it interacts with adaptive harness designs, is left for future work.
- Curator scaling: Results show stronger curators (GPT-5.4-mini) yielding higher scores than smaller ones with the same executor, raising the question of how curator capability, model scale, and skills transferability relate.
Target Audience
Researchers and engineers working on LLM agents, self-evolving agents, agent memory and skill libraries, and agent harness systems. It will be most useful to readers already comfortable with agent trajectories, tool-use loops, and retrieval baselines, and to practitioners who need to decide how a deployed agent should turn its growing log of past executions into guidance for the next task.
Authors’ abstract
Learning from experience in LLM agents has become a key paradigm for developing self-evolving agents that continuously learn and expand their capabilities. Within this paradigm, synthesizing the agent skill has emerged as a promising solution for transforming accumulated experience into reusable procedural knowledge, serving as an important layer for the harness system that supplies agents at runtime. Despite its potential, existing approaches largely abstract past experience into fixed procedural knowledge before downstream demands are known, which risks discarding knowledge that later becomes critical while retaining instance-specific details irrelevant to future tasks. In this paper, we reframe agent skill synthesis as a dynamic navigation problem over past experience, where agents actively explore accumulated trajectories on demand for the current task with targeted and fine-grained access to experience knowledge. To this end, we propose ExpVoyager, a novel framework in which a skill curator navigates raw experience across different views and resolutions, continually identifying reusable procedural knowledge from what it observes while tracking remaining knowledge needs that guide where to navigate next. Extensive experiments demonstrate both the effectiveness and versatility of ExpVoyager, showing consistent improvements in downstream task performance, continual gains as the experience space scales, and practical compatibility with existing skills under efficient experience access.