Research
EVOKE: Eliciting World Knowledge in Agents for Transferable Decision-Making
Overview Research area: Natural language processing and LLM-based agents — specifically post-training methods for multi-step decision-making in interactive text environments, and generalization to uns

- arXiv
- 2609.38334
- Published
- 2026-09-29
- Authors
- Yuhan Guo, Jinming Liu, Liang Xu, Ziqiang Li, Jianguo Huang, Zhicheng Wang, Hu Zhu, Qiuyu Chen, Yuntao Wei, Xin Jin, Wenjun Zeng
AI summary
Overview
- Research area: Natural language processing and LLM-based agents — specifically post-training methods for multi-step decision-making in interactive text environments, and generalization to unseen environments.
- Technical level: Advanced. The paper assumes familiarity with reinforcement learning from preferences, contrastive ranking losses, LoRA fine-tuning, imitation learning dataset aggregation (DAgger), and world-model approaches for agents.
- Scope in one sentence: The paper proposes EVOKE, a post-training method that trains a language-model policy to rank candidate actions under alternative goals at fixed states, in order to elicit world knowledge already present in pretrained models rather than teaching it through a separate prediction objective.
What This Paper Is About
LLM agents perform well in the environments they are trained in but transfer poorly to environments they have not seen. The usual fix is world-model training: teach the agent to predict the consequences of its actions. The authors argue that for agents operating in digital environments such as websites and search engines, this consequence knowledge is already internalized during pretraining, so the real problem is not acquiring it but getting the policy to actually use it at decision time. EVOKE supplies that pressure by holding the environment state and interaction history fixed while swapping in different goals and training the policy to rank the same candidate actions correctly under each goal.
Key Contributions
- A reframing of world knowledge for LLM agents. The paper argues the problem should be treated as eliciting knowledge already in pretrained models rather than acquiring it through an additional prediction objective, which costs extra training and compounds errors when predicted futures are used for planning.
- The EVOKE method. A post-training procedure that collects states from the policy's own rollouts, poses alternative goals at those states, executes and assesses the same candidate actions under each goal, and trains the policy by contrastive ranking — repeatedly, in a DAgger-style aggregation loop.
- Empirical evaluation across three backbones and three task families. Text-based household tasks (ALFWorld), web navigation (WebShop), and search-based QA, with reported gains in task performance, generalization to unseen environments, and data efficiency, plus comparisons against world-model and dynamics baselines.
- Controlled analyses of the mechanism. Ablations isolating goal interventions and ranking, data-efficiency curves, iterative-training curves, linear probes testing whether action consequences are already decodable, and a goal-pair diagnostic measuring "habitual errors."
Main Findings
- Best average on every benchmark and backbone: EVOKE is reported as achieving the best average on every benchmark and backbone tested, including all world-model and dynamics methods on Qwen2.5-7B-Instruct, and it remains the best on unseen ALFWorld games. Margins are largest on ALFWorld and WebShop, which require multi-step interaction, and on the smallest backbone, Qwen3-1.7B.
- ALFWorld seen games (avg over Pick/Look/Clean/Heat/Cool/Pick2): 91.4 for Qwen2.5-3B-Instruct, 96.0 for Qwen2.5-7B-Instruct, and 91.1 for Qwen3-1.7B. For comparison on the 3B backbone, SDAR reaches 80.1 and RLSD 78.4; on 7B, the strongest listed baseline is EnvRL-GiGPO at 95.2.
- ALFWorld unseen games (avg): 91.1 for Qwen2.5-3B-Instruct (vs PCSD 85.9 and AHEAD 82.6), 93.1 for Qwen2.5-7B-Instruct (vs AHEAD 83.8), and 84.6 for Qwen3-1.7B (vs GRSD 84.0).
- WebShop: EVOKE reports Score 87.4 / Succ 82.8 on Qwen2.5-3B-Instruct, 90.6 / 85.9 on Qwen2.5-7B-Instruct, and 88.4 / 80.5 on Qwen3-1.7B.
- Search-based QA (Avg over NQ, TriviaQA, PopQA, HotpotQA, 2WikiMultiHopQA, MuSiQue, Bamboogle): 45.3 on Qwen2.5-3B-Instruct, 49.2 on Qwen2.5-7B-Instruct, and 43.9 on Qwen3-1.7B. The paper notes this is where baselines are closest because each question involves only a few retrieval decisions, but EVOKE still leads the strongest baseline on each backbone.
- Goal interventions drive the gain: Source only, which uses the same ranking objective but only the original goals, drops from 91.8% to 82.8% unseen success and needs 40.7% more actions per game. Source replay repeats the original-goal data until it matches EVOKE's updates and still reaches only 85.9%, taking 34.4% more actions.
- Ranking beats imitation, so more goal data is not the explanation: Positive SFT receives exactly the same goal contexts, positive actions, and updates as EVOKE but imitates positives instead of ranking them; it nearly matches on seen games but falls 7.5 points behind on unseen games.
- Data efficiency: Across nested subsets of 20–100% of labeled states with three training orders each, EVOKE outperforms SFT in all 15 paired runs, and with 60% of the states already exceeds SFT trained on all of them.
- Iteration keeps improving the policy: Over two rounds of aggregation, unseen success rises by 20.1–31.3 points on every backbone, with a clear gain in every round. On Qwen2.5-3B, the two rounds together supervise about 9.1% as many decisions as the demonstrations behind the initial SFT policy (1,482 supervised decisions vs 16,247 demonstration steps).
- The knowledge is already there: Linear probes predict five consequences (whether the target object is held, lies in the receptacle named by the action, or is hot, clean, or cool) with almost perfect accuracy in every model, including the original Qwen2.5-3B backbone before any ALFWorld training. On transitions where the same action text leads to different outcomes, a probe on the action text is at chance while model representations remain above 97.0%, against 65.0–68.0% for the same architecture with random weights. Yet the initial policy succeeds on only 60.4% of unseen games.
- EVOKE turns knowledge into goal-directed decisions: On 812 goal pairs constructed from the 134 unseen games (where two goals share state, history, and available actions but have disjoint progress-making actions, and random choice succeeds on 0.3% of pairs), the initial policy solves only 19.5% and more than half its errors are habitual. EVOKE makes the fewest habitual errors at 6.4% of decisions, 4.3 points fewer than Positive SFT (95% CI [3.0, 5.6]).
- Better decisions, not more trial and error: EVOKE solves 91.0% of unseen games within 20 actions, versus at most 80.6% for the other trained variants, and needs 24.4–28.9% fewer actions per game. On the 70 games that all five variants solve, EVOKE still takes the fewest actions, makes less than half as many invalid actions as Positive SFT, and about a third as many revisits as the other trained variants.
Methodology in Plain English
The method starts from a policy obtained by supervised fine-tuning on successful demonstration trajectories broken into per-step input–action pairs (16,247 examples from 2,700 training games). The training loop then repeats four steps:
- Collect decision states. Let the current policy interact with training environments, recording the goal and visible history at each step so the same state can be revisited. States are sampled both where the policy makes progress and where it detours, repeats actions, or heads toward failure.
- Intervene on goals. For each collected state, keep the original goal and add alternative goals that are achievable from the same state. An LLM annotator proposes candidates, and each is kept only if it is compatible with the scene, not yet satisfied, and achievable. Crucially, the state, visible history, and list of available actions stay identical across goals.
- Assess actions against actual outcomes. Execute each candidate action from the same state and record what actually happens. The annotator then labels actions that advance the goal as positives, actions that are worse as competitors, and leaves actions with insufficient evidence unlabeled. Because the annotator judges observed outcomes rather than predicting them, labels reflect how the environment really responds.
- Train by contrastive ranking. The policy scores an action by the mean token log-probability of its completion. The loss combines a listwise term that raises the probability mass of the whole positive set with a pairwise term that separates every positive from every competitor, weighted by λ = 0.25 with temperature τ = 1. Competitors are chosen to be "hard negatives" — actions the policy itself favors but the annotator judges worse.
LoRA adapters are updated while the backbone stays frozen. Each round repeats the whole procedure with the current policy and trains on the aggregate of all rounds, following dataset aggregation in imitation learning, so hard negatives stay aligned with the current policy's mistakes. At deployment the result is an ordinary policy: no world-model module, no inference-time planning, and no annotator. The same pipeline applies to WebShop and search-based QA, with goals, states, and actions taking the form appropriate to each environment.
Evaluation uses ALFWorld Seen (140 games) and Unseen (134 games) with a 50-action limit and greedy free-form generation without legal-action masking or output repair; search-based QA accuracy over seven datasets; and WebShop score and success rate. Results are averaged over training seeds, with micro success (fraction of successful games) and macro success (unweighted mean over the six task types) reported.
Why This Matters
Impact on research. The paper challenges a common assumption in agent research: that improving transferability requires adding a prediction objective. It argues that for LLM agents in digital environments, the consequence knowledge is already encoded and measurable (the probe results support this), and that the bottleneck is whether single-goal supervision gives the policy any reason to use it. This reframes the design space from "what to predict" to "how to supervise decisions," and the controlled analyses aim to show the gain comes from the supervision structure rather than from more data or more training. The result is also comparatively cheap: the two rounds supervise about 9.1% as many decisions as the initial demonstration set.
Real-world applications (drawn from the environments the paper evaluates):
- Web navigation and shopping agents that must complete multi-step purchase or search tasks on sites they were not trained on.
- Search-based question answering systems that issue retrieval queries and compose answers across single-hop and multi-hop questions.
- Assistants for text-based household or task environments, where the same observed state may call for different actions depending on the user's current goal.
- Any deployed agent that must follow changing user instructions at a fixed interface state, which is exactly the setting the paper's goal-intervention diagnostic models.
Industry relevance. The method needs no separate world model and no inference-time planning, so it adds no runtime cost at deployment — attractive for production agents. The data-efficiency result (matching full-data SFT with 60% of labeled states) matters because the expensive part of this pipeline is executing and annotating candidate actions. The pipeline does rely on an LLM annotator and environment execution during training, which the paper notes is not used at deployment.
Future Directions
- Extension to physical and embodied environments. The paper explicitly says learning consequences through prediction is important in physical environments, where dynamics such as contact and motion are hard to capture without grounding; whether EVOKE-style elicitation helps there is not tested.
- Reducing the cost of the annotation pipeline. Goal proposal and outcome assessment currently rely on an LLM annotator plus real environment execution; the paper reports data efficiency gains but does not report a method that removes the annotator.
- Understanding how rich the alternative-goal set must be. The theory the authors cite requires competence across a "sufficiently rich" set of goals; the paper does not report a study of how goal diversity or goal quality trades off against performance.
- Extending beyond the three evaluated task families. The paper evaluates ALFWorld, WebShop, and search-based QA on three backbones; transfer to other digital environments is not reported.
Target Audience
Researchers and practitioners working on LLM agents, agent post-training, and reinforcement learning or preference-based optimization for language models. It is most useful for readers already comfortable with contrastive ranking losses, LoRA fine-tuning, and imitation learning aggregation, and for engineers building deployable agents in web navigation or retrieval settings who care about generalization to unseen environments and about inference-time cost. Readers looking for an introduction to agent training or for embodied-robotics results will find this paper assumes substantial background and does not address physical environments.
Authors’ abstract
Large language models (LLMs) are increasingly deployed as agents for multi-step decision-making, yet transfer poorly to unseen environments. World-model methods address this by training agents to predict future observations, at the cost of additional training and errors that compound when predictions are used for planning. However, for LLM agents operating in digital environments, much of this world knowledge is already internalized during pretraining, which shifts the problem from acquiring it to eliciting it. We argue that typical post-training provides little pressure for such elicitation, since supervision under a single goal at each visited state inadvertently drives policies to rely on superficial contextual habits. We introduce EVOKE, a post-training method that supplies this pressure through goal diversity at fixed states. Motivated by theory showing that an agent competent across diverse goals must encode a world model recoverable from its action preferences, EVOKE holds the environment state and interaction history fixed and ranks the same candidate actions under alternative goals, forcing action preferences to change, so that a policy relying on contextual habits or single-goal correlations cannot order them correctly. This implicitly elicits the policy's pretrained world knowledge to inform decisions. We evaluate EVOKE across diverse tasks in three backbones, demonstrating improved task performance, unseen environment generalization, and data efficiency. We further conduct controlled analyses to better understand what drives these gains. These findings offer a new perspective on eliciting internalized world knowledge for transferable action through direct decision supervision.