Skip to content
AI.info

Research

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Overview Research area: Security and training of LLM-based web agents, specifically defenses against indirect prompt injection, using reinforcement learning inside a learned web world model. Technical

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model
arXiv
2610.08773
Published
2026-10-06
Authors
Sarim Hashmi, Mukul Ranjan, Kshitij Mishra, Mikhail Kuznetsov, Praneeth Vepakomma, Nils Lukas

AI summary

Overview

Research area: Security and training of LLM-based web agents, specifically defenses against indirect prompt injection, using reinforcement learning inside a learned web world model.

Technical level: Advanced. The paper assumes familiarity with reinforcement learning (group-relative advantages, low-rank adapters), LLM agents, and adversarial training.

Scope: The paper presents AdvSim2Real, a two-stage framework that co-evolves a task-generating curriculum, an injection adversary, and a web-agent executor inside a frozen web world model, and measures whether this produces a Qwen3.5-4B agent that is both more capable and more robust.

What This Paper Is About

Web agents complete user requests by reading and acting on pages written by third parties, so an instruction planted in page content can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. AdvSim2Real addresses this by training the agent, the task generator, and the attacker against each other inside a frozen web world model, so that both the tasks and the attacks keep adapting to the agent's current skill level.

Key Contributions

  1. AdvSim2Real, a two-stage framework that co-evolves a task curriculum, an injection adversary, and a web agent inside a frozen web world model, so that both tasks and attacks track the current agent.
  2. The success-flip reward, which replays an accepted clean run to the injection step and credits the adversary only when the continuation fails, so failures the agent makes on its own earn nothing.
  3. A benchmark of 150 form-filling web tasks in five skill strata, with protected fields and forbidden controls, a reactive adversary that chooses when and what to inject, and a deterministic browser check of the submitted form; the authors release it with all checkpoints and trajectories.
  4. Evidence that the curriculum stage buys capability and the adversarial stage buys robustness: without the curriculum stage, the final agent loses 4.67 clean points but only 0.44 points of attacked completion.

Main Findings

  • Clean capability improves inside the world model. On the 150 tasks in WebWorld-14B, clean judged completion rises from 74.89% (Base) to 78.00%, 77.11%, and 79.33% across the three Stage-1 rounds, and the final checkpoint reaches 81.33%.
  • Robustness improves under the learned adversaries. The attack mean over Adv v1–v3 rises from 48.07% (Base) to 54.74% after the first Stage-1 round (which never saw an injection), then to 57.48% at Robust iter 3. Stage 2 alone adds 3.14 points under attack relative to Capability iter 3 while also raising clean completion by 2.00 points.
  • The adversarial stage is not free at first. The first Stage-2 round lowers both metrics, by 0.64 points under attack and 2.00 points clean; the second and third rounds recover and exceed the starting point.
  • The attacked gain concentrates on the first adversary: 5.78 points against Adv v1, 2.54 against Adv v2, and 1.11 against Adv v3.
  • Robustness holds against an adversary that never trained against the agent. Under Kimi-K3, a frontier model, the base agent falls from 74.89% clean to 23.00%, and Robust iter 3 reaches 30.72%, a 7.72-point absolute gain and a 33.6% relative gain. The three Stage-2 rounds lie within 1.05 points of one another.
  • Capability gains transfer to a real browser. Without any world-model call, strict deterministic browser success rises from 25.56% (Base) to 31.78%, 43.56%, and 44.44%, and correct final field values from 52.50% to 73.24%. The share of tasks solved in both the world model and the browser rises from 23.1% to 42.9%, while the share judged good only in the world model falls from 51.8% to 36.4%.
  • A large gap remains. A 23.85-point clean-to-attacked gap persists at the final checkpoint, and every checkpoint still loses at least 47.66 points of clean completion to a frontier-model adversary limited to one injection per trajectory.
  • The 4B checkpoint matches a hosted 9B model. On matched task and seed identities, Robust iter 3 reaches 57.40% under attack against Qwen3.5-9B's 56.96%, and 81.33% clean against 78.22%, a 3.11-point clean advantage; the 0.44-point attacked difference is smaller than the 4B checkpoint's 0.56-point seed spread.
  • Removing Stage 1 mostly costs capability, not robustness. The ablation reaches 76.67% clean and 57.04% under attack; the complete pipeline leads clean by 5.78, 2.00, and 4.67 points across rounds, but leads under attack by only 0.89, trails by 0.52, then leads by 0.44.
  • Kimi-K3 attacks more aggressively than the learned adversaries: it requests an injection in 95.0% to 98.0% of episodes versus 66.5% to 81.0%, and the world model renders the requested injection in 72.3% to 80.9% of episodes.
  • Base fails by running out of budget. Under Kimi-K3, Base exhausts the 12-action budget without a final message in 28.3% of episodes, against 14.0% to 15.4% for the Robust checkpoints and 3.7% of clean episodes pooled over all checkpoints.
  • Stratum-specific gains vary. Robust iter 3 gains most over Base on conditional-policy (+14.3 points) and conjunctive-selection (+12.8) tasks and least on relational joins (+1.6); conjunctive selection remains the hardest stratum at 40.6%.
  • Compute is only partially reported. One recorded Stage-1 iteration reserved two 96 GB RTX 6000 Pro GPUs for 20 h 46 min, or 41.54 allocated GPU-hours; token counts and API charges for the complete sequence were not measured.

Methodology in Plain English

The authors build a simulator: a frozen world model (WebWorld-14B) that predicts the next web page given the goal, the page history, and an action, and can also render a requested injected notice in the style of the current page. A frozen language-model judge (Qwen3.8-27B) decides whether a trajectory completed the task. Inside this simulator, three policies are trained against each other.

In Stage 1, a curriculum policy writes tasks, and the executor attempts them several times. The curriculum is rewarded for tasks the executor solves about half the time — the uncertainty reward borrowed from R-Zero, which peaks at a success rate of one half and vanishes at 0 or 1 — with a penalty for lexically similar proposals. The executor then trains on fresh tasks from the updated curriculum, with advantages scaled so that tasks near its competence frontier get more weight.

In Stage 2, the curriculum is frozen and an adversary joins. The adversary does not simply try to make the agent fail; it is rewarded only for a success flip: the authors save clean runs the judge accepted, replay one to a chosen step, insert one injection there, and credit the adversary only if the continuation fails. The executor trains on a mixture of new attacks, attacks from earlier rounds, and clean tasks, so an injection that stops working is not forgotten. Each role trains a separate low-rank adapter. Each stage runs three rounds; checkpoints are named Capability iter 1–3 and Robust iter 1–3, with Adv v1–v3 being the saved adversaries.

Evaluation uses 150 form-filling tasks from ten templates in five skill strata (conditional policy, derived values, conjunctive selection, relational join, ordered repair). Each task specifies required final field values, two protected fields that must stay unchanged, and two forbidden controls, Save draft and Reset all. Attacked results are judge verdicts inside the world model, where injections are rendered; the capability checkpoints are additionally run in a real Chromium browser with live DOM observations and a deterministic check of the submitted form.

Why This Matters

Impact on research. The paper argues that defenses trained on injections fixed before training never meet an adaptive attacker, and that existing adversarial-training methods adapt the attack but keep the tasks fixed, so a task stops teaching once the agent solves it. It shows that a world model can remove both limits at once, and that a 4B agent trained this way can match a hosted 9B model. It also supplies a benchmark, checkpoints, and trajectories for evaluating injection defenses against adversaries trained on the defended agent.

Real-world applications:

  • Back-office form completion, such as updating a customer record from values shown on a page, where a planted notice could otherwise order the agent to discard its work.
  • Any agent that reads third-party content while holding an authorized goal, where the page must be used for data and controls but must not be obeyed as an instruction source.
  • Browser-automation deployments where a deterministic check of the submitted form is the real success criterion, not a model's judgment.
  • Red-teaming workflows: the released adversarial policies let defenders test agents against an attacker that adapts to them.

Industry relevance. The endpoint gains are 6.44 clean points and 9.41 attacked points on the same 150 tasks, and the capability gain survives the sim-to-real jump to a real browser. The paper's threats and mitigations are directly relevant to agent providers who must keep authorized completion working despite competing page instructions.

Future Directions

  • Close the gap to a real adaptive attacker. The training adversary proposes from the initial page, while the evaluation adversary reacts to the trajectory; the authors state that a procedure hardening web agents against injections should ultimately be evaluated against an adversary that adapts to the trained agent.
  • Move beyond judge verdicts to executable task completion. The world model can display a value the executor never entered and the judge can accept a wrong calculation; only the capability checkpoints were verified in a browser, and a judged robustness gain is not yet a gain in executable task completion.
  • Isolate the mechanisms. The paper attributes gains to the checkpoint sequence as a whole and states it did not isolate any of its three intended mechanisms — zero curriculum reward at judged rates 0 and 1, the success-flip reward's exclusion of self-inflicted failures, and historical attack inputs with fresh labels.
  • Run compute-matched component controls, blinded human audits, and executable task checkers, and measure end-to-end cost, since token counts and API charges for the complete sequence were not measured.

Target Audience

Researchers and engineers working on LLM agent security, reinforcement learning for agents, and web automation. It suits readers who already understand adversarial training and policy optimization and want a specific, benchmarked design for co-evolving tasks, attacks, and defenders inside a simulator. Practitioners deploying browser agents will find the task-preserving robustness framing and the deterministic browser results directly applicable.

Authors’ abstract

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a task stops teaching once the agent solves it. We introduce AdvSim2Real, which co-evolves a task curriculum, an injection adversary, and the agent inside a frozen web world model. The curriculum is rewarded for tasks the agent solves about half of the time, and the adversary only for a success flip, an injection that turns a judged success into a failure. Training in the simulator makes a 4B agent both more capable and more robust: its completion rises with and without attacks, holds against a frontier-model adversary it never trained against, and its capability gain carries over to a real browser. On 150 web tasks, AdvSim2Real raises completion under this unseen adversary by 33.6\% relative to the base agent.

Read the original paper