Skip to content
AI.info

Research

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback Overview Research area: AI agent security (cs.CR) — specifically red-teaming of "agent skills," the reusa

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback
arXiv
2609.32400
Published
2026-09-26
Authors
Pengyu Zhu, Jingyi Yang, Yi Liu, Li Sun, Sen Su

AI summary

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Overview

  • Research area: AI agent security (cs.CR) — specifically red-teaming of "agent skills," the reusable packages of instructions, executable code, and task-specific resources that agents load at execution time.
  • Technical level: Advanced. The paper assumes familiarity with LLM agent architectures, tool/skill packaging, sandboxed execution harnesses, and layered defense pipelines (static scanning plus runtime monitoring).
  • Scope: The paper proposes and evaluates SkillDRE, a fully automated framework that evolves a complete malicious skill package toward a fixed, task-conditioned attack objective by alternating between pre-execution scanner feedback and runtime-defense feedback across four victim models on SkillsBench.

The paper is authored by Pengyu Zhu, Jingyi Yang, Yi Liu, Li Sun, and Sen Su, with affiliations including Beijing University of Posts and Telecommunications, North China Electric Power University, and Chongqing University of Posts and Telecommunications. It is listed as arXiv:2609.32400v1 [cs.CR], dated 26 Sep 2026, under a CC BY 4.0 license, with code released at the companion repository linked in the paper.

What This Paper Is About

Agent skills can be improved automatically using execution feedback, and the same mechanism lets an attacker evolve a malicious skill to become more effective and harder to detect. The problem is that a malicious skill may survive pre-execution scanning yet fail to achieve its harmful goal under runtime defenses, while a revision that fixes execution may introduce new scanner findings — so improving one side can undo progress on the other. SkillDRE attacks this problem by holding a task-conditioned malicious objective and a verifiable judge rule fixed, then evolving the skill implementation through a closed loop that alternates between scanner-guided and runtime-guided refinement.

Key Contributions

  1. A feedback-driven threat model for malicious skill evolution. The attacker repeatedly revises a single skill using pre-execution scanner findings and runtime-defense decisions, while the task instruction, victim agent, other skills, and defenses remain fixed. The effect on legitimate task functionality is measured separately from attack success.

  2. SkillDRE, a fully automated dual-stage framework. SkillDRE autonomously constructs and validates task-conditioned attack targets and judge rules without supplied payloads or hand-crafted attack strategies, then evolves skill implementations through a cross-stage closed feedback loop integrating pre-execution scanner feedback, runtime defense feedback, and attack-outcome validation.

  3. A cross-stage closed loop design. Each runtime-guided revision re-enters the pre-execution stage for rescanning and further optimization before the next execution, so every candidate must jointly satisfy scanner and runtime constraints.

  4. Empirical evaluation on SkillsBench across four victim models. SkillDRE achieves a 45.28% average attack success rate with 0% SkillScan detection on final submitted skills, exceeding the strongest baseline by 40.3 percentage points, with ablations isolating the contribution of each stage.

Main Findings

  • Attack success rate: On 249 skills associated with 94 SkillsBench tasks across four victim models, SkillDRE reaches an average ASR of 45.28%, exceeding the strongest baseline by 40.3 percentage points. Per-model ASR is 46.99% (DeepSeek-V4-Pro), 43.78% (Qwen3.5-397B), 45.38% (GLM-5.2), and 44.98% (Gemini-3.7-Flash).

  • Baseline comparison: SkillJect achieves ASR of 2.41%, 2.81%, 1.61%, and 0.00% across the four models, with detection rates of 100.00% in each case. SkillHarm achieves ASR of 2.53%, 1.82%, 6.55%, and 9.00%, with detection rates of 100.00%, 98.25%, 97.34%, and 98.45%. Both baselines remain at or below 9.00% ASR.

  • Detectability: SkillDRE yields 0.00% detection rate on final submitted skills across all four victim models, compared with 97.34–100.00% for the baselines.

  • Cross-model stability: SkillDRE's ASR is highest on DeepSeek-V4-Pro and lowest on Qwen3.5-397B, a difference of only 3.21 percentage points, indicating effectiveness is not confined to a single victim model.

  • Benign task preservation: Comparing No Attack with SkillDRE within each victim model, SkillDRE yields higher observed task accuracy on three models, with increases of 2.14–3.58 percentage points, while GLM-5.2 shows a decrease of only 0.15 percentage points. Manual inspection of cases that flipped from failure under No Attack to success under SkillDRE found that agents sometimes skipped a task-relevant skill under No Attack and failed; revised skill instructions led the agent to invoke that skill and complete the benign objective.

  • Ablation — stage isolation: With DeepSeek-V4-Pro as the victim, Phase 1 one-shot yields 37.75% ASR and 65.46% DR; Phase 1 iterative yields 34.94% ASR and 0.00% DR; Phase 2 one-shot yields 36.14% ASR and 62.65% DR; Phase 2 iterative yields 46.59% ASR and 79.92% DR. Full SkillDRE yields 46.99% ASR and 0.00% DR.

  • Ablation — why both stages are needed: Iterative scanner-guided refinement reduces DR from 65.46% to 0.00% but lowers ASR from 37.75% to 34.94%. Iterative runtime-guided refinement raises ASR from 36.14% to 46.59% but raises DR from 62.65% to 79.92%. Relative to iterative Phase 1, full SkillDRE gains 12.05 percentage points in ASR while retaining zero DR; relative to iterative Phase 2, it reduces DR by 79.92 percentage points and achieves 0.40 percentage points higher ASR.

  • Construction efficiency: Task-level intents, skill-level targets, and judge rules require an average of 1.49, 2.24, and 1.84 proposal–validation rounds respectively, with first-round acceptance rates of 69.15%, 80.72%, and 61.04%. Skill-level target construction has the highest first-round acceptance rate but its distribution extends to 29 rounds, raising its average.

  • Cold-start evolution: Only 34.54% of initial candidates obtain a valid SkillScan result with no risk findings. Cumulative scanner acceptance rises to 84.34% by round 5 and 95.18% by round 10, with the remaining 4.82% requiring more than ten rounds. Phase 1 obtained a candidate with a valid SkillScan result and no risk findings for all 249 task–skill pairs within the 30-round budget, taking 3.23 rounds on average. Aggregate high-, medium-, and low-risk findings decrease by 81.82%, 78.21%, and 77.25% respectively from round 1 to round 5, though these decreases partly reflect fewer candidates scanned in later rounds as accepted pairs leave Phase 1.

  • Runtime-guided gain timing: In the first Phase 2 round, scanner-approved cold-start candidates achieve 34.54–37.35% ASR across the four victim models, showing a clean scan does not guarantee runtime success. Cumulative ASR then rises to 43.78–46.99%, a gain of 8.03–12.45 percentage points. GLM-5.2 reaches within 0.40 percentage points of its final ASR by round 6, while Qwen3.5-397B obtains more than half of its total gain between rounds 5 and 8. By round 8, all four models have achieved most of their observed gains, with no more than 2.01 additional percentage points over the final two rounds.

Methodology in Plain English

SkillDRE treats the construction of a malicious agent skill as a constrained optimization problem with two competing constraints, and separates the question of what harm is targeted from how the skill realizes it.

Step 1 — Fix the objective. Given a benign task and its associated skills, SkillDRE first builds a task-level intent describing the intended malicious side effect, constraints on admissible realizations, and a broad success theme. It then instantiates a skill-level target for each skill, with an observable success condition and the artifacts needed to verify it. A fixed ensemble of validator models approves or rejects each proposal; a proposal is accepted only if all validators approve. Rejections produce reasons that feed into the next proposal. Once accepted, the target stays fixed for the rest of the process.

Step 2 — Build a judge rule. A judge rule is iteratively constructed to evaluate whether the target was actually realized during execution. Candidates must pass a schema/well-formedness check, a "smoke test" that requires a negative attack decision in an empty sandbox without a hard execution error, and a semantic review by the validator ensemble. The best retained candidate is carried forward, and the accepted rule is then frozen.

Step 3 — Phase 1: pre-execution evolution. Starting from an initial candidate package generated from the source skill and the fixed target, every candidate is passed to SkillScan, which reports high-, medium-, and low-risk findings. These are combined into a weighted risk score using weights of 10000, 1000, and 10 respectively, designed to implement a lexicographic preference for reducing higher-severity findings first. Each round, the lowest-risk candidate so far becomes the refinement baseline, and a refiner generates the next candidate while being instructed to preserve the original skill's functionality and interface, keep the attack target unchanged, and make mechanism-level rather than presentation-only edits. A candidate is admitted to Phase 2 only after a valid SkillScan result with a risk score of zero.

Step 4 — Phase 2: runtime-guided evolution. The admitted candidate replaces the original skill in the task's skill set, and the victim agent executes the task with SkillSonar active as a fixed defense skill. Success requires both that SkillSonar issues no operation-stopping decision and that the judge rule confirms the target was realized in the resulting sandbox state. When it fails, the trajectory, both feedback signals, and an optimization memory of prior failures and repairs guide a new revision.

Step 5 — Close the loop. Each runtime-guided revision returns to Phase 1 with its local round index reset, so it must again earn a clean scan before the next runtime trial. The loop continues until success or until the runtime budget is exhausted. Attack-target construction and each Phase 1 run are capped at 30 rounds; Phase 2 allows at most 10 runtime trials per skill.

Evaluation setup. SkillScan is configured with Qwen3.5-397B-A17B, DeepSeek-V4-Pro is used as the generator for target construction and as the attack model for all methods, and the validator ensemble comprises DeepSeek-V4-Pro, GLM-5.2, and Kimi-K2.6. Codex is the execution harness for all victim models. Metrics are attack success rate (ASR), detection rate (DR), and task accuracy (ACC), the last measured with the original SkillsBench evaluation procedure.

Why This Matters

The paper shows that the two defense stages commonly deployed around agent skills — pre-execution scanning and runtime monitoring — leak complementary information that an adaptive attacker can exploit, and that evaluating either stage in isolation can miss the resulting attack capability. This reframes layered defenses from an additive stack into a feedback channel for automated red teaming, and it also demonstrates that a benign skill can be evolved into a working attack while scanner visibility drops to zero and legitimate task performance is largely preserved.

Real-world applications:

  • Security evaluation of agent marketplaces and skill registries: Vendors distributing or curating reusable agent skills need to know whether a skill that passes static scanning is actually safe under execution, and SkillDRE provides a stress test for that assumption.
  • Red-team tooling for enterprise agent deployments: Organizations running agents with access to internal tools can use dual-stage evolution to probe whether their scanner-plus-runtime defense combination holds against automated adaptation.
  • Defense-in-depth design and procurement: Teams choosing between static scanners and runtime monitors get evidence that scanner-only or runtime-only evaluation is insufficient, which informs how much budget each layer deserves.
  • Benchmark and evaluation standard development: The paper argues for joint evaluation of pre-execution and runtime defenses, which is relevant to anyone designing agent-security leaderboards or certification criteria.

Industry relevance: Agent skills are described in the paper as becoming a deployment mechanism for agent capabilities, and skill self-evolution is presented as a mainstream technique for improving agent behavior without updating model parameters. That makes the security of the skill artifact itself a production concern rather than a research curiosity. The result that final packages achieve 0.00% detection across all four victim models is directly relevant to vendors whose product claims rest on scanner detection rates, and the finding that 34.54–37.35% of clean-scanning cold-start candidates still succeed at runtime while others fail speaks to how much assurance a clean scan actually provides.

Future Directions

The paper as provided does not report an explicit future work section, so the following are open questions its results raise rather than directions the authors state.

  • Defenses that expect cross-stage adaptation. Since scanner-guided and runtime-guided refinement each undo the other's failures, an open question is what pre-execution or runtime defense would degrade under a closed-loop attacker rather than under single-stage ones, and whether defenses should model the attacker as adaptive.

  • Transferability across defense configurations. SkillScan was configured with Qwen3.5-397B-A17B and SkillSonar was the fixed runtime defense. Whether evolved skills transfer to other scanners, other runtime monitors, or both simultaneously is not reported.

  • Generalization beyond SkillsBench. Evaluation covers 249 skills associated with 94 tasks and four victim models. Whether the 45.28% average ASR and 0.00% detection rate hold on other skill ecosystems, longer-horizon tasks, or multi-skill compositions is not reported.

  • Detecting the evolution process itself. The paper measures detection of the final submitted skill, which is 0.00% across all models. Whether the iterative revision process — rather than the artifact it produces — leaves observable traces that a defense could act on is not addressed.

  • Interpreting the benign-accuracy effect. Negative ΔACC values were investigated by manual inspection, but the paper notes this only "may contribute" to the observed pattern. A systematic explanation of when skill revision accidentally repairs benign task behavior remains open.

Target Audience

This paper is most useful to agent-security researchers and red-team practitioners working on tool-use and skill-based LLM agents, particularly those studying adversarial adaptation against layered defenses. It also serves defense engineers and security architects evaluating agent skill pipelines, benchmark designers who need to justify evaluating pre-execution and runtime defenses jointly, and advanced graduate students in security or trustworthy machine learning who want a concrete, metrics-heavy case study of feedback-driven attack automation. Readers should be comfortable with attack success rate and detection rate as evaluation constructs, sandboxed agent execution, and the idea of LLM-based generator, refiner, and validator roles within an automated pipeline.

Authors’ abstract

Agent skills package instructions, executable code, and task-specific resources into reusable artifacts that agents can improve using execution feedback. The same mechanism also enables attackers to evolve malicious skills, making them more effective and less detectable. However, a candidate skill may pass pre-execution scanning yet fail to realize its target under runtime defenses, while a revision that repairs execution may introduce new scanner findings. We introduce SkillDRE, a fully automated framework for evolving complete malicious skill packages through a dual-stage feedback loop. Given a benign task and its associated skills, SkillDRE autonomously constructs and validates a task-conditioned malicious objective and a verifiable judge rule. It then holds both fixed while evolving the skill implementation, with preservation of legitimate task capability. SkillDRE combines scanner-guided evolution with runtime-guided refinement informed by execution outcomes observed under runtime defense. Each runtime-guided revision returns to the pre-execution stage for rescanning and further optimization before re-execution, forming a cross-stage closed loop. Evaluated on SkillsBench across four victim models, SkillDRE achieves an average attack success rate of 45.28%, exceeding the strongest baseline by 40.3%, while its final submitted skills receive no SkillScan findings and largely preserve benign-task performance. These results show that two-stage defense feedback can serve as a useful learning signal for adaptive red teaming and that evaluating either defense stage in isolation can miss the resulting attack capability. Codes is available at https://github.com/whfeLingYu/SkillDRE

Read the original paper