Skip to content
AI.info

Research

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses

Overview Research area: LLM agent self-evolution, agent runtime safety, and reversibility of model-generated system modifications. Technical level: Advanced. The paper uses formal state notation, type

EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
arXiv
2608.28363
Published
2026-08-28
Authors
Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah

AI summary

Overview

Research area: LLM agent self-evolution, agent runtime safety, and reversibility of model-generated system modifications.

Technical level: Advanced. The paper uses formal state notation, typed observational-equivalence contracts, a 2×2 factorial design, and exact McNemar tests with Holm–Bonferroni correction.

Scope: The paper defines and empirically studies "recoverability-constrained self-evolution" — the requirement that capability-improving modifications an LLM agent makes to its own harness must be verifiably reversible across counterfactual states, not just the state in which the mutation was drafted.

What This Paper Is About

LLM agents increasingly rewrite their own prompts, tools, middleware, routing, files, and resources at runtime, and most self-improvement procedures only check whether a change improves the forward task objective. The problem is that an improvement can overwrite configuration, reorder middleware, shadow tools, or leak resources, and a static inverse operation may be unable to restore the earlier state — especially in states different from the one the mutation was created in. EvoUndo is a framework that represents, synthesizes, diagnoses, and independently verifies recovery for model-generated self-modifications before they are permanently admitted.

Key Contributions

  1. Formalizes recoverability-constrained self-evolution and introduces EvoUndo, which couples harness mutations with witness capture, counterfactual verification, typed diagnosis, and closed-loop recovery synthesis (the forward mutation m is locked; only the witness w, recovery program u, and effect contract C_e may be revised).

  2. Establishes a natural failure benchmark from 600 self-evolution tasks, identifying 197 capability-positive recovery failures where standard iterative repair achieves 0/197 success under the base recovery representation.

  3. Decouples natural recovery failures into grounding and expressivity bottlenecks: state-grounded diagnostic feedback recovers 38/48 (79.2%) failures in the S_0 stratum, while an extended recovery calculus enables recovery on 142/143 (99.3%) failures in the S_1 stratum.

  4. Observes a non-monotonic interaction on the primary configuration between diagnostic granularity and recovery-language capacity: state-grounded feedback degrades recovery under L_1 on gpt-oss-120b, while this negative interaction did not reproduce under the Qwen3.8-27B constrained-decoding replication configuration.

Main Findings

  • Controlled repairs work when the representation fits: On a controlled benchmark of 120 deliberately corrupted but recoverable mutations at budget B=4, independent regeneration recovers 4/120 (3.3%), generic verifier feedback 104/120 (86.7%), raw verifier traces 101/120 (84.2%), typed diagnosis 114/120 (95.0%), and prescriptive hints 117/120 (97.5%).

  • Natural failures break conventional repair: On the frozen 197-task natural failure cohort, all four verifier-guided repair modes achieve 0/197 (0.0%) recovery at B=4 under the base L_0 representation, while independent regeneration recovers 6/197 (3.0%). Baseline taxonomy classification placed 180/197 natural failures outside the predefined controlled defect taxonomy; the overlap with the 180/197 tasks rescued by D_0L_1 is 164 tasks, with 16 taxonomy-only and 16 rescue-only tasks.

  • Expressivity is a hard limit before grounding is: A zero-generation constructive oracle audit shows only 48/197 mutations are oracle-recoverable under L_0, rising to 191/197 under the extended language L_1. This partitions the bank into S_0 (|S_0|=48) and S_1 (|S_1|=143); six tasks remain unrecovered because of implemented oracle heuristics and are excluded from stratum contrasts.

  • C1 — grounding unlocks S_0: On S_0, D_0L_0 yields 0/48 while D_1L_0 yields 38/48 (79.2%), a paired risk difference of +79.17 pp (95% bootstrap CI [+66.67, +89.58]), exact McNemar p_raw = 2^-37 = 7.28 × 10^-12 (p_Holm = 1.46 × 10^-11).

  • C2 — expressivity unlocks S_1: On S_1, D_0L_0 yields 0/143 while D_0L_1 recovers 142/143 (99.3%), a paired risk difference of +99.30 pp (95% CI [+97.90, +100.00], p_Holm = 1.08 × 10^-42).

  • C3 — a non-monotonic interaction on gpt-oss-120b: On S_1, adding state-grounded feedback to the rich language drops recovery from 142/143 (99.3%) under D_0L_1 to 133/143 (93.0%) under D_1L_1 (Δ = −6.29 pp, 95% CI [−11.19, −2.10], p_Holm = 0.0117). The full-cohort interaction is I = (D_1L_1 − D_0L_1) − (D_1L_0 − D_0L_0) = −25.89 pp (95% CI [−34.01, −18.27]). Trace analysis of the 11 discordant tasks suggests changes in semantic decomposition and operation ordering rather than systematically longer recoveries; the paper labels this analysis descriptive rather than causal.

  • Full factorial table (N=197, B=4): D_0L_0 0/197 (0.0%), 0/48, 0/143; D_1L_0 38/197 (19.3%), 38/48 (79.2%), 0/143; D_0L_1 180/197 (91.4%), 38/48 (79.2%), 142/143 (99.3%); D_1L_1 167/197 (84.8%), 34/48 (70.8%), 133/143 (93.0%).

  • End-to-end 600-task accounting: 478/600 (79.67%) produced capability-positive mutations — 281 already admissible and 197 recovery-defective (197/478, 41.2% of capability-positive). Under D_0L_1 at B=4, EvoUndo rescues 180/197 (91.37%), yielding 461/478 (96.44%) admissible capability-positive mutations and 461/600 (76.83%) total benchmark yield, a +30.00 pp absolute increase over the initially admissible 281/600 (46.83%). Zero forward-capability regressions were introduced by recovery repair because m_0 is executed verbatim.

  • Cross-model replication: On a frozen 96-task Qwen3.8-27B constrained-structured-decoding replication, C1 is +47.92 pp (p < 0.001) and C2 is +37.50 pp (p < 0.001), while C3 is +6.25 pp (95% CI [−2.08, +16.67], p_adj = 0.375) — the negative interaction did not reproduce. Model, decoding protocol, and sample size all differ, so the cause is not isolated.

  • Fresh-holdout precision–coverage trade-off: On 300 fresh tasks (50 per family) with 100 hidden states, increasing diverse development states from Q_2 to Q_20 reduced observed false positives from 16.8% to 0.0% but lowered Rescue@4_100 from 49.7% to 35.0% (Δ = −14.67 pp, p = 5.76 × 10^-6) and increased the conservatism gap from 6.6% to 24.1%. Under Q_20, Listeners had 0/50 development passes but 32/50 hidden-criterion passes.

  • Effect-contract auditing: In a controlled audit of 300 intentionally incomplete effect contracts across all six harness families, independent dynamic effect tracking detected every injected undeclared semantic effect (300/300) with no false positives on 300 matched complete contracts.

  • Cost: Fresh-holdout inference ranged 2.39–3.09 calls and 3.37k–4.30k generated tokens per task; the final strict-L_0 D_1L_0 run averaged 3.51 calls/task.

  • Snapshot baseline comparison: Effect-scoped snapshots are the stronger baseline in the evaluated serializable in-memory selective-undo regime, recovering 300/300 tasks versus 243/300 for EvoUndo under different-surface changes and 159/300 versus 131/300 under same-surface changes. There are no EvoUndo-only wins; superiority in settings where exact effect-scoped restoration is unavailable is not established.

Methodology in Plain English

The researchers model an agent harness as a persistent state plus an execution policy, and a mutation as a function that maps one state to another. A candidate self-modification is packaged as a four-part object: the forward mutation, a witness program that records prior values and existence flags before the mutation runs, a recovery program that uses that witness plus a mutation-produced resource receipt, and a declared effect contract that says which state surfaces the recovery must cover.

For each candidate, the system runs a counterfactual round trip — capture witness, apply mutation, run recovery — across development states and hidden states, and checks whether the recovered state is observationally equivalent to the pre-state under typed equality rules (sorted keys, canonicalized pointers, SHA-256 file hashes, and so on). The forward mutation is immutable: only the witness, recovery program, and effect contract can be revised over up to B = 4 synthesis turns. The runtime independently computes the actual effect set and requires the declared contract to cover it, so shrinking the contract cannot be used to evade verification, and the contract cannot be omitted.

Recovery is bounded by two formal recovery languages. L_0 covers configuration, prompt, tool-registry, and routing state; L_1 extends it with middleware sequences, event listeners, sandboxed files, managed sockets, and ordered (LIFO) composition of recovery operations. The team built a controlled benchmark of 120 injected defects, a natural failure bank from 600 single-sample zero-shot tasks across six 100-task families, and a deterministic constructive oracle audit to establish which failures are recoverable in principle under each language. A protocol-locked 2×2 factorial then crossed diagnostic granularity (coarse typed versus state-grounded bundle) with recovery language (L_0 versus L_1), with three pre-registered paired contrasts. Admissibility used a split-Wald rule requiring the normal-approximation lower endpoint at τ_R = 0.85 to pass independently on IID and OOD hidden splits of n = 20 each, so at least 19/20 successes are needed per split.

Why This Matters

The work reframes agent self-improvement: forward capability gain alone is insufficient for long-lived autonomous systems, and recovery correctness is a relational property that depends jointly on the state distribution, witness semantics, observational contract, and recovery-language expressivity. It shows that when recovery representations align with the underlying state structure, LLMs can synthesize state-dependent recovery programs across counterfactual states — but also that adding the most granular diagnostics is not universally better.

Real-world applications:

  • Agent platforms that let models edit their own tools and middleware, where a bad edit must be undone without corrupting later state.
  • Autonomous coding and DevOps harnesses that modify routing, configuration, or sandboxed files and need auditable rollback.
  • Resource-managing agents that allocate sockets, listeners, or files and must release exactly what they introduced.
  • Safety and compliance review pipelines that need independently verified evidence that a persistent modification is reversible before admission.

Industry relevance: the paper argues for co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting. It also reports that straightforward effect-scoped snapshots outperform EvoUndo in the evaluated serializable in-memory selective-undo regime, which is useful guidance for teams deciding whether to build snapshot infrastructure or recovery synthesis.

Future Directions

  • Joint optimization of the forward mutation and recovery, since this study holds m = m_0 fixed and never tests capability-equivalent but more reversible forward mutations.
  • Adaptive diagnostic strategies that escalate from coarse indicators to granular state traces only when coarse repair fails, since no diagnostic curriculum or address-use constraint was evaluated and the paper does not establish causality for the negative C_3 interaction.
  • Resolving the precision–coverage trade-off, where strict development gates (20/20 round trips under Q_20) drove false positives to 0.0% while cutting Rescue@4_100 to 35.0% and leaving Listeners at 0/50 development passes but 32/50 hidden-criterion passes.
  • Extending the modeled state surfaces — distributed and external state, third-party APIs, and unmanaged processes are unmodeled — and adding compensation for irreversible effects. Only gpt-oss-120b and Qwen3.8-27B were evaluated, and no low-, medium-, or high-reasoning-effort comparison was performed.

Target Audience

Researchers in LLM agent architectures, self-improving and self-evolving systems, and reversible or bidirectional computation; systems engineers building agent runtimes with rollback, undo, or effect-scoped snapshot infrastructure; and safety, evaluation, and governance practitioners who need verified guarantees that persistent model-generated modifications can be reversed. The formal notation, factorial design, and statistics make it most accessible to readers with a background in agent systems or program verification.

Authors’ abstract

LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail recoverability verification. Under the original recovery representation, conventional repair strategies recover 0/197 of these natural failures. Deterministic oracle analysis recovers 48/197 under the original recovery language L0, while the extended recovery calculus increases empirical oracle recovery to 191/197. A protocol-locked 2x2 grounding-by-expressivity intervention then separates two bottlenecks: exact state-address grounding increases successful recovery from 0/48 to 38/48 (79.2%) when the original language is sufficient, while extending the recovery language enables recovery on 142/143 (99.3%) failures in the oracle-defined S1 stratum. On the primary gpt-oss-120b backbone, adding exact-address diagnostics to the richer language reduces recovery to 133/143 (93.0%); a Qwen3.8-27B replication preserves the grounding and expressivity effects but not this negative interaction, indicating that the latter is model-dependent. These results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.

Read the original paper