Research
EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses
Overview Research area: LLM agent self-evolution, agent runtime safety, and reversibility of model-generated system modifications. Technical level: Advanced. The paper uses formal state notation, type

- arXiv
- 2608.28363
- Published
- 2026-08-28
- Authors
- Tanmay Sah, Dolly Sah, Harshul Jain, Tanya Sah
AI summary
Overview
Research area: LLM agent self-evolution, agent runtime safety, and reversibility of model-generated system modifications.
Technical level: Advanced. The paper uses formal state notation, typed observational-equivalence contracts, a 2×2 factorial design, and exact McNemar tests with Holm–Bonferroni correction.
Scope: The paper defines and empirically studies "recoverability-constrained self-evolution" — the requirement that capability-improving modifications an LLM agent makes to its own harness must be verifiably reversible across counterfactual states, not just the state in which the mutation was drafted.
What This Paper Is About
LLM agents increasingly rewrite their own prompts, tools, middleware, routing, files, and resources at runtime, and most self-improvement procedures only check whether a change improves the forward task objective. The problem is that an improvement can overwrite configuration, reorder middleware, shadow tools, or leak resources, and a static inverse operation may be unable to restore the earlier state — especially in states different from the one the mutation was created in. EvoUndo is a framework that represents, synthesizes, diagnoses, and independently verifies recovery for model-generated self-modifications before they are permanently admitted.
Key Contributions
-
Formalizes recoverability-constrained self-evolution and introduces EvoUndo, which couples harness mutations with witness capture, counterfactual verification, typed diagnosis, and closed-loop recovery synthesis (the forward mutation
mis locked; only the witnessw, recovery programu, and effect contractC_emay be revised). -
Establishes a natural failure benchmark from 600 self-evolution tasks, identifying 197 capability-positive recovery failures where standard iterative repair achieves 0/197 success under the base recovery representation.
-
Decouples natural recovery failures into grounding and expressivity bottlenecks: state-grounded diagnostic feedback recovers 38/48 (79.2%) failures in the
S_0stratum, while an extended recovery calculus enables recovery on 142/143 (99.3%) failures in theS_1stratum. -
Observes a non-monotonic interaction on the primary configuration between diagnostic granularity and recovery-language capacity: state-grounded feedback degrades recovery under
L_1on gpt-oss-120b, while this negative interaction did not reproduce under the Qwen3.8-27B constrained-decoding replication configuration.
Main Findings
-
Controlled repairs work when the representation fits: On a controlled benchmark of 120 deliberately corrupted but recoverable mutations at budget
B=4, independent regeneration recovers 4/120 (3.3%), generic verifier feedback 104/120 (86.7%), raw verifier traces 101/120 (84.2%), typed diagnosis 114/120 (95.0%), and prescriptive hints 117/120 (97.5%). -
Natural failures break conventional repair: On the frozen 197-task natural failure cohort, all four verifier-guided repair modes achieve 0/197 (0.0%) recovery at
B=4under the baseL_0representation, while independent regeneration recovers 6/197 (3.0%). Baseline taxonomy classification placed 180/197 natural failures outside the predefined controlled defect taxonomy; the overlap with the 180/197 tasks rescued byD_0L_1is 164 tasks, with 16 taxonomy-only and 16 rescue-only tasks. -
Expressivity is a hard limit before grounding is: A zero-generation constructive oracle audit shows only 48/197 mutations are oracle-recoverable under
L_0, rising to 191/197 under the extended languageL_1. This partitions the bank intoS_0(|S_0|=48) andS_1(|S_1|=143); six tasks remain unrecovered because of implemented oracle heuristics and are excluded from stratum contrasts. -
C1 — grounding unlocks
S_0: OnS_0,D_0L_0yields 0/48 whileD_1L_0yields 38/48 (79.2%), a paired risk difference of +79.17 pp (95% bootstrap CI [+66.67, +89.58]), exact McNemarp_raw = 2^-37 = 7.28 × 10^-12(p_Holm = 1.46 × 10^-11). -
C2 — expressivity unlocks
S_1: OnS_1,D_0L_0yields 0/143 whileD_0L_1recovers 142/143 (99.3%), a paired risk difference of +99.30 pp (95% CI [+97.90, +100.00],p_Holm = 1.08 × 10^-42). -
C3 — a non-monotonic interaction on gpt-oss-120b: On
S_1, adding state-grounded feedback to the rich language drops recovery from 142/143 (99.3%) underD_0L_1to 133/143 (93.0%) underD_1L_1(Δ = −6.29 pp, 95% CI [−11.19, −2.10],p_Holm = 0.0117). The full-cohort interaction isI = (D_1L_1 − D_0L_1) − (D_1L_0 − D_0L_0) = −25.89pp (95% CI [−34.01, −18.27]). Trace analysis of the 11 discordant tasks suggests changes in semantic decomposition and operation ordering rather than systematically longer recoveries; the paper labels this analysis descriptive rather than causal. -
Full factorial table (N=197, B=4):
D_0L_00/197 (0.0%), 0/48, 0/143;D_1L_038/197 (19.3%), 38/48 (79.2%), 0/143;D_0L_1180/197 (91.4%), 38/48 (79.2%), 142/143 (99.3%);D_1L_1167/197 (84.8%), 34/48 (70.8%), 133/143 (93.0%). -
End-to-end 600-task accounting: 478/600 (79.67%) produced capability-positive mutations — 281 already admissible and 197 recovery-defective (197/478, 41.2% of capability-positive). Under
D_0L_1atB=4, EvoUndo rescues 180/197 (91.37%), yielding 461/478 (96.44%) admissible capability-positive mutations and 461/600 (76.83%) total benchmark yield, a +30.00 pp absolute increase over the initially admissible 281/600 (46.83%). Zero forward-capability regressions were introduced by recovery repair becausem_0is executed verbatim. -
Cross-model replication: On a frozen 96-task Qwen3.8-27B constrained-structured-decoding replication, C1 is +47.92 pp (
p < 0.001) and C2 is +37.50 pp (p < 0.001), while C3 is +6.25 pp (95% CI [−2.08, +16.67],p_adj = 0.375) — the negative interaction did not reproduce. Model, decoding protocol, and sample size all differ, so the cause is not isolated. -
Fresh-holdout precision–coverage trade-off: On 300 fresh tasks (50 per family) with 100 hidden states, increasing diverse development states from
Q_2toQ_20reduced observed false positives from 16.8% to 0.0% but loweredRescue@4_100from 49.7% to 35.0% (Δ = −14.67 pp,p = 5.76 × 10^-6) and increased the conservatism gap from 6.6% to 24.1%. UnderQ_20, Listeners had 0/50 development passes but 32/50 hidden-criterion passes. -
Effect-contract auditing: In a controlled audit of 300 intentionally incomplete effect contracts across all six harness families, independent dynamic effect tracking detected every injected undeclared semantic effect (300/300) with no false positives on 300 matched complete contracts.
-
Cost: Fresh-holdout inference ranged 2.39–3.09 calls and 3.37k–4.30k generated tokens per task; the final strict-
L_0D_1L_0run averaged 3.51 calls/task. -
Snapshot baseline comparison: Effect-scoped snapshots are the stronger baseline in the evaluated serializable in-memory selective-undo regime, recovering 300/300 tasks versus 243/300 for EvoUndo under different-surface changes and 159/300 versus 131/300 under same-surface changes. There are no EvoUndo-only wins; superiority in settings where exact effect-scoped restoration is unavailable is not established.
Methodology in Plain English
The researchers model an agent harness as a persistent state plus an execution policy, and a mutation as a function that maps one state to another. A candidate self-modification is packaged as a four-part object: the forward mutation, a witness program that records prior values and existence flags before the mutation runs, a recovery program that uses that witness plus a mutation-produced resource receipt, and a declared effect contract that says which state surfaces the recovery must cover.
For each candidate, the system runs a counterfactual round trip — capture witness, apply mutation, run recovery — across development states and hidden states, and checks whether the recovered state is observationally equivalent to the pre-state under typed equality rules (sorted keys, canonicalized pointers, SHA-256 file hashes, and so on). The forward mutation is immutable: only the witness, recovery program, and effect contract can be revised over up to B = 4 synthesis turns. The runtime independently computes the actual effect set and requires the declared contract to cover it, so shrinking the contract cannot be used to evade verification, and the contract cannot be omitted.
Recovery is bounded by two formal recovery languages. L_0 covers configuration, prompt, tool-registry, and routing state; L_1 extends it with middleware sequences, event listeners, sandboxed files, managed sockets, and ordered (LIFO) composition of recovery operations. The team built a controlled benchmark of 120 injected defects, a natural failure bank from 600 single-sample zero-shot tasks across six 100-task families, and a deterministic constructive oracle audit to establish which failures are recoverable in principle under each language. A protocol-locked 2×2 factorial then crossed diagnostic granularity (coarse typed versus state-grounded bundle) with recovery language (L_0 versus L_1), with three pre-registered paired contrasts. Admissibility used a split-Wald rule requiring the normal-approximation lower endpoint at τ_R = 0.85 to pass independently on IID and OOD hidden splits of n = 20 each, so at least 19/20 successes are needed per split.
Why This Matters
The work reframes agent self-improvement: forward capability gain alone is insufficient for long-lived autonomous systems, and recovery correctness is a relational property that depends jointly on the state distribution, witness semantics, observational contract, and recovery-language expressivity. It shows that when recovery representations align with the underlying state structure, LLMs can synthesize state-dependent recovery programs across counterfactual states — but also that adding the most granular diagnostics is not universally better.
Real-world applications:
- Agent platforms that let models edit their own tools and middleware, where a bad edit must be undone without corrupting later state.
- Autonomous coding and DevOps harnesses that modify routing, configuration, or sandboxed files and need auditable rollback.
- Resource-managing agents that allocate sockets, listeners, or files and must release exactly what they introduced.
- Safety and compliance review pipelines that need independently verified evidence that a persistent modification is reversible before admission.
Industry relevance: the paper argues for co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting. It also reports that straightforward effect-scoped snapshots outperform EvoUndo in the evaluated serializable in-memory selective-undo regime, which is useful guidance for teams deciding whether to build snapshot infrastructure or recovery synthesis.
Future Directions
- Joint optimization of the forward mutation and recovery, since this study holds
m = m_0fixed and never tests capability-equivalent but more reversible forward mutations. - Adaptive diagnostic strategies that escalate from coarse indicators to granular state traces only when coarse repair fails, since no diagnostic curriculum or address-use constraint was evaluated and the paper does not establish causality for the negative
C_3interaction. - Resolving the precision–coverage trade-off, where strict development gates (20/20 round trips under
Q_20) drove false positives to 0.0% while cuttingRescue@4_100to 35.0% and leaving Listeners at 0/50 development passes but 32/50 hidden-criterion passes. - Extending the modeled state surfaces — distributed and external state, third-party APIs, and unmanaged processes are unmodeled — and adding compensation for irreversible effects. Only gpt-oss-120b and Qwen3.8-27B were evaluated, and no low-, medium-, or high-reasoning-effort comparison was performed.
Target Audience
Researchers in LLM agent architectures, self-improving and self-evolving systems, and reversible or bidirectional computation; systems engineers building agent runtimes with rollback, undo, or effect-scoped snapshot infrastructure; and safety, evaluation, and governance practitioners who need verified guarantees that persistent model-generated modifications can be reversed. The formal notation, factorial design, and statistics make it most accessible to readers with a background in agent systems or program verification.
Authors’ abstract
LLM agents increasingly modify their own prompts, tools, middleware, resources, and execution harnesses at runtime. Such self-evolution can improve capability, but a successful mutation may leave persistent effects that cannot be safely reversed in states different from the one in which it was created. We introduce EvoUndo, a framework for representing, synthesizing, diagnosing, and independently verifying recoverability of model-generated self-modifications across counterfactual states. Across 600 unseen one-shot self-evolution tasks, we identify 197 capability-improving mutations that fail recoverability verification. Under the original recovery representation, conventional repair strategies recover 0/197 of these natural failures. Deterministic oracle analysis recovers 48/197 under the original recovery language L0, while the extended recovery calculus increases empirical oracle recovery to 191/197. A protocol-locked 2x2 grounding-by-expressivity intervention then separates two bottlenecks: exact state-address grounding increases successful recovery from 0/48 to 38/48 (79.2%) when the original language is sufficient, while extending the recovery language enables recovery on 142/143 (99.3%) failures in the oracle-defined S1 stratum. On the primary gpt-oss-120b backbone, adding exact-address diagnostics to the richer language reduces recovery to 133/143 (93.0%); a Qwen3.8-27B replication preserves the grounding and expressivity effects but not this negative interaction, indicating that the latter is model-dependent. These results indicate that reliable agent self-evolution requires co-designing verification, state grounding, witness semantics, and recovery-language expressivity rather than relying on iterative prompting alone.