Research
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite
Overview Research area: AI agents / terminal-task agents, harness engineering, and self-improvement through supervised finetuning on rewritten trajectories (LLM agent training). Technical level: Inter

- arXiv
- 2610.02826
- Published
- 2026-10-02
- Authors
- Zongxia Li, Yucheng Shi, Zhongzhi Li, Junyao Yang, Ruhan Wang, Chengsong Huang, Fuxiao Liu, Haitao Mi, Jordan Boyd-Graber, LeoweiLiang
AI summary
Overview
Research area: AI agents / terminal-task agents, harness engineering, and self-improvement through supervised finetuning on rewritten trajectories (LLM agent training). Technical level: Intermediate to Advanced. Scope: The paper introduces Recursive Self-Rewrite (RSR), a framework that collects successful terminal-task trajectories from three different agent harnesses and rewrites them into verified, harness-agnostic demonstrations that are used to finetune a single 27B base model.
What This Paper Is About
A model's ability to complete hard terminal tasks depends heavily on the harness it runs in — the loop that feeds it observations, tracks progress, verifies results, and decides when to stop. Different harnesses solve complementary sets of tasks, but simply pooling their successful trajectories for finetuning mixes in harness-specific control logic that will not exist at inference time. RSR's goal is to convert those harness-assisted successes into reusable capability by having the same base model reconstruct each success as a runbook, screen out leakage, and re-solve the task from scratch under a single general harness.
Key Contributions
- A multi-harness discovery study. The authors run one fixed base model (Qwen-3.8-27B) under three harnesses — Terminus 2, StateM, and Recursive Self-Reflect Terminus (RSRT) — over an approximately 3K-task terminal pool, and quantify how much each harness and their union solve, including per-domain and per-task-type coverage.
- The Recursive Self-Rewrite pipeline. Three roles of the same base model — a planner that turns a source trajectory into a runbook, a critic that filters verifier and solution leakage, and an executor that re-solves the task in a fresh sandbox under the general Terminus 2 harness — plus deterministic and model-based filtering and rejection sampling.
- Scaling of training data and measured gains. Reconstruction expands the training set from 2,001 to 11,094 high-quality trajectories, and finetuning on them outperforms both the base model and Direct SFT on five benchmarks plus process reward on Long-Horizon Terminal Bench.
- An analysis of rewriting behavior. Command-similarity metrics within versus across runbooks, and two case studies (Markdown inline parsing, OpenFOAM PitzDaily) showing that runbooks transfer useful procedure but do not guarantee success.
Main Findings
- Harnesses are complementary in task coverage. Across the pooled 2,929 tasks, Terminus 2 solves 565 (19.3%), RSRT 512 (17.5%), and StateM 461 (15.7%); the union of all three solves 759 tasks (25.9%), adding 194 tasks over the strongest individual harness, a 34.3% relative increase. The union improves over the strongest individual harness by 19.4% on RST and 33.8% on SWR.
- Complementarity survives equal sampling. On the 1,247 tasks with equal rollout counts across harnesses (2,074 rollouts per harness), Terminus 2, RSRT, and StateM solve 249 (20.0%), 285 (22.9%), and 227 (18.2%) tasks; their union solves 352 (28.2%), adding 67 tasks over RSRT, a 23.5% relative increase, with 30, 56, and 21 exclusive solves respectively.
- Exclusive solves are common. Of the 759 tasks solved by at least one harness, 288 are solved exclusively by one harness (129 only Terminus 2, 91 only RSRT, 68 only StateM), 163 by exactly two harnesses, and 308 by all three.
- No single harness dominates every domain. Under 87 fine-grained labels (35 RST task categories, 52 SWR tool families), Terminus 2, RSRT, and StateM solve at least one task in 49, 46, and 42 domains, with a union of 51; across 88 task types they cover 56, 52, and 49 types, with a union of 60. RSRT leads RST scripting and automation (35.7%) and SWR media and signal processing (13.1%); StateM leads RST software development and debugging (26.2%); combining harnesses raises SWR media and signal coverage to 22.3% and earth and geospatial to 19.8%, versus best individual rates of 13.1% and 12.2%.
- Discovery rollouts. The pool contains 14,598 rollouts with 2,001 successful trajectories covering 759 distinct tasks. Per harness: Terminus 2 5,405 rollouts, 766 passing, 14.2% success rate; RSRT 3,777 rollouts, 599 passing, 15.9%; StateM 5,416 rollouts, 636 passing, 11.7%; pooled 13.7% success rate.
- Execution styles differ measurably. On passing trajectories, Terminus 2, RSRT, and StateM have mean turns of 19.4, 23.2, and 24.0 (medians 17.0, 18.0, 24.0); completion claims per trajectory of 1.72, 2.20, and 1.11; exploration command shares of 55.5%, 50.6%, and 35.3%; and state-management shares of 0.0%, 0.0%, and 29.6%. RSRT passes after a rejected completion in 4.5% of passing trajectories. In a fixed sample of 200 failed trajectories per harness, RSRT averages 67.1 turns versus 22.8 for Terminus 2 and 23.4 for StateM.
- Rewrite throughput. With K=4 candidate runbooks and M=4 fresh executions at temperature 0.7 per successful source rollout, RSR collects 12,893 rollouts covering 975 unique tasks, with 86% passing, yielding 11,094 passing rewrites.
- Training gains over Direct SFT (pass@3). RSR improves over Direct SFT by 20.8, 4.1, 4.6, 7.0, and 3.0 percentage points on TB2, TB3, TB4, TBH, and SWR100 respectively. Mean per-run gains over Direct SFT are 26.3, 5.4, 2.9, 11.5, and 2.0 percentage points on the same five benchmarks.
- Absolute results. Pass@3 goes from 57.0% (51/89) for Base and 53.4% (48/89) for Direct SFT to 74.2% (66/89) for RSR on TB2; on TB3 from 0.0% (0/74) and 5.4% (4/74) to 9.5% (7/74); on TB4 from 1.5% (1/66) and 4.5% (3/66) to 9.1% (6/66); on TBH from 39.0% (39/100) and 56.0% (56/100) to 63% (63/100); on SWR100 from 3.0% (3/100) for both to 6.0% (6/100). Means over three runs are TB2 51.7% / 43.8% / 70.1%, TB3 0.0% / 1.4% / 6.8%, TB4 0.0% / 3.2% / 6.1%, TBH 31.0% / 42.5% / 54.0%, SWR100 3.0% / 3.0% / 5.0%.
- Direct SFT can hurt. Direct SFT falls by 3.6 percentage points on TB2 pass@3 and its TB2 mean per-run rate drops from 51.7% to 43.8%, a 7.9 percentage point decrease; the authors attribute this to circular and dead-end behavior in raw trajectories from RSRT, many of which exceed 100 steps.
- Process reward improves without full solves. LHTB process reward rises from 0.21 (Base) and 0.25 (Direct SFT) to 0.29 (RSR), but all three groups complete zero of the 46 tasks, so the gain reflects greater partial progress rather than more solved tasks.
- Harness choice alone changes benchmark results. With no training, the base model scores TB2 51.7 (46), TBH 33.0 (33), TB3 0.0 (0) under Terminus 2; TB2 55.1 (49), TBH 66.0 (66), TB3 4.1 (3) under RSRT; TB2 58.4 (52), TBH 46.0 (46), TB3 1.3 (1) under StateM; and TB2 68.5 (61), TBH 70.0 (70), TB3 4.3 (3) under the union of all three.
- Runbooks shift command-level behavior. For Markdown (12 rewrites, three runbooks, 18 same-runbook and 48 different-runbook pairs), same-runbook pairs score higher on exact command overlap (19.1 vs 12.1), command order (24.6 vs 15.7), tool-transition overlap (34.0 vs 18.9), and action-sequence similarity (53.6 vs 42.0). For OpenFOAM (eight rewrites, two runbooks, 12 and 16 pairs) the differences are weaker for exact overlap (12.6 vs 10.4) and command order (18.2 vs 14.7), with tool-transition overlap nearly equal (28.4 vs 28.5) and action-sequence similarity slightly lower within runbooks (61.0 vs 63.1).
- Rewriting transfers procedure but not guaranteed success. On the Markdown task, 7 of 12 rewrites pass (58.3%) and the median passing rewrite is 32 turns, a 50.0% reduction from the 64-turn source. On OpenFOAM, 7 of 8 rewrites pass (87.5%) and the median passing rewrite is 29 turns, with the source parsed at 38 turns (registry entry 39; Markdown registry entry 65).
Methodology in Plain English
The authors start with a single base model, Qwen-3.8-27B, and an approximately 3K-task pool of terminal tasks: roughly 2.5K self-constructed SWR tasks spanning 50 domains (software usage, biology, chemistry, physics, hardware, operations, security) plus 420 filtered and modified tasks from the RST dataset with increased difficulty.
They run that one model under three different harnesses. Terminus 2 is a general terminal interaction loop and is also the target "general harness" used later for training and inference. StateM adds persistent workflow states and checked transitions. Recursive Self-Reflect Terminus extends Terminus 2 with a discovery protocol that gives up to 180 minutes per task, runs the verifier on each completion claim, and on failure reports only the failure outcome so the model can review, debug, and continue.
They pool every passing trajectory from all three harnesses. For each one, a planner produced by the same base model compacts the source (keeping the task instruction, model actions, and environment observations while stripping harness control messages) and reconstructs it as a runbook describing end state, milestones, checks, recovery strategies, and pitfalls, without handing over the finished deliverable. A critic — again the same model — screens candidate runbooks, first with deterministic checks (schema validation, removal of known artifacts, rejection of unsupported tool references), then with a model-based judgment that sees only the public task instruction and the runbook. Accepted runbooks guide an executor that re-solves each task in a fresh sandbox under Terminus 2, with the runbook given as private guidance that never appears in the public trajectory. Trajectories are then filtered for values that cannot be derived from the task or environment, and only public interaction histories are kept for finetuning.
The finetuning comparison has three arms on the same base model: Base (no training), Direct SFT (finetuning on passing source rollouts with harness-specific insertion phrases removed but no rewriting), and RSR (finetuning on verified rewrites plus the 766 direct passing Terminus 2 trajectories). Evaluation uses pass@3 and mean per-run pass rate over three runs on Terminal-Bench 2, Terminal-Bench 3, Terminal-Bench 4, Terminal-Bench Hard, and Software Terminal 100, plus process reward on Long-Horizon Terminal Bench.
Why This Matters
The paper argues that harnesses need not only be inference-time scaffolding but can serve as discovery tools whose successes are distilled back into model weights. That reframes harness diversity as a data-generation strategy rather than a permanent crutch, and it targets the distribution-mismatch
Authors’ abstract
Successful trajectories on difficult tasks provide valuable supervision for model improvement, but specialized harnesses introduce interventions that may be unavailable during deployment. We propose Recursive Self-Rewrite (RSR), a framework that uses one base model, Qwen-3.8-27B, to discover successful solutions under diverse harnesses and reconstruct them as training trajectories under a general harness. A planner extracts procedures into runbooks, a critic screens for verifier and solution leakage and guides recursive revision, and an executor follows qualified runbooks in fresh sandboxes. Across approximately 3K self-curated terminal tasks, three harnesses jointly solve 759 tasks, 34.3% more than the strongest individual harness in the recorded pool. RSR expands 2,001 successful source trajectories into 11,094 rewritten trajectories for supervised finetuning. Training on these trajectories outperforms both the base model and direct trajectory SFT. Compared with the base model, pass@3 increases from 57.0% to 74.2% on Terminal-Bench 2, from 1.5% to 9.1% on Terminal-Bench 4, from 39.0% to 63.0% on our self-curated Terminal-Bench Hard, and from 3.0% to 6.0% on our Software Terminal-Bench. Process reward on Long-Horizon Terminal-Bench rises from 0.21 to 0.29. These results show how diverse harness-assisted experiences can be reconstructed into reusable capabilities for a model operating under a general harness.