Skip to content
AI.info

Research

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins

Overview Research area: LLM agent harness optimization and recursive self-improvement (artificial intelligence / agent systems). Technical level: Advanced. Scope: The paper introduces PluginRSI, a met

PluginRSI: Recursive Improvement of Agent Harnesses with Reusable Plugins
arXiv
2609.32423
Published
2026-09-26
Authors
Yaorui Shi, Yuchun Miao, Yuxin Chen, Jiayuan Zhang, Yueqing Sun, Xierui Song, Xiang Wang, An Zhang

AI summary

Overview

  • Research area: LLM agent harness optimization and recursive self-improvement (artificial intelligence / agent systems).
  • Technical level: Advanced.
  • Scope: The paper introduces PluginRSI, a method that represents an agent harness as a composition of independently evolved, reusable plugins and shows gains over prior harness-optimization baselines on software engineering, command-line interaction, and question-answering tasks.

What This Paper Is About

The "harness" around a language model — the tools, roles, skills, memory and the workflow coordinating them — heavily determines agent performance. Prior optimization methods rewrite whole harness programs, which couples mechanisms together so individual ideas are hard to isolate, evaluate, or reuse across iterations. PluginRSI instead treats each mechanism as a separate plugin with a standardized interface, improves plugins one at a time, stores the wins in a shared library, and recomposes new harnesses from that library.

Key Contributions

  1. A plugin-parameterized harness formulation. A harness is written as H = ℋ(W, 𝒫), where 𝒫 is a set of plugins (each with a YAML specification and Python implementation) and W is the workflow that coordinates them. This lets a single mechanism change while everything else stays fixed.
  2. The PluginRSI algorithm, which alternates two stages per iteration: plugin mutation (parallel exploration of individual plugin changes, evaluated on balanced minibatches, with improving variants added to the shared library) and harness recomposition (selecting plugins from the updated library and revising the workflow).
  3. A shared plugin library that accumulates and reuses mechanisms across iterations, seeded from prior research and open-source agent systems (75 components: 20 Roles, 16 Tools, 31 Skills, 8 Memory).
  4. An empirical study across SWE-bench Verified, Terminal-Bench-2.1, and five QA benchmarks, including cross-solver transfer, optimization-dynamics comparisons, ablations, and a plugin-reuse experiment.

Main Findings

  • SWE-bench Verified gains. On 400 held-in / 100 held-out tasks, PluginRSI reached held-out resolve rates of 65.00% (optimized with GPT-5.6 Terra) and 69.00% (optimized with Kimi-K3), exceeding the best baseline by 10.00 and 6.00 percentage points. The corresponding held-in gains were 4.00 and 2.50 points, so held-out gains were larger than held-in gains. The strongest baseline reported was Meta-Harness (55.00% and 63.00% held-out in those settings).
  • Cross-model transfer without further optimization. The GPT-5.6 Terra-optimized harness reached 60.00% held-out with Kimi-K3 and 52.00% with GLM-5.2, exceeding the strongest baselines by 7.00 and 4.00 points. The Kimi-K3-optimized harness reached 64.00% with both GPT-5.6 Terra and GLM-5.2, improving over the strongest baselines by 3.00 and 5.00 points. The authors note some held-in cross-model cases where PluginRSI scored 58.00% against a previous best of 58.25%, calling the gaps modest and the results comparable.
  • Terminal-Bench-2.1 gains. With Kimi-K3, PluginRSI scored 69.5% on the 59 held-in tasks and 73.3% on the 30 held-out tasks, versus Meta-Harness at 64.4% and 66.7%. Improvements were larger on held-out tasks.
  • QA gains. Optimized and evaluated on Kimi-K3, PluginRSI raised held-in average accuracy from 0.753 to 0.820 and overall accuracy from 0.638 to 0.675, with held-out accuracy rising more modestly from 0.568 to 0.588. The authors attribute the smaller held-out gain to limited overlap among QA domains (e.g., math strategies may not transfer to law questions).
  • Faster optimization. PluginRSI reached a held-in resolve rate that Meta-Harness did not attain within the evaluated budget, using fewer cumulative solver rollouts; Meta-Harness's held-out performance fluctuated and remained lower at the end.
  • Both stages matter. Removing plugin mutation dropped the Kimi-K3 held-out resolve rate from 69.00% to 57.00%, and removing harness recomposition dropped it to 59.00%; both ablations also lowered held-out scores with GPT-5.6 Terra and GLM-5.2.
  • The library alone is not the source of gains. Adding an initialized or empty plugin library to Meta-Harness gave Kimi-K3 held-out rates of 64.00% and 62.00%, close to its original 63.00%. Removing PluginRSI's initial plugins still left Kimi-K3 held-in scores similar (66.75% vs 67.00%) and reduced held-out scores by only 1.00–2.00 points across the three solvers.
  • Plugin accumulation and reuse. The library grew from 75 to 132 plugins over 15 evolution steps. PluginRSI candidates reached approximately 1,300 Python lines of code by the final step versus 400 for Meta-Harness, while the workflow remained below 300 lines. Reusing the evolved library from the initial harness reached a held-out resolve rate of 69% after two evolution steps, compared with 65% without reuse, approaching the 70% reached after three steps of continued optimization from the evolved harness.
  • Advantage over additional baselines. With Kimi-K3 as both proposer and solver, PluginRSI improved held-out resolve rates over SkillOPT, A-Evolve, and AHE by 9.00, 12.00, and 6.00 points respectively.

Methodology in Plain English

The researchers start with a working agent harness built from a library of 75 pre-built plugins (roles, tools, skills, memory) and an initial workflow that coordinates four of them. Each optimization iteration has two steps:

  1. Plugin mutation: the current harness is run on the validation tasks, then eight mutation candidates are generated in parallel. Each candidate gets its own minibatch of 20 tasks balanced with 10 solved and 10 unsolved problems, so different candidates see different feedback. The proposer agent picks one plugin in the harness, rewrites it, and the harness is recompiled with the rewritten plugin and everything else unchanged. If the change improves the score on that minibatch, the new plugin version is saved into the library.
  2. Harness recomposition: the proposer looks at past execution feedback and the mutation results, then selects a new plugin set from the updated library and revises the workflow describing when plugins run and how information flows. One new candidate is built and evaluated on all held-in tasks.

The search runs for 15 iterations against a 400-task held-in set (each iteration costs 400 + 8 × 20 = 560 solver rollouts), and the best harness on the validation set is frozen and then tested on 100 held-out tasks. Baselines (GEPA, DGM, Meta-Harness, and expert-curated ReAct and ACE) run for 25 iterations under the same splits, API endpoints, and sandboxes. The proposer uses reasoning effort xhigh with a 16,384-token output cap; solvers use none reasoning effort (or disabled for GLM-5.2) with an 8,192-token cap.

Why This Matters

Harness design currently sits somewhere between hand-crafting and opaque whole-program rewriting. PluginRSI shows that agents can improve their own scaffolding in a way that leaves behind inspectable, transferable artifacts, and that those artifacts transfer to solver models that were never used during optimization.

Real-world applications:

  • Software engineering agents: harnesses for repository-level bug fixing, where the paper reports the largest held-out gains.
  • Command-line and terminal automation: improving agents that operate shells and system tools, as measured on Terminal-Bench-2.1.
  • Domain question answering: assembling different plugin combinations for math, law, medicine, economics, or science questions.
  • Reusable agent infrastructure: teams could maintain a growing library of vetted plugins instead of rebuilding agent stacks per task.

Industry relevance: the method adapts agent behavior without weight updates, so the costs are API rollouts rather than training runs. The authors report cross-vendor transfer across GPT-5.6 Terra, Kimi-K3, and GLM-5.2, which is relevant for organizations deploying agents on whichever model is cheapest or available at a given time. The paper also notes that evolved harnesses can generate code and commands that introduce errors or unintended changes, so deployment beyond benchmark environments should include code review and appropriate execution permissions.

Future Directions

  • Longer-term evolution under changing task distributions, plus management of a growing plugin library (the paper lists this as a limitation).
  • Broader cross-domain and cross-model transfer, since gains were smaller across dissimilar QA domains and could improve with multi-domain, multi-model optimization.
  • Joint plugin–workflow optimization, because evaluating each mutation with the rest of the harness fixed may miss mechanisms that require coordinated changes elsewhere.
  • Extending to long-horizon tasks, which the authors name as future work alongside more significant cross-domain transferability.

Target Audience

Researchers and engineers working on LLM agents, agent harness design, and self-improving systems. It is most useful to readers who already understand agent loops, tool use, and harness-optimization baselines such as Meta-Harness, DGM, and GEPA, and who want to know whether decomposing a harness into independently evolvable, reusable components improves optimization efficiency and transfer.

Authors’ abstract

The harness surrounding a language model is a central determinant of agent performance. Recent methods optimize harnesses by searching over complete programs, where individual mechanisms are difficult to isolate and reuse. We introduce PluginRSI, which represents a harness as a composition of atomized plugins and organizes harness evolution around these plugins. Individual plugins are improved independently and accumulated in a shared library, then recombined into new harnesses at each iteration. PluginRSI improves over existing harness optimization methods across software engineering, command-line interaction, and question-answering tasks. The resulting harnesses retain their advantage when transferred to other solver models without further optimization. The evolved plugin library accelerates subsequent optimization from the initial harness, which helps faster and higher convergence on unseen tasks. These results show that accumulating reusable mechanisms provides an effective basis for continued harness improvement.

Read the original paper