Research
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives
Harness Engineering for Software Engineering via Modular Executable Dev-Primitives Overview Research area: Software engineering agents built on large language models — specifically the harness (the sc

- arXiv
- 2610.07832
- Published
- 2026-10-06
- Authors
- Haibo Jin, Xinjie Li, Peng Kuang, Haohan Wang
AI summary
Harness Engineering for Software Engineering via Modular Executable Dev-PrimitivesOverview
Research area: Software engineering agents built on large language models — specifically the harness (the scaffolding around a model) rather than the model itself.
Technical level: Advanced. The paper assumes familiarity with LLM agents, repository-level code editing, benchmarks such as SWE-bench, and multi-agent coordination concepts.
Scope: The paper introduces Dev-Primitives (modular executable interfaces that pair each repository component with its own resident LLM) and HERMES, a harness-engineering framework that activates, coordinates, executes, and diagnoses these primitives, evaluated on four software engineering benchmarks.
What This Paper Is About
LLM agents with terminal access still struggle on long-horizon software engineering workflows because the program state they need is scattered across interdependent source files, configurations, tests, dependencies, and runtime behavior. Agents must repeatedly reconstruct this scattered state inside a single growing interaction history, which leads to context explosion, semantic drift, and degraded reasoning. The paper's goal is to stop forcing one central agent to reconstruct everything, and instead let individual repository components reason about their own responsibilities and communicate relevant requirements directly to one another.
Key Contributions
-
Dev-Primitives. A modular, executable abstraction that wraps a single repository artifact (source file, configuration file, or test file) together with a resident LLM. Each primitive reads its own implementation, can be addressed in natural language, exchanges implementation requirements with other primitives, and can modify its own artifact. Formally, a primitive computes
(a_i', m_i) = P_i(a_i, x_i, C_i) = M([a_i; x_i; C_i]), wherea_i'is the optionally modified artifact andm_iis a natural-language message to other components. -
HERMES, a harness-engineering framework. HERMES (Harness Engineering for software engineeRing via Modular Executable Dev-primitiveS) instantiates Dev-Primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised.
-
Extensive evaluation across four benchmarks. SWE-bench Verified, SWE Refactor Bench, Terminal-Bench 4.0, and DevOps-Gym, spanning issue resolution, whole-repository refactoring, terminal-based tasks, and DevOps workflows.
-
A scaling and cost analysis across harness components. The paper studies heterogeneous configurations that allocate different model sizes to activation, Dev-Primitives, and diagnosis, and reports the resulting cost-performance trade-off.
Main Findings
-
Average improvement over matched baselines: HERMES improves over matched baseline harnesses by 12.4 percentage points on average across the four benchmarks.
-
SWE-bench Verified (500 instances): HERMES improves the resolved rate for 19 of the 20 backbones, ties on the remaining one (Claude Opus 5), and outperforms both baselines for the two backbones evaluated against two harnesses. It reaches 97.0% with GPT-5.6 Sol versus 96.2% under mini-SWE-agent (+0.80), and the largest gain is +19.4 points for GPT-5 Mini (60.80 to 80.20). HERMES reaches 80.6% with Qwen3-8B.
-
Lower latency in some configurations: With Claude Sonnet 5, HERMES improves resolution by 6.2 points while reducing average latency from 962.37 s to 824.63 s; DeepSeek-V4 behaves similarly.
-
SWE Refactor Bench (20 whole-repository migration tasks): The largest gains appear here. With GPT-5.6 Sol the composite score rises from 6.5% to 31.0% at medium effort (+24.5) and from 19.0% to 36.5% at high effort (+17.5). GPT-5.6 Luna rises from 0.0% to 16.0% at medium effort. Repeating the default GPT-5.6 Sol configuration three times yields 31.0 ± 1.3, indicating the gains are not run-to-run variation.
-
Terminal-Bench 4.0 (66 executable tasks, five trials per configuration): HERMES improves GPT-5.6 Sol by +18.8 (33.0 to 51.8) at medium effort and +18.2 (37.3 to 55.5) at max effort, and GPT-5.6 Terra by +26.3 and +26.7. Claude Opus 5, whose baselines are already strong, gains only 1.5 and 1.2 points. Claude Sonnet 5 is a notable exception in cost: HERMES improves resolution while reducing both token consumption and inference cost at either effort level.
-
DevOps-Gym (build/configuration, monitoring, issue resolving, test generation): Gains appear in all four categories. With GPT-5.6 Sol the average rises from 49.89% (Codex) to 55.41% (+5.52). HERMES reaches 37.48% with Qwen3-8B.
-
Comparison with related multi-agent frameworks: MAGIS reports 13.9 resolution on SWE-bench while HERMES with GPT-5.6 Luna reaches 36.6%. On NL2Repo-Bench, CodeTeam reports 34.6% (prompting) and 42.3% (fine-tuning) average test pass rates, while HERMES reaches 44.9%.
-
Diagnosis scaling matters more than activation scaling. Keeping Dev-Primitives fixed to Qwen3-8B, using GPT-5.5 only for diagnosis raises Terminal-Bench resolution from 27.6% to 46.1% (+18.5), versus 36.7% (+9.1) for activation alone; the same ordering holds on SWE-bench Verified (+5.8 versus +3.8).
-
Scaling the Dev-Primitives themselves adds comparatively little. It contributes between 1.2 and 4.5 points across the four benchmarks, so primitives can run on a small model without forfeiting most of the gain. With strong activation and diagnosis models, HERMES using Qwen3-8B Dev-Primitives stays within 4.5 percentage points of the homogeneous GPT-5.6 Sol configuration.
-
Cost efficiency on Terminal-Bench 4.0: GPT-5.6 Sol activation / Qwen3-8B primitives / GPT-5.5 diagnosis reaches 49.7% resolution, 2.1 points below the homogeneous GPT-5.6 Sol configuration, at 37.4% lower total inference cost ($2.13k versus $3.40k). Recovering those 2.1 points costs a further $1.27k, split between diagnosis (+0.9 points for $0.38k) and the Dev-Primitives (+1.2 points for $0.89k). The same setting buys 22.1 points over the homogeneous Qwen3-8B configuration for an additional $1.43k.
-
Selection quality: Dynamic activation recall is 95.4% on SWE-bench Verified, 83.3% on SWE Refactor Bench, 84.5% on Terminal-Bench 4.0, and 85.7% on DevOps-Gym, with precision of 72.6, 68.4, 64.9, and 70.3 respectively. The average number of activated primitives ranges from 4.7 to 12.6 per task. Solved tasks show higher selection recall than failed tasks by 16.5 to 24.7 points.
-
Ablations (GPT-5.6 Sol): Replacing Dev-Primitives with a centralized editor over the same selected components is the most damaging, costing 14.2 points on SWE Refactor Bench and 13.0 on Terminal-Bench; a compute-matched editor still trails HERMES by 6.3 points. Removing diagnosis feedback costs 4.4 (SWE-bench), 11.0 (Refactor), 10.6 (Terminal), and 8.68 (DevOps). Removing execution feedback costs 3.2, 9.5, 8.2, and 7.24. Removing inter-primitive communication costs 2.8, 7.5, 5.7, and 5.07. Removing on-demand activation costs at most 3.0 points but raises inference cost by 1.61–1.88×.
-
Three revision rounds suffice. Most of the gain is obtained within the first three revision rounds; increasing the budget B from 3 to 5 adds at most 0.9 points on any benchmark, so B = 3 is the default.
-
Note on backbone reporting: The setup text states that evaluations use seven backbones (GPT-5.6 Sol, Terra, and Luna, Claude Sonnet 5, Claude Opus 5, DeepSeek-V4, and Qwen3-8B), while the result tables additionally list other model names such as Claude Fable 5, GPT-5.5, GPT-5.4 Mini, GPT-5 Mini, Gemini variants, and Grok variants. The provided content does not reconcile this.
Methodology in Plain English
The authors do not change the underlying language models. They change what the models are attached to and how they talk to each other.
Step 1 — Attach a model to each piece of the repository. Every repository component (a source file, a config file, a test file) gets its own small LLM instance, called a Dev-Primitive. Because that model reads the code it is attached to, the component can answer questions about its own implementation and dependencies. It holds its own task objective, can edit itself, and produces a message describing what it now requires from neighboring components. Messages are directed, not broadcast — a primitive sends only to the components it implicates, which in practice are its callers, callees, configurations, and tests, so the communication topology follows the repository's dependency structure. Messages reach their recipients within the same round, so a requirement discovered while editing one file can still be satisfied by the components that depend on it before the round ends. A primitive can participate — contributing information — without modifying its own artifact.
Step 2 — Activate only what the task needs. Since only a small subset of a repository matters for a given issue, the ACTIVATE function localizes the components likely to implement the reported behavior and expands along their dependencies to callers, callees, configurations, and tests, producing a plan of (component index, local objective) pairs. Only those primitives are instantiated, so cost tracks the size of the activated set rather than the size of the repository.
Step 3 — Run the repository for real. Instead of asking a model to judge whether the change is correct, HERMES executes the modified repository in an executable environment and collects command outputs, exit codes, test outcomes, runtime errors, logs, and stack traces. Held-out evaluation tests are never used during solving.
Step 4 — Attribute failures back to owners. One round often leaves the issue unresolved, and execution reports only that the repository still fails, not which component is responsible. The DIAGNOSE step returns a pass/fail verdict and, on failure, structured feedback: the observed failure, the suspected root cause with the implicated components, and revision guidance. The plan is then revised — retuning objectives, dropping components, or activating new ones the evidence implicates — and the loop repeats until the repository state is accepted or a budget of B rounds is exhausted (B = 3 by default).
Evaluation approach. Experiments follow each benchmark's official leaderboard settings and protocols. Each result table compares a baseline harness against HERMES on the identical backbone and reasoning effort, so only the harness differs, and the reported Δ is the percentage-point gain over that row's baseline. Backbone assignments, execution isolation, primitive runtime, and baseline provenance are given in the appendices.
Why This Matters
Impact on research. The paper reframes where agent capability comes from. Rather than treating the harness as an implementation detail, it argues that harness design — where reasoning lives, how components are selected, and how failures are routed back — is a first-order factor in translating a model's raw capability into working software engineering behavior. It also offers a concrete alternative to append-only context management: instead of compressing a single growing history, keep task-relevant state with the artifact that owns it. The ablation results support this reading, since the largest loss comes from removing artifact-local primitives entirely.
Real-world applications:
- Large codebase maintenance and modernization: whole-repository migrations are where HERMES helps most, which maps directly onto multi-file refactoring, framework upgrades, and API migrations.
- CI/CD and DevOps automation: the DevOps-Gym results cover build and configuration, monitoring, issue resolving, and test generation — the stages of a production delivery pipeline.
- Terminal and operations work: Terminal-Bench 4.0 measures interactive terminal tasks, relevant to incident response and infrastructure operations.
- Cost-sensitive deployment: running lightweight Qwen3-8B primitives alongside stronger activation and diagnosis models offers a deployment path that trades a modest accuracy gap for meaningfully lower inference cost.
Industry relevance. The cost analysis matters for teams that operate agents at scale: the paper reports a configuration that lands 2.1 points below a homogeneous frontier-model setup at 37.4% lower total inference cost, and reports a 26.2% inference-cost reduction on Terminal-Bench 4.0 while remaining within 4.5 points of the homogeneous GPT-5.6 Sol configuration. It also reports latency reductions in some configurations, not just accuracy gains. The finding that diagnosis scaling buys more than activation scaling is a practical, immediately actionable design guideline for anyone building an agent harness.
Future Directions
-
Closing the remaining gap from scaling the Dev-Primitives. Scaling primitives themselves still adds 1.2 to 4.5 points across the benchmarks — the only component of HERMES not fully substituted by a small model. What specifically do stronger primitives contribute, and can it be recovered another way?
-
Improving activation precision and evaluating localization fairly. Precision ranges from 64.9% to 72.6% while recall ranges from 83.3% to 95.4%. The authors note that reference-patch overlap is a conservative measure, since a task can be resolved along a different file-level path from the developer patch, but how to measure localization quality without that bias remains open.
-
Pushing past the three-round revision budget. Gains plateau quickly, with B = 5 adding at most 0.9 points. Whether a different diagnosis or feedback representation would make additional rounds productive, especially on whole-repository migration, is unanswered.
-
Extending beyond the evaluated scope. The paper states that limitations are discussed in Appendix I. The provided content does not report the specifics of those limitations, the mechanism specifications in Appendix B, or the details of target construction for selection metrics in Appendix F.4.
Target Audience
Researchers and engineers working on LLM-based software engineering agents will benefit most — particularly those designing agent harnesses, context-management strategies, or multi-agent coordination schemes, and those evaluating agents on repository-scale or long-horizon tasks. Practitioners deploying coding agents in production will find the scaling and cost analyses directly actionable. Readers without background in LLM agents or repository tooling will need to consult the cited prior work (SWE-agent, OpenHands, CodeAct, Agentless, AutoCodeRover) to follow the comparisons.
Authors’ abstract
Large language models (LLMs) equipped with terminal access have demonstrated strong capabilities in automating software engineering tasks. However, existing agents remain brittle on long-horizon workflows, where they must repeatedly reconstruct program state scattered across source files, configurations, tests, dependencies, and runtime behavior, leading to increasingly long interaction histories, context explosion, and semantic drift. Large repositories further complicate the identification of task-relevant components. To address these challenges, we introduce \textbf{Dev-Primitives} (\emph{Development Primitives}), a modular and executable abstraction that transforms repository components from passive software artifacts into active participants in software engineering. Each Dev-Primitive pairs a repository artifact with a resident LLM, which gives the artifact an agent-native interface grounded in its own implementation and dependencies, enabling natural-language reasoning, inter-component communication, and localized self-modification. Building on Dev-Primitives, we propose \textbf{HERMES}, a Harness Engineering framework for software engineeRing via Modular Executable Dev-PrimitiveS, which instantiates these primitives at repository scale through a dependency-aware dynamic activation mechanism and a bug diagnosis mechanism that maps execution evidence back to the components that must be revised. Extensive experiments on four software engineering benchmarks demonstrate that HERMES outperforms matched baseline harnesses by 12.4\% on average. Moreover, when paired with strong activation and diagnosis models, HERMES, even with Qwen3-8B Dev-Primitives, remains within 4.5\% of the homogeneous GPT-5.6 Sol configuration across all four benchmarks, while reducing inference cost by 26.2\% on Terminal-Bench 4.0, highlighting the importance of harness design in software engineering agents.