Skip to content
AI.info

Research

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents

Overview Research area: Test-time compute scaling for LLM-driven terminal agents (natural language processing / agentic AI). Technical level: Advanced — the paper assumes familiarity with sampling-bas

Mid-Harness: Scaling Actions Between Model and Harness for Terminal Agents
arXiv
2609.39982
Published
2026-09-30
Authors
Minki Kang, Ryo Hachiuma, Shaokun Zhang, Subhashree Radhakrishnan, Yonggan Fu, Jindong Jiang, Mingjie Liu, Ehsan Hosseini-Asl, Yi Dong, Yu-Chiang Frank Wang, Byung-Kwan Lee

AI summary

Overview

Research area: Test-time compute scaling for LLM-driven terminal agents (natural language processing / agentic AI).

Technical level: Advanced — the paper assumes familiarity with sampling-based inference, verification methods (listwise, pointwise, pairwise), process reward models, LoRA distillation, and agent evaluation benchmarks.

Scope: A systematic study of "action scaling," which spends extra inference compute on sampling and verifying candidate terminal commands before execution, while holding the action-generating model and the execution harness fixed.

What This Paper Is About

Terminal agents commit to long-horizon sequences of stochastically generated commands, and one bad command (such as the wrong package install) changes the environment in ways that hinder all later steps — even when the model could have generated a better alternative. The paper asks whether allocating test-time compute at the boundary between the model and the harness can improve action reliability and trajectory success, and what makes that allocation effective. To study this, the authors introduce Mid-Harness, which samples several candidate actions from the unchanged generator, applies a verifier, and forwards only one chosen candidate to the unchanged harness.

Key Contributions

  1. Demonstrating that verification governs the benefit of action sampling. With a TMAX-9B generator on TerminalBench-Lite, wider sampling yields little benefit under a weak (zero-shot) verifier, whereas a strong verifier unlocks substantially more successful trajectories from the same generator.

  2. Comparing verification mechanisms and training a verifier by distillation. Pairwise verification performs best among the evaluated mechanisms, and distilling pairwise responses from GPT-5.6 Sol into the generator model improves trajectory success without changing the action generator. An offline analysis identifies command semantics and execution feasibility as persistent sources of disagreement with the teacher verifier.

  3. Providing Mid-Harness as a controlled experimental framework. By varying sampling and verification at a fixed model–harness boundary, the paper examines when additional computation improves trajectory success, and shows compatibility with parallel (Best-of-T) and sequential (Sequential Refine) trajectory scaling.

  4. Showing transfer across models, benchmarks, and harnesses. Gains appear for TMAX-4B/9B/27B, Qwen3.5-9B, Nemotron3.5 Lightning, and Nemotron3 Ultra, across TerminalBench-Lite, Terminal-Bench 2.1, SWE-bench-Verified (Mini subset), and FeatureBench-Mini, under the Vanillux2 and Terminus-2 harnesses.

Main Findings

  • The generator already contains useful alternatives. With TMAX-9B fixed as the generator on TerminalBench-Lite, verification by GPT-5.6 Sol raises Pass@1 from 50.00% for the base agent to 64.63% at N=4 and 68.03% at N=8 sampled actions — without changing or further post-training the generator.

  • More candidates do not compensate for weak verification. Under zero-shot listwise verification, doubling width from N=4 to N=8 changes Pass@1 only from 49.32% to 51.02% and Pass@3 from 66.33% to 67.35%. The frontier verifier uses the same listwise mechanism but achieves much higher success, suggesting the weaker verifier struggles to distinguish actions when comparing the full candidate set at once.

  • Pairwise verification performs best among the evaluated mechanisms. With the generator and N=8 fixed, pointwise improves Pass@1 only slightly over listwise and leaves Pass@3 unchanged, while pairwise reaches 54.76% Pass@1 and 71.43% Pass@3.

  • Distillation narrows the verifier-quality gap. Using 117k frontier-verifier pairwise responses collected from 732 trajectories across 244 difficult TMAX-15k tasks and LoRA training, distilled pairwise verification at N=8 raises Pass@1 from 54.76% to 57.14% and Pass@3 from 71.43% to 75.51%. At N=4, Pass@1 rises from 54.42% to 55.44% and Pass@3 from 68.37% to 70.41%. The fine-tuned LoRA is activated only for the verifier, not the generator.

  • Pairwise verification adds substantial decoding cost. At N=8, zero-shot pairwise responses add an estimated 26.6k parallelized output tokens (POT), versus at most 1.3k for listwise and pointwise. Total verifier output tokens at N=8 are 0.6k for listwise, 7.0k for pointwise, and 125.7k for zero-shot pairwise verification (which requires 28 calls for eight candidates, versus 1 listwise call and 8 pointwise calls).

  • Effective verification may not require explicit reasoning. At N=8, both the zero-shot and distilled TMAX-9B decision-only verifiers (emitting only an A/B preference) improve Pass@1 over their reasoning counterparts while lowering reference-priced token cost by 20.9% and 24.1%, respectively. At 4B and 27B, however, the evaluated decision-only variants have lower estimated token cost but also lower Pass@1 than their reasoning counterparts.

  • Distillation transfers teacher behavior, but command effects remain hard to judge. On an offline benchmark of stored TMAX-9B trajectories from 21 held-out TMAX-15K tasks with N=8, distillation reduces score MAE from 2.59 to 1.05, raises pairwise agreement from 59.01% to 74.58%, and raises verification agreement from 38.52% to 57.79%.

  • Agreement improves across the trajectory but stays lower later on. Verification agreement against the frontier verifier (measured over 1,355 states with valid comparisons for both models, 83.4% of states) improves in every turn bin, reaching 68.13% at turns 1–4 and 54.07% at turns 17–32.

  • Remaining disagreements center on command semantics and execution feasibility. Using GPT-5.6 Terra to classify clear verifier failures, the number of such cases falls from 3,328 for zero-shot verification to 1,810 after distillation, with fewer cases in every category. Candidate semantics and execution feasibility account for 67.4% of the distilled-verifier cases.

  • Action scaling composes with trajectory scaling. With TMAX-9B, Best-of-T with T=3 reaches 55.10% Pass@1; using zero-shot or distilled Mid-Harness to generate those runs raises Pass@1 to 61.22% and 66.33% — an 11.23 percentage-point gain with the same three environment executions. Using distilled Mid-Harness for both source trajectories and refinements raises Sequential Refine Pass@1 from 55.10% to 60.20% and Pass@3 from 71.43% to 75.51%.

  • Combining scaling axes improves the cost–success trade-off. SR plateaus at 55.10%, 56.46%, and 55.78% Pass@1 over R=1, 2, 3. Distilled Mid-Harness at N=8 matches Best-of-T at T=5 (57.14%) at about one-third of its reference-priced token cost. Combining it with SR or Best-of-T at T=3 reaches 60.20% or 66.33%, both exceeding Best-of-T at T=7 (59.18%) at lower reference-priced token cost. Decision-only verifiers reduce the total reference-priced token cost of these compositions further by 22–24%.

  • Gains appear across model scales. With Mid-Harness, Pass@1 goes from 38.78% (base) to 41.50% (zero-shot) to 43.88% (distilled) at 4B, and from 71.09% to 73.13% to 76.19% at 27B. TMAX-9B distillation reaches 57.14%.

  • Gains appear on other benchmarks, models, and harnesses. On Terminal-Bench 2.1 with TMAX-9B, zero-shot verification raises Pass@1 from 21.72% to 27.34%. On FeatureBench-Mini with TMAX-9B, the base agent succeeds in only 1.45% of runs, yet zero-shot and distilled verification raise Pass@1 to 5.80% and 7.25%; with TMAX-27B, distilled verification raises Pass@1 from 17.39% to 23.19% and Pass@3 from 26.09% to 39.13%. Both variants also improve TMAX-9B on SWE-bench-Verified. Across all seven settings reported in the transfer table, zero-shot Mid-Harness improves Pass@3 and matches or improves Pass@1. Two caveats are noted: for TMAX-9B, distillation lowers FeatureBench-Mini Pass@3 relative to zero-shot verification and does not improve over zero-shot on Terminal-Bench 2.1.

Methodology in Plain English

Mid-Harness is inserted as a wrapper around the model call inside an otherwise unchanged harness loop. Instead of requesting one action per step, the wrapper requests N candidates conditioned on the same interaction history, passes them to a verifier together with the task and observed history, and returns exactly one candidate for the harness to execute. The remaining candidates are discarded and the next set is generated from the updated history.

Three verification mechanisms are compared. Listwise verification puts all candidates in one prompt and asks the verifier to choose (one call, but the verifier must rank everything at once). Pointwise verification scores each candidate independently and takes the argmax (N parallel calls, but scores must be placed on a comparable scale). Pairwise verification compares two candidates at a time and aggregates margin-weighted win rates (focused comparisons, but many more calls). Under self-verification, the generator and verifier are the same model.

The main experiments fix a TMAX-9B generator and the Vanillux2 harness on TerminalBench-Lite, sampling at temperature 0.8 with at most 64 steps, a 65,536-token context, and a 16,384-token output limit per step. Three runs per task are evaluated, with Pass@1 as the average exact-success rate and Pass@3 as the fraction of tasks solved by at least one of three runs. Cost is estimated using parallelized output tokens (an idealized decoding-latency proxy), total verifier output tokens, and reference token prices applied to generator and verifier tokens.

For distillation, the authors collect 117k pairwise responses from GPT-5.6 Sol, then use LoRA to train TMAX-9B to reproduce the teacher's reasoning, scores, and preference label, activating the LoRA only for verification. An offline benchmark of stored trajectories is used to measure agreement with the teacher verifier — explicitly a measure of verification agreement, not action correctness or trajectory success.

Why This Matters

Impact on research: The paper reframes action-level compute scaling as a distinct and complementary axis to trajectory-level scaling (Best-of-T, Sequential Refine). It shows that the value of sampling is bounded by verification quality, and provides a controlled framework (fixed generator, fixed harness) for studying that interaction. It also connects pairwise/generative verification ideas from reasoning tasks to agentic settings where actions irreversibly change the environment.

Real-world applications:

  • Software engineering agents that must choose among candidate shell commands, code edits, or package installs before committing them to a repository.
  • Data science and scientific computing pipelines, where an incorrect command can corrupt intermediate state or produce silently wrong results.
  • Automated harness development and evaluation tooling, since Mid-Harness requires no change to model weights or serving architecture.
  • Long-running automation where an environment run is expensive, making in-trajectory verification cheaper than spawning additional full trajectories.

Industry relevance: Mid-Harness is a drop-in wrapper around the model call, so it can be deployed without retraining the generator or rewriting the harness. The cost analysis matters commercially: combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone, and decision-only verifiers cut reference-priced token cost further. The paper's ethics statement explicitly warns that improved terminal-agent capability does not establish command safety and that deployment should retain restricted permissions, environment isolation, and human approval for sensitive or irreversible actions.

Future Directions

  • Closing the distillation gap. The distilled verifier reaches 57.14% Pass@1 versus 68.03% for frontier verification with the same TMAX-9B generator and eight candidates; the authors motivate reinforcement learning for verifiers, world models that predict action effects, and jointly training generation and verification.
  • Evaluating action verification directly. The current evaluation lacks gold action labels, limiting measurement of verification correctness and candidate coverage. Because executing a different action changes all subsequent states, richer evaluation would need environment cloning or restoration to support branching, tree, or beam search.
  • Reducing verification compute. Open questions include adaptive verification that varies candidate width and comparison effort across steps, and whether action scaling is complementary to simply increasing a single candidate's reasoning effort at matched inference cost.
  • Extending beyond terminal agents. The paper suggests testing whether Mid-Harness improves agents for computer use and robotics, and whether verifier rationales can feed back into trajectory refinement.

Target Audience

Researchers and engineers working on LLM agents, test-time compute scaling, and verifier or process-reward-model design. It is most useful to those building terminal or software-engineering agents who need to decide where to spend inference budget — action sampling and verification versus additional full trajectories — and to practitioners who want a deployment-friendly method that leaves the generator and harness untouched. Readers seeking step-level ground-truth validation of verification correctness will find the paper identifies that gap rather than closing it.

Authors’ abstract

Terminal agents act through stochastic model generations, yet the ability to generate a useful action does not ensure its reliable execution. A poor command (e.g., wrong package install) can change the environment in ways that hinder subsequent progress, even when the model could generate a better alternative. We investigate whether allocating test-time compute at the model-harness boundary can improve action reliability and trajectory success, and what makes this allocation effective. To study these questions, we introduce Mid-Harness, which samples and verifies candidate actions before forwarding one for execution, while keeping the generator and harness unchanged. With a TMAX-9B generator, more action sampling yields little benefit under weak verification, whereas a capable verifier can exploit useful alternatives from the same generator. On TerminalBench-Lite, a GPT-5.6 Sol verifier raises Pass@1 from 50.00% for the base agent to 68.03% with 8 sampled actions. When the same TMAX-9B model serves as the verifier, pairwise verification performs best among the evaluated verification mechanisms. Distilling responses from the stronger verifier into TMAX-9B further improves Pass@1, while leaving the action generator unchanged. With TMAX-9B on TerminalBench-Lite, combining action and trajectory scaling reaches higher success at lower estimated token cost than generating more trajectories alone. Mid-Harness also improves performance across additional models, benchmarks, and harnesses. These findings identify action scaling as a promising target for test-time compute scaling in terminal agents.

Read the original paper