Skip to content
AI.info

Research

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?

Overview Research area: Large language model agents, agent harnesses, and reliability of external procedural guidance. Technical level: Advanced (the paper combines benchmark design with counterfactua

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
arXiv
2609.39578
Published
2026-09-30
Authors
Minghan Wang, Boyuan Wang, Jinhang Zuo, Yuxin Tao, Fang kong

AI summary

Overview

  • Research area: Large language model agents, agent harnesses, and reliability of external procedural guidance.
  • Technical level: Advanced (the paper combines benchmark design with counterfactual supervised fine-tuning and outcome-based reinforcement learning).
  • Scope: The paper defines "thinking outside the box" as selective reliance on external guidance, introduces Box²-Bench to measure it, and tests whether two open-weight models can learn it from bad workflows alone.

What This Paper Is About

Agent harnesses improve language models by injecting human-designed workflows covering planning, feedback, and tool use, but this assumes the human guide is more reliable than the model. The paper asks whether models can do the opposite of blind compliance: follow external procedural guidance when it helps, resist it when it misleads, and move beyond it when a better strategy emerges. To answer this, the authors build Box²-Bench, which holds the model and task fixed while varying only workflow reliability, and then test whether this capability can be trained.

Key Contributions

  1. Formalizes "thinking outside the box" as selective workflow reliance, distinct from task-solving capability, and introduces Box²-Bench, which varies workflow availability and reliability across five matched conditions: No workflow (∅,∅,∅,∅,∅), Good (+,+,+,+,+), Partial (+,+,+,∅,∅), Mixed (+,+,+,−,−), and Bad (−,−,−,−,−), where + is useful, − is misleading, and ∅ is absent guidance.
  2. Provides matched paired workflow effects — utilization (Δ_use = S_G − S_0), robustness (Δ_bad = S_B − S_0), and recovery when guidance stops (Δ_stop = S_P − S_G) or becomes misleading (Δ_switch = S_M − S_P) — evaluated across four frontier model–task settings and two open-weight settings.
  3. Tests two training strategies on bad workflows only, reserving good workflows for evaluation: counterfactual supervised fine-tuning (SFT) and outcome-based reinforcement learning (RL_env), showing they shape complementary aspects of reliance.
  4. Shows preliminary transfer beyond workflows to multi-agent collaboration in Economy of Minds (EoM) and to memory-augmented reasoning on LongMemEval-V2-Small, without setting-specific training.

Main Findings

  • Utilization and robustness are separate capabilities. Good workflows improve performance in eight of the twelve frontier model–task pairs, including gains of 5.9, 13.7, and 16.8 points on AutomationBench. Bad workflows reduce performance in all twelve pairs, with drops ranging from 2.3 to 43.3 points. The open-weight models show the same separation: good workflows improve all six model–task pairs, while bad workflows hurt five of six.
  • Scaling does not reliably remove sensitivity to misleading guidance. On AIME 2026, Qwen3-0.6B, Qwen3-4B, and Qwen3-8B all show positive Δ_use (+6.7, +13.3, +6.7) and negative Δ_bad (+3.3, −20.0, −26.7). On WebShop, Qwen3.5-4B, Qwen3.5-9B, and Qwen3.5-27B show Δ_use of +15.0, +10.4, and +12.2 with Δ_bad of −5.6, −9.8, and −11.4.
  • Reliance can persist after guidance becomes unreliable. Mixed workflows underperform their matched Partial workflows in ten of the twelve frontier model–task pairs. Because the two conditions share the same useful prefix and change point, their difference isolates the effect of continuing with misleading guidance rather than receiving no further guidance.
  • Counterfactual SFT reduces dependence on misleading workflows. On AIME, the bad-workflow penalty shrinks from 20.0 to 6.7 points; on WebShop, it shrinks from 9.8 to 0.6 points. SFT is accompanied by weaker reliance on held-out good workflows (Δ_use on AIME falls from +13.3 to −3.3), showing it first makes workflow guidance defeasible rather than mandatory.
  • Outcome-based RL recalibrates rather than reverses SFT. On AIME, RL_env restores positive utilization of held-out good workflows, raising Δ_use from −3.3 to +6.7 while retaining a robustness improvement over the base model. On WebShop, the balance established by SFT remains largely stable.
  • Training improves adaptation when reliability changes. On AIME, the Mixed–Partial gap improves from −14.4 for the base model to −4.4 after SFT and +1.1 after RL. On WebShop this gap is already near zero and remains small across training stages.
  • Selective reliance extends to peer information. In the adaptive EoM environment with four agents sharing weights, SFT raises episode success from 18.00% to 22.67%, while SFT+RL achieves 18.67%. SFT produces 11 repairs and one corruption, compared with two of each for Base; SFT+RL produces five repairs and three corruptions. Gains are not monotonic across training stages.
  • Selective reliance extends to stored memory. On LongMemEval-V2-Small, the base model scores 21.1% with no memory brief, 95.6% with good memory, and 10.0% with bad memory. After SFT these become 28.9%, 92.2%, and 14.4%; after SFT+RL_env they become 25.6%, 94.4%, and 13.3%. The base model consults the archive less often under corrupted memory (65.6% tool rate) than reliable memory (72.2%), whereas training removes and eventually reverses this gap.

Methodology in Plain English

The authors fix the task, environment, evaluator, and model, and change only the workflow given to the model. A fixed external model (Qwen3.5-Plus-02-15) acts as a scalable proxy for a human workflow designer. Using only information available during normal execution, it generates a useful workflow W⁺(x) and, for each change point k, a plausible but misleading continuation. These define the Good, Bad, Partial (good prefix only), and Mixed (good prefix followed by bad suffix) conditions. An independent verifier checks useful workflows for validity and task relevance, misleading workflows for plausibility and task relevance, and all workflows for answer leakage; accepted workflows are frozen and shared across target models.

For training, the authors use only bad workflows and hold out good ones. Counterfactual SFT pairs a problem and its misleading workflow with a verified correct solution (2,555 prompt–completion pairs for math; 8,203 turn-level examples for WebShop), so reducing the loss increases the probability of success despite the misleading guidance. Because the objective only constrains behavior under bad workflows, both selective rejection and blanket ignoring can fit the data — preserving held-out utilization must arise through generalization.

Outcome-based RL then continues from the SFT checkpoint using task-level rewards only (binary answer correctness for AIME, environment return for WebShop), never rewarding agreement with the workflow. Training uses Group Relative Policy Optimization with group-relative advantages, a clipped token-level loss, and an optional KL penalty, retaining only groups with non-degenerate rewards.

Math training data comes from 2,060 InT-SFT problems, with Qwen3-8B sampled sixteen times per problem; workflows are five steps, and 111 bad-workflow prompts are used for RL. WebShop training uses the official split with global task indices 1500 to 12086, eight-step workflows, 2,023 workflow pairs, 153 retained RL tasks (139 for training, 14 for in-domain monitoring). Evaluation uses all 30 AIME 2026 problems with eight completions each at temperature 0.7 and top-p 0.95 (240 generations aggregated into 30 predictions) and 500 official WebShop test tasks with greedy decoding and at most 14 model actions. The training experiments use Qwen3-4B on AIME 2026 and Qwen3.5-9B on WebShop.

Why This Matters

The paper argues that raw task accuracy does not capture whether an agent can regulate its reliance on external guidance, and that harness effectiveness is highly dependent on the model and task — a workflow that helps one executor can become harmful in another. As models grow more capable, unreliable guidance can increasingly constrain execution rather than support it. This reframes selective reliance on fallible external information as a dimension of agent reliability in its own right.

Real-world applications:

  • Coding agents that receive repository-specific or team-prescribed workflows and must decide when a prescribed procedure is worse than their own plan.
  • Web search and information-retrieval agents that follow human-designed browsing procedures, which may be incomplete or actively misleading for a given query.
  • Tool-use and automation agents that operate from human-authored playbooks and need to detect when a step violates task constraints.
  • **Multi-agent systems and memory-a

Authors’ abstract

Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two complementary training strategies: counterfactual supervised fine-tuning improves robustness, while outcome-based reinforcement learning can shift the balance toward greater use of helpful workflows. We further find that this behavior extends beyond workflows to other forms of external information, improving peer correction and robustness to corrupted memory. Together, our results identify selective reliance on fallible external information as a dimension of agent reliability not captured by task performance alone.

Read the original paper