Skip to content
AI.info

Research

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions

Overview Research area: LLM-based tool-using and retrieval agents, specifically the sequential decision of when to abandon a source that keeps returning useless results. Technical level: Intermediate.

Judged Useless, Queried Anyway: Tool-Using Agents Rarely Turn Their Own Evidence Judgments into Stopping Decisions
arXiv
2610.06191
Published
2026-10-05
Authors
Chubin Zhang, Zhenglin Wan, Xingrui Yu, Jingxuan Wu, Yaxin Zhou, Ivor Tsang, Bo An

AI summary

Overview

Research area: LLM-based tool-using and retrieval agents, specifically the sequential decision of when to abandon a source that keeps returning useless results.

Technical level: Intermediate. The paper assumes familiarity with the ReAct-style agent loop, retrieval-augmented question answering, and basic statistics (Mantel–Haenszel risk difference, Holm correction, bootstrap intervals), but it defines its own central measures and explains its environment concretely.

Scope: The paper introduces a controlled retrieval environment with injected source failures to test whether seven agents turn their own judgments that a source has failed into the action of stopping, and whether prompting or an enforced harness rule can supply the missing step.

What This Paper Is About

Tool-using agents often keep calling a search tool that has stopped returning anything useful. In the paper's opening example, Claude Haiku 4.5 calls every result irrelevant on a two-hop question, recalls at step 6 that Barlow is in a band, and still searches until its budget runs out. The paper's goal is to separate three things that earlier work conflated — how agents judge a result, what they believe about continuing, and what they do — and to test whether telling agents more (permission to answer from memory, a stated budget, a stated stopping rule or call price) makes them stop on the evidence.

Key Contributions

  1. A controlled retrieval environment with known failures. Questions come from the HotpotQA distractor development set (Yang et al., 2018), each with its own ten-paragraph knowledge base in which two paragraphs hold the supporting facts. Six failure regimes — clean, persistent, recover_after_c for c in {1,2,3}, and late_onset_from_3 — let the authors know both whether each result is useful and whether the source will recover. A transposed setting (292 balanced FEVER claims, Thorne et al., 2018) is used for transfer.

  2. A measurement design that separates judgments, beliefs and actions. A side channel asks a one-word USEFUL/USELESS question on a scratch copy of the conversation; a separate scratch copy elicits the agent's probability of success if it continues versus answers now. The central statistic, the time-matched contrast Δ, compares answering at the same step after an unbroken run of self-judged useless results against answering after a history containing a result judged useful.

  3. Evidence of a judging-doing dissociation. Every model whose judgments were recorded calls a failing source's results useless 97–100% of the time, yet unaided stopping rarely follows those judgments: every unaided open model has a negative Δ, and after five consecutive useless judgments five of the seven agents answer on at most 7% of questions.

  4. A minimal enforced fix plus replication. Letting the harness act after k = 5 consecutive self-judged useless results raises mean6 success for every model tested, keeps the stopping point fixed when the budget doubles, and replicates on 300 fresh questions (all nineteen pre-registered tests pass) and on FEVER.

Main Findings

  • Judgments are accurate. Through the side channel, every model whose judgments were recorded calls a failing source's results useless 97–100% of the time on HotpotQA and FEVER. Models call a gold paragraph useful 77–98% of the time when first retrieved, except Qwen2.5-7B at 57%, so the pattern is not a blanket "useless" bias.

  • Actions do not follow judgments. After judging five consecutive results useless, unaided Qwen3-8B answers on only 1 of 292 questions. Five of the seven agents answer on at most 7% of such questions. Hazard modelling gives each further useless judgment a change of −0.02 to −0.01 in the unaided 7–8B agents' answering probability, against +0.15 to +0.17 under the enforced rule.

  • The same holds for judgments stated in reasoning. Stated judgments agree with the side channel on 87–96% of decision points; computing Δ from them stays below zero, and after five stated useless judgments the agents answer on 0–0.4% of questions (Qwen3-32B: 6%).

  • Stopping would often have paid. Answering at the five-useless-judgment point succeeds 15–17% of the time versus at most 7% for continuing (7–8B models, exploratory). Forced to answer there, Claude Haiku 4.5 is correct on 114 of 284 questions, whereas unaided it answers only one of them.

  • Beliefs point in different directions. Qwen3-8B overestimates the value of continuing, Qwen2.5-7B overestimates the value of answering now, and Llama-3.1-8B underestimates it. Where their own reports favor answering now, Llama-3.1-8B still answers at only 32 of 595 decision points and Qwen2.5-7B at 267 of 2,443.

  • Permission yields indiscriminate stopping. With permission to answer from memory, Qwen3-8B answers before any real result on 55% of questions whose source recovers after three failures, and its Δ drops to −0.32.

  • A stated budget moves stopping to the deadline. Budget-cued 7–8B models answer on the last action (51–79% of answers on a failing source), and their Δ stays negative. This condition matches or exceeds the enforced rule's success for Qwen3-8B and Llama-3.1-8B and exceeds it for Qwen2.5-7B, so success alone cannot distinguish the two. With the budget doubled, these agents move their answers from the eighth to the sixteenth action and roughly double their tool calls with no gain in success.

  • Stating the rule or the price is at most partly followed. With the rule stated in words, only Llama-3.1-8B and Qwen3-32B follow it, and only partly (Δ = +0.09 and +0.08, against +0.37 and +0.33 when enforced). With a price of 0.05 per call, the 7–8B models still answer at the deadline; only Qwen3-32B stops partly on its judgments. Claude Haiku 4.5 behaves similarly — with the rule stated its stopping follows the evidence only partly, and with the price stated it answers on almost every question regardless of evidence.

  • A backup tool and a running count do not fix it. With a second, intact backup_search, no unaided agent becomes more likely to leave as useless results accumulate (Δ ≤ +0.02); disabling the failing tool after five useless judgments does make leaving evidence-driven (Δ = +0.26 and +0.30 for Llama-3.1-8B and Qwen3-8B). Showing agents their running count of useless judgments still yields Δ of −0.14 to −0.24 for the three 7–8B models.

  • The enforced rule works. Its positive Δ (+0.33 to +0.40) is expected by construction, but mean6 rises for the four open models and Claude Haiku 4.5 (+0.026 to +0.097, Holm-significant), it is non-inferior where the source recovers or never fails in 18 of 20 tests (margin −0.05; Qwen3-8B fails narrowly twice), and the answering action stays fixed when the budget doubles. Against a rate-matched random signal, the enforced rule wins by 0.011 to 0.029; a cheap lexical detector does about as well as the agent's own judgments.

  • Deadline and integration combine. Budget plus enforced rule is non-inferior to both components in 14 of 14 pre-registered tests and beats the enforced rule by +0.138 for Qwen2.5-7B, whose main failure is answering too early.

  • Reading judgments from reasoning avoids extra calls. The enforced rule makes three to five judgment calls per question; a keyword reader on the agent's own thought keeps success within 0.01 of the enforced rule for Llama-3.1-8B and the Qwen3 models, and combined with the budget statement beats the budget condition alone at 0.05 per call for Llama-3.1-8B and Qwen3-8B. It does not help Qwen2.5-7B, which rarely states a judgment.

  • Replication and transfer. On 300 fresh questions with Δ as the pre-registered primary endpoint, all nineteen tests pass; Δ stays below zero unenforced and reaches +0.30 to +0.35 under the enforced rule, which beats the unaided agent by +0.030 to +0.061. On FEVER the enforced rule raises mean6 by +0.129 to +0.201, and unaided agents answer after five useless judgments on at most 3% of claims.

  • Harder, on-topic failures weaken judgments but not the dissociation. With the question's own distractor paragraphs (plausible) or its own pages with supporting facts removed (answerless), judgments are 84–97% and 76–97% useless respectively, and unaided agents answer on 2–23% and 5–21% of questions after five useless judgments; the enforced rule still raises success in both.

  • Larger, reasoning and search-trained agents rarely integrate either. Qwen3-32B answers after five useless judgments on only 19 of 284 questions. Claude Sonnet 5 is the only unaided model whose stopping partly follows the evidence (design-label contrast +0.12 [+0.05, +0.18], exploratory), yet it exhausted its budget on 45% of questions where its enforced rule fired; its enforced-rule gain over six regimes on 100 questions is not significant (+0.022, p = 0.051), though on a failing source its success rises from 0.40 to 0.51. Search-R1 (Jin et al., 2025), trained with PPO on HotpotQA and NQ, answers from memory after one or two searches on 27% of questions but searches to the end of its budget on 45%. Qwen3-32B in thinking mode answers from memory on 97% of failing-source questions (11% without thinking).

Methodology in Plain English

The authors build a small, fully known world. Each question comes with ten paragraphs, two of which actually contain the answer, and an agent that follows the ReAct format with three tools: search, lookup and finish. A budget of eight actions is enough to answer a two-hop question after several failed searches, and an answer counts as correct at token F1 ≥ 0.6.

Failures are injected by replacing an observation with a paragraph from a different question's knowledge base, which is almost always useless, so the authors can score whether the agent noticed. Six regimes vary when failure occurs — clean, persistent, recover_after_c for c in {1,2,3}, and late_onset_from_3 — so that neither always leaving nor never leaving is the right strategy. Success is averaged over the six regimes with equal weight (mean6).

At each decision point the authors record three things separately: the agent's judgment of the latest result (stated in its thought, or asked as a one-word USEFUL/USELESS question on a scratch copy of the conversation it never sees again), its beliefs (reported probabilities of success for continuing versus answering now, on another scratch copy), and its action.

The key trick is comparing histories at the same step. If an agent integrates its judgments, it should answer more often after an unbroken run of self-judged useless results than after a history that contains a result it judged useful. The difference is Δ, pooled over decisions after three to six results with the Mantel–Haenszel risk difference and excluding the final action (the last chance to answer). A clock- or deadline-driven agent has Δ = 0, an integrating agent has Δ > 0, and an agent that answers on earlier useful evidence has Δ < 0. The authors also report the simpler integration index I = h({4,5}) − h({1,2}), noting it is confounded with time — a simulated clock-driven agent reaches +0.5 on it — which is why Δ is the primary measure.

Conditions add one element at a time: permission to answer from memory; a stated budget (eight actions, a counter, unanswered counts as wrong); a stated rule in words; a call price of 0.05 against 1 per correct answer; a decide condition showing the running count; an enforced rule where the harness permits only finish after k = 5 consecutive self-judged useless results (k chosen on 100 development questions); and a combination of budget plus enforced rule. Random

Authors’ abstract

An agent whose tool keeps returning nothing useful should stop relying on it. In a retrieval environment with controlled source failures, we separate how agents judge results from what they do. We compare stopping at the same step after longer and shorter runs of results the agent judged useless; this contrast is zero for clock- or deadline-driven stopping. Where we record their judgments, the seven agents we test call a failing source's results useless 97-100% of the time, yet most of them rarely stop on that judgment. Prompt cues change when they stop but not what they stop on. Permission to answer from memory and a reasoning mode can bring early stops regardless of evidence, a stated budget moves the 7-8B models' stops to the deadline, and a stopping rule or call cost in the prompt is followed at most partly. Stopping follows the evidence only when the harness enforces an integration step that makes the agent answer after five consecutive results it judged useless. This step raises failing-source success for every model, keeps the stopping point fixed when the budget doubles, and needs no extra judgment call when the agent states its judgments. A pre-registered replication on 300 fresh questions confirms the dissociation and the rule's effect.

Read the original paper