Research
Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents
Overview Research area: AI agents and tool use — specifically proactive information seeking by LLM-based agents, evaluated on multi-hop question answering and customer-service tasks. Technical level:

- arXiv
- 2609.37236
- Published
- 2026-09-29
- Authors
- Ido Levy, Asaf Yehudai, Segev Shlomov, Asaf Adi, Leshem Choshen
AI summary
Overview
Research area: AI agents and tool use — specifically proactive information seeking by LLM-based agents, evaluated on multi-hop question answering and customer-service tasks.
Technical level: Intermediate. The paper assumes familiarity with retrieval-augmented agents, direct preference optimization (DPO), and multi-hop QA benchmarks, but its central idea (labeling a question by what it retrieves next) is explained concretely.
Scope: The paper defines two axes of proactive information seeking — horizontal (needs already nameable from the current state) and vertical (needs that only earlier evidence reveals) — makes them measurable with a need graph recovered from benchmark decompositions, and introduces Q&D, a training method that teaches a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge.
What This Paper Is About
Tool-using agents typically respond to what the user explicitly asks, but finishing a task often requires information the user never mentions — an order number, an account record, a value that only appears in a document the agent has not yet retrieved. Existing work on proactive agents mostly asks whether and when an agent should act on its own; this paper instead studies the content of proactivity: what information an agent should pursue unasked, and when it should stop pursuing it. The goal is to make that content measurable and trainable without a model judge.
Key Contributions
-
A measurable content axis of proactivity. The paper defines horizontal and vertical proactivity on a need graph — the units of evidence a task requires, with a prerequisite edge wherever one need becomes nameable only after another is resolved — and proposes evaluation metrics for both forms plus stopping behavior (required-evidence coverage, breadth, depth-weighted recall, deepest need resolved, out-of-order rate, stop-when-done, ask-when-not-done).
-
Q&D (questioner and drafter), a training algorithm. An agent is split into a questioner that asks one question at a time or stops, and a frozen drafter that folds retrieved evidence into a draft. Because the drafter is frozen, every change of state is caused by a question, so each question can be credited with what followed it — no reward model or judge is used.
-
Empirical results on three multi-hop QA benchmarks. At equal retrieval spend, the trained 8B questioner recovers more required evidence than the same model, prompted, and beats a prompted model 15 times larger in the same role on two of three benchmarks, with gains that persist after controlling for question volume and length.
-
Transfer to a customer-service agent. Without further training, the questioner is placed in an interactive agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail outperforms the 15 times larger model with fewer follow-up turns from the customer.
Main Findings
-
More required evidence at equal spend. After the same number of questions, the trained questioner recovers more of each task's required evidence than the same model, prompted, by 11.2 percentage points on MuSiQue, 7.0 on StrategyQA and 5.0 on 2WikiMultiHopQA (Table 2). Each training seed alone is decided on all three suites, and all four training runs agree within a standard deviation of 0.8 points.
-
The gain is deep, which is vertical proactivity. Depth-weighted recall rises by 12.5, 5.4 and 5.5 points on the three suites, most on MuSiQue, whose chains run deepest. The deepest need a run resolves lies deeper by 0.22 levels on MuSiQue and 0.10 on StrategyQA. Breadth rises as well, so the questioner goes deeper without narrowing its search.
-
Coverage table values (Table 2, equal spend): MuSiQue required-evidence coverage 78.3 to 89.5, StrategyQA 78.4 to 85.5, 2WikiMultiHopQA 88.2 to 93.1. Against GPT-OSS-120B the coverage figures are 77.8 versus 84.8 (MuSiQue), 80.7 versus 85.0 (StrategyQA), and 92.4 versus 88.9 (2WikiMultiHopQA).
-
One side effect on MuSiQue. The trained questioner more often finds a deep piece of evidence before the one it depends on, by 5.7 points (the out-of-order rate). The paper attributes this to paragraphs retrieved by chance: the questions themselves skip ahead no more often than the prompted model's, by 1.2 points spanning zero.
-
Cheaper evidence. For each required need found, the trained questioner spends 40% fewer tokens on MuSiQue, 11,610 against 19,354, and fewer on the other two suites as well.
-
The gain comes from what is asked, not how much. Replacing the trained questioner's questions with as many drawn at random from other tasks of the same benchmark loses 34.7 points of coverage on MuSiQue and 63.3 on StrategyQA. Hiding the evidence from its state costs 8.4 and 8.7 points of coverage on MuSiQue (per training seed). Its questions are longer, 17.5 words against 13.1 on MuSiQue, but when the prompted model may spend as many question tokens it still recovers 14.9 and 4.7 points less on MuSiQue and StrategyQA development tasks.
-
A trained 8B questioner can lead a 15 times larger prompted model. At equal spend it beats GPT-OSS-120B on MuSiQue and StrategyQA, by 7.0 and 4.3 points, and trails it on 2WikiMultiHopQA by 3.5 points. The lead is largest where chains run deepest and reversed where needs lie at most one step deep.
-
Efficiency is visible under call budgets. On MuSiQue the trained questioner leads both prompted models at every budget from 4 to 24 calls, and allowed only 4 calls it recovers 4.1 points more of the required evidence than GPT-OSS-120B allowed 24. Elsewhere prompted models catch up at larger budgets: on StrategyQA it ties the same model with fewer questions and trails GPT-OSS-120B once that model spends more, and on FRAMES it trails GPT-OSS-120B at budgets of 16 and 24.
-
A hand-designed retrieval algorithm does not close the gap. PAR2-RAG never beats plain prompting of its own model at equal spend on any of three models, and the trained questioner leads it on every model on MuSiQue. The one comparator ahead at equal spend is Claude Opus 5, prompted plainly, on MuSiQue, by 2.9 points.
-
Stopping on its own (at most eight calls). The trained questioner asks fewer questions than either prompted model on every suite and keeps its lead on MuSiQue, by 9.6 points over the same model, prompted, and 10.1 over GPT-OSS-120B. It uses its whole budget in only 34% of MuSiQue runs when allowed 4 calls and 8% when allowed 24, against 13% to 69% for the prompted questioners across both suites and budgets. It stops more often once the work is done, by 29 to 33 points over the same model, prompted, on every suite and seed, and by 25, 43 and 12 points over GPT-OSS-120B. Where evidence is still missing it keeps asking as often as the prompted model on MuSiQue and 2WikiMultiHopQA, but 7 points less often on StrategyQA.
-
The extra evidence does not yet reach answers. On MuSiQue the trained questioner finds the evidence the answer rests on 9.7 points more often, which at the prompted model's answer rates implies an answer gain of about 4.3 points — below the 4.7 the sample can detect and consistent with the 2.4 observed. On StrategyQA it stops before finding that evidence, by choice in 92% and 86% of the runs that end without it, against 39% for the prompted model.
-
What each training signal teaches (Table 3, development tasks). Imitation copies good decisions (+4.9 coverage on MuSiQue, +8.0 on StrategyQA) but drops 24,538 unfinished states where no sampled question clears the floor. Question pairs raise coverage most (+16.7 and +8.4) but stop too early. Stop contrasts alone almost never stop (10.1% and 9.7% stop-when-done) and give no coverage gain. The final stage, trained on both, raises coverage about as much (+15.7 and +15.4 for the two seeds on MuSiQue; +8.2 and +7.8 on StrategyQA) while asking when work remains, though it stops less often than imitation once the work is done.
-
Human raters recognize the behavior but disagree on the best question. Two raters agree on 88% of candidate questions on whether they reach a need the request never states (Cohen's kappa = 0.76), but on only 68% of the pairs both judged about which question is the better move (kappa = 0.24), and each agrees with the paper's rule on 65% and 68% of pairs, against 50% by chance.
-
Customer-service transfer (tau-squared-bench). In retail, success rises from 13% and 12% to 34% and 32% under two prompt variants, better on ten of twenty-five tasks and worse on none in each; in airline it rises by 6.9 and 6.4 points. It asks 1.3 and 1.4 fewer questions per retail dialogue. GPT-OSS-120B, prompted, completes 17% and 19% of retail tasks, and the trained questioner beats it by 17.0 and 12.7 points of success with 1.6 and 1.8 fewer follow-up turns from the customer, both decided. Against the same model, prompted, follow-up turns do not differ detectably, and in airline no difference from GPT-OSS-120B is decided.
-
Composed and synthetic tasks. On synthetic tasks built from two to four independent lines of inquiry, each several steps deep, the trained questioner advances more lines than the same model, prompted, when a task has two or three lines, ties with four, and recovers more of the deep evidence in every case. When two held-out MuSiQue questions are joined into one task whose answer needs both chains, it advances more of the two chains past their first step and recovers 15.2 points more of the required evidence.
-
Need graph statistics (Table 4). MuSiQue: 800 graphs, 2,660 nodes, 1,860 edges, depth histogram 0:1,196, 1:800, 2:530, 3:134, with 55.0% of needs behind an edge. StrategyQA: 2,290 graphs, 6,720 nodes, 4,483 edges, depths 0:3,721, 1:2,320, 2:615, 3:60, 4:4, 44.6% behind an edge. 2WikiMultiHopQA: 12,576 graphs, 31,120 nodes, 11,956 edges, depths 0:19,164 and 1:11,956, 38.4% behind an edge.
Methodology in Plain English
Making the invisible measurable. The researchers noticed that multi-hop QA benchmarks already ship a decomposition: a later step refers to an earlier one (for example, MuSiQue writes "What is the birthplace of #1?"). They mechanically recover a graph from that structure — every one of MuSiQue's 1,464 needs below the surface carries such a reference, and none of the 1,196 on the surface does; StrategyQA carries the same references on 2,999 of its 6,720 steps; 2WikiMultiHopQA links a need to another whose subject is the first's object. No person and no model wrote any node or edge. Needs at depth zero can be named from the request; deeper needs can only be named after the evidence above them. A run is then scored from its transcript against this graph, with no model judge.
Splitting the agent in two. Q&D separates the agent into a questioner that decides what to ask or whether to stop, and a frozen drafter that rewrites the draft from the retrieved evidence. A frozen answerer writes the final answer in every arm under a fixed word cap, so only the questioner differs between policies and answer length favors none of them. The questioner sees the task, the evidence so far, the current draft and its past questions; the draft exposes the frontier, so at each step it can go deeper along the chain the last evidence opened or open a need already nameable beside it.
Training on consequences. A recorded run is forked at a step, eight alternative questions are sampled there, and each is continued to the end by the policy that produced the run. Because the candidates share the task, the evidence and the history, they differ only in the question asked. Pairs are ordered by consequence alone: first by whether the run answered the task (which decides 6% of the pairs trained on), then by which reached the complete evidence sooner, then by how much evidence each turn added. At a state the need graph marks unfinished, asking is ranked above stopping. Training proceeds in three stages: imitation of good decisions, then direct preference optimization on question pairs, then question pairs and stop contrasts together.
Setup. The questioner is Qwen3-8B with a low-rank adapter, reported over two training seeds of its final stage. The main comparator is the same model, prompted with the same template; beside it are GPT-OSS-120B (a model 15 times larger) prompted, Claude Opus 5 prompted plainly, a baseline that never asks, and PAR2-RAG reimplemented with the paper's retriever and drafter. Retrieval is BM25 over each task's released pool of 20 paragraphs (10 on 2WikiMultiHopQA), returning the top 5, 2 and 3 on MuSiQue, StrategyQA and 2WikiMultiHopQA, and cost is counted in retrieval calls, one per question. Results are read on held-out test splits of 200 tasks per benchmark, at two rollout seeds each; checkpoints were selected on 132 MuSiQue and 333 StrategyQA development tasks. Two readings are kept apart: equal spend (both policies read at the lower of their two question counts) and own stop (each policy decides for itself, with at most eight calls). Intervals are bias-corrected and accelerated bootstraps over tasks at 10,000 resamples, and a cell is called decided only when intervals from three independent 50,000-resample bootstrap runs all exclude zero.
Why This Matters
Impact on research. The paper argues that proactivity has a content dimension that prior work — which mostly studies whether and when an agent acts — leaves unmeasured. By grounding horizontal and vertical proactivity in a need graph recovered mechanically from existing benchmark decompositions, it offers a scoring
Authors’ abstract
An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence reveals. A need graph, recovered from a benchmark's own decomposition, records which needs depend on which, so both forms, and whether the agent stops at the right time, can be scored from a transcript without a model judge. To learn this behavior, we propose Q&D (questioner and drafter), which trains a questioner to prefer the question whose continuation retrieves more of the required evidence, with no reward model or judge. On held-out splits of three multi-hop question-answering benchmarks, at equal retrieval spend, the trained questioner improves both forms of proactivity over the same model, prompted, and outperforms a prompted model $15\times$ larger in the same role on two of the three, and the gain persists after controlling for question volume and length. Without further training, we place the questioner in an interactive customer-service agent with a simulated customer, where it completes more tasks while asking fewer questions, and in retail it outperforms the $15\times$ larger model with fewer follow-up turns from the customer. These results show that proactivity depends not only on whether an agent acts without being asked, but also on what it chooses to pursue and when it stops.