Research
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search Agents Overview Research area: Multimodal AI agents — specifically efficient reasoning for multimodal search agents built on sm

- arXiv
- 2610.01892
- Published
- 2026-10-01
- Authors
- Feiyu Gavin Zhu, Xiaoyu Zhu, Jiqi Yang, Rui Yang, Arnab Kumar Mondal, Yancheng Wang, Xinke Deng, Jean Oh, Reid Simmons, Joerg Liebelt, Xiang Kong, Zhongyu Jiang
AI summary
Selection-Based Structured Reasoning: Toward Efficient Multimodal Search AgentsOverview
Research area: Multimodal AI agents — specifically efficient reasoning for multimodal search agents built on small vision-language models.
Technical level: Intermediate. The core idea (reasoning as selection rather than generation) is intuitive, but the paper assumes familiarity with ReAct-style agents, autoregressive decoding, KV caches, and reinforcement learning objectives such as GRPO.
Scope: The paper proposes and evaluates Selection-based Structured Reasoning (SSR), a framework that replaces freeform token-by-token reasoning in multimodal search agents with selection from a small library of reusable natural-language reasoning candidates, and measures both task success and inference latency across seven benchmarks, four training objectives, and two model sizes.
Note on the content: The version provided is truncated. Several appendices (dataset details, the reasoning library text, hyperparameters, and pseudocode continuations) are referenced but not fully reproduced, so some specifics cannot be verified from the supplied text and are flagged as not reported below.
What This Paper Is About
Multimodal search agents typically "think" by generating a freeform reasoning trace in natural language before every action, and for small on-device models this generated text can be long and unhelpful while costing substantial inference time. SSR changes the representation of reasoning itself: instead of generating reasoning, the agent picks one of a handful of pre-written reasoning candidates based on how likely that candidate is given the current context, then generates only the action. The goal is to keep task success competitive with state-of-the-art search agents of the same scale while cutting reasoning latency dramatically.
Key Contributions
-
Reasoning by selection instead of generation. SSR replaces freeform reasoning in ReAct-style agents with selection among reusable natural-language reasoning candidates conditioned on the current interaction history.
-
Parallel reasoning decoding. The language model itself scores each candidate by its length-normalized likelihood, using no additional classification head and no SFT warm start. Because every candidate text is pre-specified, teacher-forced prefilling computes all token likelihoods concurrently within a candidate and scores candidates concurrently using a shared KV cache — removing autoregressive generation of the reasoning trace entirely.
-
Competitive performance with large efficiency gains. Across GRPO, GSPO, SAPO, and SFT, the 4B agents maintain competitive performance while reducing mean reasoning latency per turn by over 90% and mean model latency per question by 28–54%.
-
A characterization of what makes selection work. Ablations show that scoring the full reasoning text outperforms scoring only a candidate index, that additional library entries improve success monotonically, and that the same checkpoint produces different reasoning-selection distributions across different benchmarks.
Main Findings
-
Average success is competitive with leading 4B agents. The 4B SSR model with GRPO reaches 61.37% average success across seven benchmarks, described as comparable to the best 4B state-of-the-art model (TAPO with GSPO at 61.25%), and it outperforms the 4B MMSearch-R1 (57.66%) and SenseNova-MARS (51.65%) baselines.
-
Smaller backbone also works. The 2B SSR agent reaches 51.26% average success, which the paper describes as comparable to the zero-shot 8B model (Qwen3-VL-8B-Instruct at 51.33%).
-
Latency reduction is large and consistent. Per-turn reasoning latency falls by more than 90% across objectives and sizes, and per-question model latency falls by 28–54%. Example pairings: GRPO 4B goes from 0.896 s to 0.061 s per turn (-93.2%) and from 5.456 s to 2.514 s per question (-53.9%); SFT 4B goes from 0.770 s to 0.062 s (-91.9%) and from 5.308 s to 3.318 s (-37.5%).
-
Success rates are comparable but not uniformly higher. SSR improves average success over freeform GRPO at both sizes (58.65% to 61.37% at 4B; 49.98% to 51.26% at 2B) and matches freeform SFT (54.55% to 54.57%). It shows modest decreases under GSPO (60.45% to 58.60%) and SAPO (61.26% to 60.46%). The authors hypothesize this comes from the selector-gradient approximation, which becomes less accurate across multiple gradient updates per rollout.
-
Effective reasoning throughput rises by orders of magnitude. Throughput for GRPO-SSR 4B is reported as 3029.9 tokens/s versus 72.7 tokens/s for GRPO-freeform 4B (+4070.2%), consistent with processing all candidate tokens in parallel at prefill speed.
-
SSR beats four efficient-reasoning baselines. On the profiling subset, SSR reaches 58.36% success versus Chain-of-Draft 54.06%, Sketch-of-Thought 56.07%, efficiency-reward RL 56.28%, and Probe & Prefill 55.49%. SSR delivers roughly 7x or greater speedup in reasoning over all baselines; its mean model latency of about 2.5 s per question is about 30% lower than Sketch-of-Thought, the fastest baseline on that metric.
-
Latency is far more predictable. SSR's p95 reasoning time is 0.071 s, only 10 ms above its mean, while baseline p95 reasoning latency ranges from 0.798 s to 1.892 s. At the question level, SSR's p95 is 4.6 s, about a 30% reduction from the fastest baseline's 6.6 s.
-
Full-text scoring beats index-only scoring. Selecting using only the candidate index lowers average success by 3.9% (61.37% to 58.98%), affecting all seven benchmarks — evidence that the semantics of the reasoning tokens matter for context-dependent selection.
-
Autoregressive generation of library entries only works with an SFT warm start. With SFT warm start followed by RL with a format reward, the 4B model reproduces library entries and achieves average success of 60.01%, comparable to parallel decoding. Without SFT, the model degenerates to freeform reasoning (60.23% average, but not constrained to the library), showing the interface is enforced directly by selection rather than learned.
-
Bigger libraries help monotonically. Success rises monotonically from 37.24% with a single generic candidate ("I am a helpful assistant") to 61.37% with six candidates. The jump from one generic candidate to two candidates (answer or tool call) accounts for 17.27 percentage points.
-
The same checkpoint adapts its reasoning choices per benchmark. SimpleVQA shows the highest rate for "Answer from image" and the lowest for "Look up a fact," while the high-resolution-centric HR-MMSearch shows the highest rate for "Examine image detail." In 99.9% of turns, the generated action type matched the expected action type of the selected candidate.
-
Per-benchmark outcomes vary. 4B SSR scores 61.40 on MMSearch, 30.16 on HR-MMSearch, 70.22 on FVQA-test, 71.08 on SimpleVQA, 57.19 on LiveVQA, 78.67 on MAT-Search, and 60.90 on InfoSeek. The 2B model scores 51.18, 17.06, 61.16, 61.83, 51.54, 62.59, and 53.44 respectively.
Methodology in Plain English
The researchers start from a ReAct-style agent loop: at each turn, the agent sees its interaction history (the original question plus all prior reasoning, actions, and tool observations), produces a reasoning trace, then produces an action such as a text search, a reverse-image search, an image crop, or a final answer.
Step 1 — Build a small library of reasoning candidates. Instead of letting the model write reasoning, they hand-write six reusable natural-language candidates covering the general cases a search agent encounters: (1) do an image search to identify the entity when unsure, (2) look up specific information via text search, (3) zoom into part of the image for targeted inspection, (4) answer from gathered evidence, (5) answer from the image, and (6) answer with general knowledge.
Step 2 — Score all candidates at once. At each turn, the model is used directly as a classifier over the library. Each candidate receives a length-normalized log-likelihood score (with length-normalization exponent α = 1, so candidates are compared by mean token log-likelihood rather than summed probability, which would otherwise favor the shortest candidate). A softmax with selection temperature τ converts scores into a categorical distribution. Because every candidate's tokens are already known, the model can be teacher-forced to compute all token log-probabilities in a single batched forward pass, and all candidates can be scored concurrently against a shared KV cache for the history. No new tokens are generated during reasoning.
Step 3 — Inject the winner and generate the action. The selected reasoning text is inserted into context as natural-language guidance, and the model then generates the action and its arguments, such as a question-specific search query. Selection and action generation share the same parameters and system prompt, with no auxiliary head.
Step 4 — Train end to end. Under SFT, the loss is a weighted sum of a categorical selection loss against the ground-truth candidate and a token-level action loss (λ_reason^SFT = 1). Under GRPO-style RL, the token-level importance ratios for reasoning are replaced by a single categorical ratio for the selection, while action tokens keep their usual ratios; the objective combines clipped surrogate terms plus a reference-policy KL regularizer, with γ_reason^RL = 1 and β = 10⁻³. SSR can be trained directly with RL without an SFT warm-up, which the authors attribute to selection constraining reasoning to the library. Models are 2B and 4B Qwen3-VL, trained in a single stage for one epoch at learning rate 10⁻⁵ using outcome and format rewards, with the same tool harness and LLM-as-judge used by prior work.
Step 5 — Measure latency carefully. Latency profiling samples 100 questions per benchmark on an NVIDIA H100 GPU with an offline SGLang inference engine. They report mean and 95th-percentile per-turn reasoning latency, per-question model latency (prefill, reasoning, action generation, and other model inference over the full trajectory), and effective reasoning throughput. Environment-side processing — tool execution and LLM-as-judge evaluation — is excluded because it introduces model-independent variation.
The theoretical argument is a critical-path comparison: freeform reasoning has depth O(L log(H+L)) for L generated tokens, while SSR has depth O(log(H+K_max) + log N), removing the factor of L through token-level parallelism.
Why This Matters
Impact on research. The paper challenges the default assumption that agent reasoning must occupy the full natural-language space. It shows that recurrent high-level reasoning in multimodal search can be compressed into a compact, reusable set of candidates without sacrificing task success, and it introduces a decoding scheme — teacher-forced parallel scoring of complete candidate texts — that sidesteps autoregressive reasoning cost rather than merely shortening it. This is orthogonal to prior efficiency work that shortens or compresses reasoning while keeping sequential token dependencies, and to work like HyperEyes that reduces the number of interaction rounds rather than the cost of the reasoning span within each round.
Real-world applications.
- On-device assistants and phones, where small 2B/4B models must respond within tight latency budgets and cannot afford long deliberation before every action.
- Visual shopping and product search, where an agent decides among looking at the image, searching by image, or looking up a fact — matching the six library candidates closely.
- High-resolution and detail-critical inspection tasks, where the agent must decide when to zoom into part of an image before answering.
- Accessibility and information-lookup tools that combine camera input with knowledge retrieval under real-time constraints.
Industry relevance. The work is affiliated with Apple and Carnegie Mellon University, and the framing — small, on-device models, energy-efficient inference, predictable latency — maps directly onto commercial constraints. The reported p95 characteristic is arguably as important as the mean: SSR's p95 reasoning time sits only 10 ms above its mean, which makes latency budgeting for a shipped product far easier than with freeform generation whose length varies per turn. The authors also note that reasoning candidates are "manually written" and shared across benchmarks, which means the approach can be adopted without training a new reasoning generator.
Future Directions
- Improving the selector-gradient approximation. The paper attributes the modest GSPO and SAPO regressions to this approximation becoming less accurate across multiple gradient updates per rollout, suggesting a more accurate treatment of selection gradients is an open problem.
- Learning or expanding the reasoning library. The library is limited to six hand-written candidates and success increases monotonically as the library grows from one to six entries; how far that trend continues, and whether candidates can be discovered automatically rather than written by hand, is unresolved.
- Extending beyond multimodal search. The method is evaluated only on search-agent benchmarks with reverse-image search, text search, and image cropping tools; whether selection-based reasoning transfers to other agentic domains such as code execution or embodied control is not tested.
- Combining with round-reduction approaches. Because SSR reduces the cost of reasoning within a turn while methods such as HyperEyes reduce the number of interaction rounds, combining the two is a natural but untested direction.
- Testing larger backbones and broader benchmarks. The results cover 2B and 4B models; whether the latency advantage persists, shrinks, or becomes less necessary at larger scales is not reported. The paper also does not report dataset sizes for training or evaluation in the supplied text, as those details live in the truncated Appendix D.
Target Audience
This paper is most useful to applied researchers and engineers building agentic systems with small vision-language models, particularly those working under latency, memory, or on-device constraints. It will also interest RL researchers studying how to adapt policy-gradient objectives (GRPO, GSPO, SAPO) when a portion of the policy is a discrete categorical choice over pre-specified text rather than a token sequence. Product teams working on multimodal search, visual assistants, or camera-based lookup will find the latency profile and the concrete six-candidate library directly practical. Readers looking for a dense theoretical treatment or for large-scale frontier-model results should look elsewhere; the contribution here is a representation change and its efficiency profile, demonstrated at the 2B and 4B scale.
Authors’ abstract
Multimodal agents commonly generate free-form reasoning before each action. For small models, limited model capacity can result in lengthy reasoning that provides little useful guidance for action generation while incurring substantial inference cost. To address this challenge, we introduce Selection-based Structured Reasoning (SSR), a framework that reformulates reasoning as selection instead of open-ended generation. SSR represents recurring high-level reasoning as pre-specified, reusable natural-language candidates. At each turn, the model selects from these reasoning candidates based on their likelihoods given the current context, without requiring an auxiliary task head. Using pre-specified reasoning traces enables parallel scoring, where teacher-forced prefilling computes token likelihoods concurrently within and across candidates using a shared context KV cache. We evaluate SSR on seven multimodal search benchmarks using 2B and 4B models. Across multiple reinforcement learning objectives and supervised fine-tuning, SSR delivers significant efficiency gains without sacrificing task performance. SSR achieves an average success rate competitive with leading search agents of the same scale, while reducing per-turn reasoning latency by over 90% and total per-question model inference latency by 28-54%. Project page: https://zfy0314.github.io/ssr-webpage/.