Research
Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
Overview Research area: Omni-modal agent systems — multi-step evidence-seeking question answering over video, audio, images, web pages, and computation. Technical level: Advanced Scope: The paper diag

- arXiv
- 2607.11433
- Published
- 2026-07-13
- Authors
- Ming Ma, Yi Zhu, Yiran Zhong, Feida Zhu, Yuhao Wang, Junhan Shi, Lingrui Mei, Tianming Yang, Steven Hoi
AI summary
Overview
Research area: Omni-modal agent systems — multi-step evidence-seeking question answering over video, audio, images, web pages, and computation.
Technical level: Advanced
Scope: The paper diagnoses planning (not perception) as the dominant bottleneck of omni-modal agents and proposes evidence-ledger planning, a system in which noisy multimodal observations are digested by a critic into a typed, task-scoped ledger that the planner reads instead of a growing dialogue history.
What This Paper Is About
Omni-modal agents answering questions grounded in heterogeneous media must keep track of which evidence they still need while noisy observations pile up in the conversation history and degrade later decisions. The paper argues that current multimodal models fall short mainly on the planning side, not the perception side, and that this is caused by unmanaged context rather than by weak perception. The goal is a system whose persistent context stays compact and verified throughout a task, so the planner always decides from a clean state.
Key Contributions
- Diagnosis. Through controlled backend replacements, the paper identifies planning as the dominant bottleneck of omni-modal agents, showing that replacing the planner causes a much larger performance loss than replacing the perception backend.
- System. Evidence-ledger planning keeps observation noise out of persistent context: each task maintains a typed ledger, the critic digests every observation into ledger updates, the ledger accepts updates only through the critic's verdicts, and the planner's persistent context contains only the ledger, with each raw observation seen once at the step under review.
- Training. A recipe requiring no step-level manual annotation: trajectories and ledger states recorded during execution are harvested as supervised fine-tuning and decision-level reinforcement learning signal, improving two planner backbones.
- Results. State-of-the-art 81.4% accuracy on open-world tool interaction in OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on long-video understanding in WorldSense.
Main Findings
- OmniGAIA accuracy: Omni-Decision reaches 81.39% overall, the best result on the benchmark, versus 79.44% for Gemini-3.1-Pro measured in the same evaluation window with the same judge — a 1.95-point gap, concentrated on Easy and Hard.
- Comparison to agent systems: The strongest publicly reported agent system, the sandboxed coding agent, reaches 75.00%; Omni-Decision is 6.4 points higher overall and 7.7 points higher on Hard. The other three agent systems run by the authors (Minimal agent, OmniGAIA base, OmniAgent) all stay below 26%.
- Harness effect: With the same pair of backends, overall accuracy ranges from 5.56% to 81.39% across harnesses.
- Planner versus perception: With Gemini-3.1-Pro perception fixed, replacing the planner with GLM-5.2, Qwen3.5-Plus, Qwen3.5-27B, and Qwen3-Omni-30B lowers accuracy from 81.39% to 75.00%, 58.06%, 46.94%, and 14.72% respectively. With the GPT-5.2 planner fixed, all three perception replacements stay at or above 58.89%.
- Plus-tier ratio: Within the Qwen3.5 family Plus tier, replacing the planner costs 23.3 points and replacing perception costs 15.8 points, a ratio of approximately 1.5.
- Saturation: Replacing both backends gives 13.89%, nearly the same as the 14.72% from replacing only the planner — once the planner is weak enough, perception quality no longer matters.
- State ablation: Evidence Ledger 81.39% overall versus Memory 68.33% and No-ledger ReAct 60.28%. The gain over No-ledger ReAct is approximately 24 points on Medium and Hard, compared with approximately 16 points on Easy.
- Training gains: Qwen3.5-27B rises from 46.94% to 49.44% with state-SFT and 50.56% after closure-aligned RL; Qwen3-Omni-30B rises from 14.72% to 18.61% with state-SFT.
- Failure attribution: Across the 67 incorrect cases, media perception accounts for 27 cases (40.3%), external-fact retrieval 22 (32.8%), verification/conflict resolution 9 (13.4%), computation 6 (9.0%), and premature stopping 3 (4.5%). Media perception and external retrieval together account for approximately three quarters of failures.
- Evidence progress: Correct cases close 66.2% of their declared needs on average, compared with 32.8% for incorrect cases; 92.5% of incorrect cases close at least one need.
- WorldSense: Omni-Decision obtains 65.0%, level with Gemini-3.1-Pro at 65.5%, and higher than the other models and agent systems in the table (OmniAgent reaches 28.06%). Cost averages approximately $0.3 per question versus approximately $0.8 for direct answering by the same model.
- OmniGAIA cost: Omni-Decision averages approximately $1.2 per question and Gemini-3.1-Pro approximately $2.8 in the same window.
Methodology in Plain English
Instead of letting tool observations accumulate in a chat history, each task keeps an evidence ledger with four fields: open evidence needs (U), fact and computation slots (F), confirmed evidence atoms with source and timing (E), and unresolved conflicts (C). The query is parsed into a checklist of needs, each marked blocking or supplementary, plus the fact and computation slots those needs must fill.
At each step the planner reads the ledger as bounded text and selects an action from a fixed action set defined by the available assets, tools, and answer format. It can invoke media grounding, retrieval, browsing, computation, or visual verification, or it can select finish. When a tool returns an observation, a critic — a separate judgment from the planner — reads that observation against the current needs and issues a verdict: supports, conflicts with, or omits. Only the critic's verdict reaches the reducer, which is the ledger's sole writer; it applies deterministic field-by-field updates that close needs, fill verified slots, append sourced evidence atoms, and record contradictions while retaining both conflicting sources. Randomness stops at the commit boundary, and no module can bypass the reducer.
Readiness requires three conditions simultaneously: no blocking need remains open, every fact and computation slot holds a verified value, and no conflict remains unresolved. On finish, a finalizer drafts an answer and checks it against the ledger; a failed check writes the diagnosis back as an event for the next step. The system stops and reports insufficient evidence when no remaining action can advance the ledger, and returns a forced best value if repeated finish attempts exhaust the revision budget.
Training needs no step-level annotation. Every planner decision is recorded with the ledger state it saw and the critic verdict that followed. State-SFT retains runs the judge marks correct, expands each step into a ledger-state/planner-action pair, and fine-tunes the planner. Closure-aligned reinforcement learning then targets the dominant weak-planner failure — abstaining or converging while needs remain open — by sampling a group of candidate decisions per reconstructed ledger state and reinforcing those scored above the group mean by rule-based closure checks and a lightweight reward judge; because the ledger lists each open need explicitly, a decision can be scored without executing it.
Configuration: planner, critic, and finalizer are one LLM reading only the ledger and current observation as text (GPT-5.2, version gpt-5.2-2025-12-11 in the main experiments), perception is Gemini-3.1-Pro (gemini-3.1-pro-preview) exposed as a tool, and runs use at most 15 steps with at most one tool call or one finalization attempt per planner call. Tools include subtitle grounding, audio scouting, clip grounding, frame confirmation, web search, and a code executor; web search blocks Hugging Face domains to prevent retrieval of benchmark data.
Why This Matters
Impact on research: The paper turns the system layer into an explicit, ablatable object, enabling attribution of bottlenecks rather than reporting only end-to-end scores. It provides direct evidence that planning is the shortfall to close first and that perception becomes the bottleneck once planning is closed, and it shows that execution traces from a ledger can serve as process supervision without human step-level annotation.
Real-world applications:
- Fact-checking and investigative Q&A over mixed video, audio, and web sources, where conflicts between sources must be recorded rather than silently resolved.
- Long-video and media-archive understanding, including the conventional multiple-choice setting tested on WorldSense where answers lie in video, audio, and subtitles.
- Enterprise assistants that must complete missing attributes from the web and compute derived values — dates, differences, arithmetic — before answering.
- Cost-sensitive deployments, since the ledger approach pays roughly 43% of Gemini-3.1-Pro's cost per question on OmniGAIA and roughly $0.3 versus roughly $0.8 on WorldSense.
Industry relevance: The finding that harness choice moves accuracy from 5.56% to 81.39% with identical backends implies that agent-framework engineering can matter more than model upgrades for this class of task.
Future Directions
- Building a critic that reads multimodal input directly: the current critic reads observations only in text form and media never enter its input, so ledger quality is bounded by the text output of the perception tools.
- Closing the remaining acquisition bottleneck, since with the ledger and strongest planner in place, failures concentrate in media perception and external retrieval.
- Extending training gains to weaker backbones, as improvements currently vary with the planner backbone's initial capability.
- Determining whether the ledger's task-scoped scope can be extended, given the paper states its scope does not extend to open-ended organization of long-term memory.
Target Audience
Researchers and engineers working on multimodal agents, tool-using LLM systems, and agent context management, as well as practitioners who need to attribute failures between a model's planning and its perception. Readers should be comfortable with agent loops, reinforcement learning terminology, and benchmark evaluation.
Authors’ abstract
Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.