Research
Omni-IO Skills: Harnessing Your Agent Omni-Native
Overview Research area: Natural Language Processing / multimodal agent systems, specifically the harness layer that sits between a general-purpose agent's reasoning core and heterogeneous multimodal e

- arXiv
- 2609.31847
- Published
- 2026-09-25
- Authors
- Yanlin Li, Mingyang Hao, Shengqiong Wu, Hao Fei, Mong-Li Lee, Wynne Hsu
AI summary
Overview
Research area: Natural Language Processing / multimodal agent systems, specifically the harness layer that sits between a general-purpose agent's reasoning core and heterogeneous multimodal execution backends.
Technical level: Advanced. The paper assumes familiarity with multimodal foundation models, agent tool-use loops, and evaluation metrics for structured generation.
One-sentence scope: The paper introduces Omni-IO Skills, a plug-and-play harness of 27 hierarchical Skills plus dependency-aware orchestration and a persistent asset registry, and evaluates whether it can make two existing general-purpose agents omni-native on the 90-instance UniM-90 benchmark.
What This Paper Is About
General-purpose agents can plan and act over long horizons, but their ability to produce things is fragmented across text, images, audio, video, documents, 3D assets, and code. Adding a new modality to a foundation model means costly model retraining, while stitching together specialist models and tools leaves open how procedures, dependencies, intermediate assets, and cross-turn revisions get coordinated.
The paper's goal is an Omni agent harness: a layer between the agent's reasoning and its execution backends that turns scattered models, tools, and procedures into selectable, composable, executable, and traceable capabilities, without modifying the host agent's reasoning core.
Key Contributions
-
Formulation of an Omni-modal Agent Harness. The authors define a system layer that extends a general-purpose agent into an omni-native one without retraining the host model, leaving the host's reasoning and planning core unchanged.
-
The Omni-IO Skills architecture. A four-layer design (Skill Entry, MCP Tool Service, Provider and Configuration, Asset Registry) integrating multi-granularity procedural knowledge, heterogeneous execution backends, dependency-aware orchestration, and persistent artifact state.
-
An application-facing system with broad coverage. 27 Skills — 19 Atomic, 2 Expert, and 6 Scenario — covering 38 representative tasks across understanding, generation, reasoning, and retrieval, spanning seven artifact modalities and domains from education and research to marketing, creative production, and software engineering.
-
Consistent gains across two different host agents. Demonstrated improvements in modality coverage, semantic quality, interleaved coherence, and structural completeness for both GPT-5.6 Sol and Claude Sonnet 5.
Main Findings
-
Modality coverage closes completely. On UniM-90, Omni-IO Skills raises the input-support rate of GPT-5.6 Sol from 40.00% to 100% and of Claude Sonnet 5 from 38.89% to 100%.
-
Relative semantic quality improves sharply. Relative Semantic–Quality Coupled Score (SQCS) rises from 26.99 to 74.94 for GPT-5.6 Sol and from 27.82 to 77.78 for Claude Sonnet 5 — gains of 47.95 and 49.96 percentage points.
-
Absolute semantic quality also improves, not just coverage. Absolute SQCS increases from 67.49 to 74.94 and from 71.53 to 77.78. The authors note that the post-harness absolute values across all 90 instances exceed the base agents' 67.49 and 71.53 measured only on their narrower supported subsets.
-
Interleaved coherence jumps in the relative setting. Relative Interleaved Coherence Score increases from 34.61 to 93.98 for GPT-5.6 Sol and from 32.00 to 83.28 for Claude Sonnet 5. Absolute ICS rises from 86.53 to 93.98 and from 82.29 to 83.28.
-
Structure scores are near-perfect after loading the harness. Strict Structure Score reaches 100.00 for GPT-5.6 Sol and 99.78 for Claude Sonnet 5; both hosts attain a Lenient Structure Score of 100.00. Absolute StS moves from 47.84 and 52.21, and absolute LeS from 72.22 and 68.57.
-
Failures stay in the denominator. For instances whose inputs can be processed but whose required output modalities cannot be generated, or whose target outputs cannot be completed, the failures remain included in metric computation.
-
Cross-agent comparisons are deliberately not interpreted. Because the default capabilities of the two agents may differ, the authors focus on within-agent gains and do not treat cross-agent score differences as a ranking of model capability.
-
Qualitative cases confirm end-to-end behavior. In an art tutorial case, the system invokes the S4 Education Sharing Skill to plan and coordinate, A2 Video Understanding to extract drawing steps, A3 Audio Understanding to recover explanations and procedural order, and A6 Image Generation to produce an 11-image tutorial. In a product promotion case, it invokes A1 Image Understanding, E1 Poster Design, E2 Complex Video Production, and A16 Code Generation. Both hosts produce all requested outputs while preserving consistency.
Methodology in Plain English
The authors build a harness rather than a model. The harness sits between the host agent and a set of external multimodal services, and it carries capabilities as loadable Skills at three levels of granularity:
- Atomic Skills perform a single operation, such as image understanding, video generation, PPT generation, or web search.
- Expert Skills target one concrete deliverable and expand into several Atomic Skills. For example, Poster Design expands to A1, A6, and A16; Complex Video Production expands to A2, A3, A6, A7, A8, A9, and A10.
- Scenario Skills address an application context with multiple related deliverables. The six implemented ones are Social-Media Post, Office Documents, Job Application, Education Sharing, Event Material, and Game Asset.
Every Skill follows a shared declarative form describing its applicability conditions, inputs, procedure, outputs, and relationships to other Skills. When a request arrives, the harness selects the relevant Skill and recursively expands it until the remaining steps are executable. Those steps become nodes in a Declare Execution Graph (DEG), where edges encode both control dependency (a node cannot start before its predecessor finishes) and data dependency (a predecessor's output is consumed downstream).
Before anything runs, the graph is validated: every dependency must resolve to a node or a registered asset, and the graph must be acyclic. An invalid graph is treated as a planning error and rebuilt rather than submitted. Valid graphs are scheduled in Waves — nodes whose predecessors have all completed form the next Wave, and nodes within a Wave are mutually independent and dispatched concurrently. Scheduling follows declared dependencies rather than modality, so image and audio generation can run in parallel when neither consumes the other, while an image-conditioned video node waits for its reference image. If a node fails, its pending descendants are cancelled but independent branches continue, and completed outputs are not rolled back.
Below the Skills, three further layers handle execution. The MCP Tool Service maps each executable node to a standardized tool interface grouped into understanding, generation, and utility tools. The Provider and Configuration layer binds each capability to a concrete provider, model, credentials, default parameters, and fallback policy, so a provider can be swapped without rewriting any Skill. The Asset Registry normalizes every user-provided, intermediate, and final artifact into a record with a globally unique asset_id, type, subtype, path, description, params, turn_id, and source_asset_id. Records are persisted in an append-only JSON registry with atomic, file-locked writes.
The evaluation uses UniM-90, a fixed 90-instance subset of UniM covering text, image, audio, video, document, code, and 3D modalities plus interleaved combinations, selected independently of any evaluated agent's native modality capabilities. Each agent is tested in two configurations — its default environment with built-in Skills and tools, and the same environment plus Omni-IO Skills and its MCP tool services, with all other settings and execution budgets unchanged. The UniM Evaluation Suite reports input-support rate (τ), absolute and relative SQCS, ICS, Strict Structure Score, and Lenient Structure Score.
Why This Matters
Impact on research. The paper argues that Omni capability need not come only from scaling a shared backbone. By placing unification at the level of task execution and artifact flow, and treating foundation models, specialist models, and media engines as replaceable backends, it opens a complementary research direction where applications evolve independently of any particular model stack. It also connects two previously separate lines — Omni-agent coordination of modality experts, and Agent Skills as packaged procedural knowledge — at their underdeveloped intersection of multi-asset execution and persistent artifact state.
Real-world applications:
- Education content production. The paper's own example: a course built from lecture recordings and reference documents, ending in slides, illustrations, narration, and an explainer video that share facts, style, and timing.
- Marketing and product promotion. A scenario Skill expands a set of product images into a poster, a promotional video with sound effects, and a landing web page, with a four-Wave schedule and later style revisions that reuse the registered product analysis.
- Game asset pipelines. Scenario Skill S6 Game Asset coordinates image, video, 3D, and code generation from a character brief, producing concept art, a showcase video, and a 3D model.
- Office and job-application workflows. S2 Office Documents coordinates audio, document, and PPT/Word/PDF/Excel generation, while S3 Job Application coordinates poster and video production with document, speech, PPT, and Word generation.
Industry relevance. The harness is plug-and-play against existing agents — the code is released at https://github.com/any2any-mllm/Omni-IO-Skill — and the Provider layer allows models and providers to be substituted without touching procedural knowledge. Any vendor whose multimodal services keep improving outside a shared backbone can be plugged in as an execution backend, and the append-only asset registry with cross-turn reuse addresses a practical production requirement: revisiting and revising earlier deliverables without regenerating validated work.
Future Directions
-
Growing the Skill catalog and scope. The current implementation covers 27 Skills across seven modalities; the architecture is presented as extensible, leaving open how far the Atomic/Expert/Scenario hierarchy scales before selection and expansion become the bottleneck.
-
Broadening host-agent evaluation. Only two hosts, GPT-5.6 Sol and Claude Sonnet 5, are tested, and the authors explicitly avoid interpreting cross-agent differences as a ranking. Whether the gains hold for other agent architectures, and how the harness interacts with agents whose native modality support already differs, is not settled.
-
Failure semantics under partial execution. The paper specifies that node failures cancel pending descendants while independent branches continue, and that completed outputs are not rolled back. How to recover, retry, or reason about partially completed multi-deliverable requests is left as an open systems question.
-
Benchmark generality. UniM-90 is a controlled 90-instance subset selected independently of any agent's native capabilities. Whether the reported input-support and structure-score gains transfer to larger or differently distributed multimodal workflow benchmarks is not reported.
The paper does not include an explicit future-work section, so the above are open questions the work raises rather than stated plans.
Target Audience
Researchers and engineers working on multimodal agents, agent harnesses, and tool-orchestration systems who want a system-level alternative to retraining foundation models for additional modalities. It is most useful to readers already comfortable with agent action loops, multimodal generation services, and structured output evaluation. Practitioners building production pipelines that must coordinate dependent media artifacts across turns — in education, marketing, game development, or document workflows — will find the architecture and asset-lifecycle design directly applicable.
Authors’ abstract
General-purpose agents can plan, reason, and act over long horizons, yet their production capabilities remain fragmented across text, images, audio, video, documents, 3D assets, and code. Extending a foundation model to additional modalities ties capability growth to costly model updates, while assembling specialist models and tools leaves unresolved how procedures, dependencies, intermediate assets, and cross-turn revisions should be coordinated. We present Omni-IO Skills, a plug-and-play Agent Harness that makes existing agents omni-native through hierarchical Skills, a standardized multimodal execution interface, dependency-aware orchestration, and a persistent Asset Registry. Multi-asset workflows are represented as Declare Execution Graphs, which schedule independent operations concurrently and register successful outputs for downstream and cross-turn reuse across replaceable execution backends. Its 27 Skills cover 38 representative tasks spanning seven artifact modalities and four capability families: understanding, generation, reasoning, and retrieval. On UniM-90, the harness raises the input-support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% to 100%, while increasing relative Semantic--Quality Coupled Score from 26.99 to 74.94 and from 27.82 to 77.78, respectively; Strict Structure Score reaches 100.00 and 99.78. These results establish harness-level capability composition as a practical route to broad, evolvable Omni systems without changing the host agent's reasoning core.