While general-purpose LLM agents exhibit strong long-horizon reasoning, their production utility remains heavily fragmented across seven heterogeneous artifact modalities: text, images, audio, video, documents, 3D assets, and code. Training unified native omni-modal foundation models incurs prohibitive compute, while naive tool chaining fails to coordinate asset dependencies and multi-turn revisions. Researchers from the National University of Singapore (NUS) introduce Omni-IO Skills, a plug-and-play agent harness transforming existing text/vision LLMs into omni-native agents. Using hierarchical skill abstractions, standardized I/O interfaces, a persistent Asset Registry, and Declare Execution Graphs, Omni-IO coordinates complex multi-asset workflows concurrently. Evaluated on the UniM-90 multimodal benchmark, Omni-IO boosts input-support rates for frontier models from ~40% to 100%, and nearly triples semantic quality coupling scores without modifying foundational model weights.

Key Takeaways

  • ✓Plug-and-Play 7-Modality Runtime: Transforms existing LLM backbones into omni-native agents across text, image, audio, video, document, 3D, and code without modifying foundational model weights.
  • ✓Declare Execution Graphs & Persistent Asset Registry: Coordinates concurrent multi-asset pipelines while persisting intermediate multimodal artifacts for robust cross-turn reuse.
  • ✓100% Modality Support on UniM-90: Propels input-modality support for GPT-5.6 Sol and Claude Sonnet 5 from ~40% to 100%, nearly tripling coupled semantic-quality scores.
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

Core Background & Industry Pain Points While frontier LLM agents excel in symbolic reasoning, real-world digital pipelines span seven heterogeneous artifact modalities: text, images, audio, video, documents, 3D assets, and code. Navigating multi-asset workflows faces two polarizing engineering dilemmas: 1. Monolithic Omni-Modal Models Incur Prohibitive Costs: Training unified foundation models to natively ingest and generate high-dimensional audio, video, 3D meshes, and code simultaneously requires colossal compute and suffers from catastrophic forgetting across modalities; 2. Loose Tool Calling Lacks Structured Orchestration: Chaining isolated specialist models via basic function calling lacks dependency-aware tracking. Intermediate assets are lost between execution turns, and disparate I/O protocols break pipeline reproducibility. ### Architectural Highlights & Underlying Mechanics Researchers from the National University of Singapore (NUS NExT++ Lab) introduce Omni-IO Skills, a plug-and-play agent harness transforming existing LLMs into omni-native powerhouses: 1. Standardized Multimodal Protocol Across 7 Modalities: Defines uniform I/O interfaces across text, image, audio, video, document, 3D, and code, featuring 27 hierarchical skills covering understanding, generation, reasoning, and retrieval; 2. Declare Execution Graphs (DEGs): Represents multi-asset workflows as declarative directed acyclic graphs. The execution engine automatically parallelizes independent modality operations, minimizing end-to-end latency; 3. Persistent Asset Registry: Assigns unique identifiers to all intermediate multimodal artifacts, enabling deterministic cross-turn revision, caching, and interchangeable execution backends without loss of state. ### Benchmark & Experimental Validation Evaluated across 38 representative multimodal tasks on the newly designed UniM-90 benchmark: - 100% Modality Support Across Frontier Models: Raises the raw input-modality support rates of GPT-5.6 Sol and Claude Sonnet 5 from 40.00% and 38.89% up to a flawless 100%; - Near-3x Leap in Coupled Quality Scores: On the rigorous Semantic-Quality Coupled Score measuring alignment and fidelity, GPT-5.6 Sol leaps from 26.99 to 74.94, while Claude Sonnet 5 advances from 27.82 to 77.78; - Flawless Structural Compliance: Achieves 100.00 and 99.78 Strict Structure Scores, verifying that Declare Execution Graphs prevent format hallucinations across chained multimodal steps. ### Engineering Takeaways & Practical Guide - Paper & System Specs: Accessible at arXiv:2609.31847; - System Architecture Blueprint: For teams building multi-asset generation engines (e.g., automated game asset creation or cross-media publishing), decouple cognitive planning from modality rendering. Treat foundation LLMs as graph orchestrators and route specialized renderers through a persistent asset registry; - Zero-Finetuning Adaptability: Can be deployed immediately as harness-level middleware without modifying core LLM checkpoints.