Autonomous deep research agents are expanding beyond text-centric search toward multimodal discovery across high-resolution imagery, complex charts, and long-horizon video. However, existing multimodal models struggle to ground fine-grained visual anchors and bridge perceptual evidence with external search verification. Researchers from Fudan University and collaborating institutions introduce OneSearch-VL (arXiv:2610.12419), a unified multimodal deep research agent for single-image, multi-image, and video environments. OneSearch-VL centers on the Visually Grounded Evidence Graph (VGEG), which explicitly encodes localized visual anchors, entity relationships, source-supported facts, and research operations into a shared dependency topology. Leveraging VGEG, the authors construct OneSearch-VL-SFT-110K and OneSearch-VL-RL-10K datasets alongside an Evidence-aware Visual-Grounded Rubric reward (EVGR) to supervise visual traceability during reinforcement learning. On the newly introduced OneSearch-MI-Bench and OneSearch-Video-Bench, OneSearch-VL-8B outperforms tool-augmented Qwen3-VL-8B by 20.2 and 17.6 percentage points respectively, while setting new state-of-the-art results across seven established image benchmarks and VideoDR, with code fully open-sourced.

Key Takeaways

  • ✓Fudan University releases OneSearch-VL, unifying image and video deep research via Visually Grounded Evidence Graphs (VGEG)
  • ✓Outperforms tool-augmented Qwen3-VL-8B by 20.2 and 17.6 points on multi-image and video benchmarks with 91.8% citation precision
  • ✓Open-sources SFT-110K, RL-10K datasets, and EVGR process rewards, setting an auditable standard for multimodal agents
🧭

Turn your technical choice into a development budget

Compare 40 dev plans & simulate token costs vs $20/mo subscriptions

🔬

In-Depth Technical Analysis

Background and the Problem

With Deep Research emerging as an industry priority, autonomous retrieval and report generation are moving beyond text toward multimodal domains such as pathology scans, satellite imagery, and long-form video archives. However, existing Vision-Language Models (VLMs) encounter three fundamental barriers: visual perception remains disconnected from external search queries, temporal and spatial cross-image dependencies drift over long horizons, and the community lacks process-supervised datasets linking visual grounding to external evidential verification.

Architecture and How It Works

To unite visual anchoring with autonomous research, researchers from Fudan University present OneSearch-VL (arXiv:2610.12419, code: appletea233/OneSearch-VL):

  1. Visually Grounded Evidence Graph (VGEG): Explicitly structures deep research workflows into directed graphs. Nodes differentiate localized visual bounding boxes, retrieved entity records, verified source facts, and synthesis steps, preserving end-to-end task dependencies.
  2. Scalable VGEG-Driven Data Engine: Constructs and filters expert multi-image and video trajectories, assembling OneSearch-VL-SFT-110K for supervised alignment and OneSearch-VL-RL-10K for post-training exploration.
  3. Evidence-Aware Visual-Grounded Rubric Reward (EVGR): Replaces scalar end-task rewards with a step-level rubric rewarding accurate visual bounding, authentic source citation, and faithful graph derivation, effectively suppressing visual hallucinations.

Benchmarks and Measured Results

Evaluated on demanding multimodal research benchmarks:

  1. 20.2 and 17.6 Percentage Point Leads: On the newly introduced OneSearch-MI-Bench and OneSearch-Video-Bench, OneSearch-VL-8B outscores tool-equipped Qwen3-VL-8B by 20.2 and 17.6 percentage points respectively.
  2. New SOTA Across Seven Image Benchmarks and VideoDR: Demonstrates state-of-the-art visual localization and deep reasoning across seven visual QA benchmarks and fine-grained video retrieval suites.
  3. 43% Reduction in Evidential Drift: Blind evaluations reveal that 91.8% of external citations generated by OneSearch-VL strictly correlate with physical visual anchors, cutting evidential hallucinations by 43% over baselines.

Getting Started for Developers

The authors have released the codebase, dataset recipes, and checkpoints on GitHub (appletea233/OneSearch-VL). When designing production multimodal research agents or visual coding assistants, developers should avoid piping raw images straight into unconstrained search tools. Adopt the VGEG protocol: enforce structured intermediate steps where the agent outputs explicit bounding coordinates and visual hypotheses before issuing search API requests, ensuring verifiable evidential rigor across multimodal research workflows.

Action HubReady to adopt this in production?

Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.