Streaming video LLMs must preserve perceptual evidence before downstream intent is revealed and respond proactively when evidence matures. Nanjing University researchers introduce OneStreamer, a unified architecture that couples query-independent factual recording with task execution through a shared proactive generation interface. Its Proactive Hierarchical Caption Memory (PHCM) synthesizes time-anchored micro-captions and event summaries, allowing inference to complement recent visual frames with cached factual records without reprocessing raw historical visual features. Trained on the broad-coverage OneStreamer-1M dataset, the 4B model establishes top SOTA across all eight streaming video understanding benchmarks.

Key Takeaways

  • ✓Resolves latency-memory trade-offs in streaming video LLMs via Proactive Hierarchical Caption Memory (PHCM) and PSTL
  • ✓Establishes new SOTA records across all eight streaming video benchmarks with a compact 4B foundation model
  • ✓Releases the million-scale OneStreamer-1M benchmark dataset and single-GPU real-time camera ingestion pipelines
🧭

Finished reading? Explore benchmark rankings & pricing

Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

🔬

In-Depth Technical Analysis

核心背景与行业痛点

Streaming video LLMs are vital for smart glasses, autonomous driving, and embodied robotics. In real-world streaming, future queries are unknown when evidence appears. Retaining extensive visual KV-caches across long temporal horizons rapidly exhausts VRAM and degrades real-time perception frame rates, while naive frame dropping irreversibly erases fine-grained facts.

架构亮点与底层机制

OneStreamer from Nanjing University unifies observation, contextual memory, and proactive reasoning through a shared generation interface:

  1. Proactive Hierarchical Caption Memory (PHCM): Dynamically produces micro-level descriptive captions and event-level summaries, transforming raw pixel streams into compact, reusable factual traces.
  2. Visual Window & Textual Memory Synergy: Combines a tight real-time visual frame window with cached text memories, completely bypassing the need to revisit high-dimensional historical visual tokens.
  3. Proactive State Transition Learning (PSTL): Filters repetitive waiting tokens, supervising only 27.5% of critical transition anchors to achieve superior state tracking without gradient dilution.
  4. OneStreamer-1M Dataset: Contributes over 1 million synthetic and cleaned streaming video records across diverse interactive scenarios.

权威 Benchmark 与实测跑分对比

Evaluated across eight streaming video understanding benchmarks:

  1. Consistent SOTA Across 8 Benchmarks: The compact 4B parameter model establishes superior accuracy over larger streaming baselines across all 8 benchmarks.
  2. Preserved Real-Time FPS: Retaining textual memories substantially improves historical retrieval QA without compromising millisecond-level perception throughput on incoming frames.
  3. Extreme Data Efficiency: PSTL outperforms dense supervision while requiring only 27.5% token coverage.

开发者实战落地与开箱指南

OneStreamer code, checkpoints, and datasets are open-sourced on GitHub. Engineers building edge devices (smart glasses, robotics, inspection drones) can ingest direct RTSP/camera feeds to achieve seamless, low-latency streaming comprehension on single-GPU hardware profiles.