Understanding long-form video requires multimodal agents to gather and synthesize temporal evidence across dozens of reasoning steps. However, existing architectures predominantly employ append-only working memory, resulting in 'semantic thrashing': as contexts expand, attention across critical facts collapses and the agent loses access to previously discovered clues. Researchers from the University of Rochester and Microsoft Research introduce VideoLoop, an agent architecture featuring two coupled loops: an outer reasoning loop coupled with an inner memory rewriting loop. After each step, the inner loop queries an unbounded filesystem of past observations and continuously rewrites a bounded working memory buffer. VideoLoop provides plug-and-play gains averaging 4.2 points across four major LVLM backbones, propelling Gemini 3.1 Pro to 88.3% on VideoMME (long) and 88.8% on VideoMMMU.
Key Takeaways
- ✓Formalizing Semantic Thrashing: Mathematically establishes that append-only memory structures inevitably fail over long horizons without an explicit rewrite operator to prune accumulated observation noise.
- ✓Coupled Dual-Loop Memory Architecture: Combines an outer multimodal reasoning loop with an inner memory rewriting loop that selectively retrieves from an unbounded observation store into a compact, bounded working context.
- ✓88.3% on VideoMME (Long) with Gemini 3.1 Pro: On the hardest quartile of VideoMME (long), blind evaluation accuracy leaps from 60.9% to 81.1%; paired with Gemini 3.1 Pro, achieves 88.3% on VideoMME and 88.8% on VideoMMMU.
Finished reading? Explore benchmark rankings & pricing
Real-world SWE-bench scores & $20/mo vs API cost break-even calculator

Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.