Google DeepMind launched agentic video understanding on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. Models dynamically search frames, audio, and transcripts instead of ingesting video at a fixed FPS. Official claims include up to 88% fewer tokens, about 66% lower analysis cost, and up to about 7% better accuracy. It is live via the Gemini API, Google AI Studio, and the Gemini Enterprise Agent Platform at standard token pricing with no extra feature fee.
Key Takeaways
- ✓Available on Gemini 3.7 Flash / 3.6 Flash / 3.5 Flash-Lite; enable with processing=agentic.
- ✓Long-form video: up to 88% fewer tokens, ~66% lower cost, up to ~7% accuracy gain.
- ✓Supports uploads and YouTube; unlocks sub-second retrieval, anomaly detection, and precise counting.