Google DeepMind is rolling agentic video understanding into the latest Gemini models. Instead of scanning an entire file, Gemini reasons over transcript, audio, and frames and dynamically adjusts frame rate — using up to 88% fewer tokens with better accuracy, especially on long-form video. The capability is live via API on 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
Key Takeaways
- ✓The model dynamically samples transcript, audio, and frames instead of scanning whole files, with the biggest gains on long-form video
- ✓Accuracy improves while using up to 88% fewer tokens, cutting the cost of multimodal agent video workflows
- ✓Rolling out via API in Google AI Studio on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite, then the Gemini app