Google DeepMind is rolling out agentic video understanding to its latest Gemini models, which reason over transcript, audio, and frames and dynamically adjust frame rate instead of scanning entire files. Accuracy improves while using up to 88% fewer tokens, with the largest gains on long-form content. The feature is live via API for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite in Google AI Studio.
Key Takeaways
- βAgentic frame selection replaces full-file scans, cutting tokens by up to 88%
- βGemini reasons jointly over transcript, audio, and frames with dynamic frame-rate control
- βRolling out now to 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite via Google AI Studio API