Google and DeepMind added agentic video understanding to the latest Gemini models. Instead of statically sampling one frame per second, Gemini reasons over transcript, audio, and frames and dynamically adjusts frame rate, cutting tokens by up to 88% on long-form video. The feature is live via the Gemini API for 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.
Key Takeaways
- ✓Replaces static 1-fps scanning with dynamic frame-rate control over transcript, audio, and visual frames
- ✓Biggest gains on long-form video, using up to 88% fewer tokens with higher accuracy
- ✓Rolling out now in Gemini API / AI Studio; coming to the Gemini app and YouTube’s Ask YouTube