Gemini's agentic video understanding picks frames and cuts tokens by up to 88 per cent
Google DeepMind released agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. The model chooses what to watch, at what speed, and whether to use frames, audio or the transcript. Google says long video uses up to 88 per cent fewer tokens, costs up to 66 per cent less and is up to 7 per cent more accurate. The Gemini app and YouTube's Ask YouTube follow.
Artificial Intelligence··Midday
Google opens agentic video understanding on three Flash models
Google DeepMind released agentic video understanding for Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Android Authority reports that Google is bringing the capability to those three Flash models and that users can upload videos for analysis. Google says standard token pricing applies with no separate feature fee. The feature works today on uploaded videos and YouTube videos through the API.[1], [2]
The model picks which frames to watch; Google says tokens fall by up to 88 per cent
Google says that instead of sampling a video at a fixed frame rate, the model decides what to watch, at what speed and whether to read frames, audio or the transcript. The company says that on long-form video this uses up to 88 per cent fewer tokens than static processing, lowers cost by up to 66 per cent and raises accuracy by up to 7 per cent. Android Authority repeats the 88 per cent token figure and the 7 per cent accuracy figure. Android Authority writes that Gemini can pinpoint split-second changes, answer complex questions across multi-hour videos, inspect videos for visual artifacts, and count and track physical movements and objects.[1], [2]
The feature is live on the API; the Gemini app and Ask YouTube are next
Google lists sub-second moment retrieval, long-form needle-in-a-haystack search, anomaly detection and counting of actions and objects as target tasks. Android Authority writes that for now the feature is available only through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Google says it is coming to the Gemini app soon and to YouTube's Ask YouTube feature in the coming months. Android Authority also says the Gemini app will follow soon and that Ask YouTube will run on this capability.[1], [2]