← News

DeepMind ships agentic video understanding for Gemini Flash

1 Sep 2026: DeepMind published that agentic video understanding is live on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite — a native-tool loop instead of fixed-FPS ingest. Company blog is the primary. Token, cost, and accuracy figures are DeepMind’s benches.

1 Sep 2026: DeepMind published “Introducing agentic video understanding with Gemini.” The lab says the feature launched on Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. That dated post is the filing event.

DeepMind contrasts it with static processing, where the model ingests video at a fixed frames-per-second rate (default 1 FPS). Agentic mode, it says, pairs Gemini’s reasoning with native video tools in a loop that searches, scans, and inspects target segments across frames, audio, and transcripts — fetching only what the query needs.

DeepMind’s figures on the same page: up to 88% fewer tokens, up to 66% lower cost, and up to 7% better accuracy versus static ingest on “standard video analysis benchmarks.” Those are DeepMind benches. This desk did not rerun them.

Availability on the post: video uploads and YouTube URLs via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform. Enable it by setting processing to “agentic” in the API config.

Pricing, per DeepMind: standard Gemini API token rates. The lab says there is no extra feature fee.

Roadmap on the same page, not today’s ship: DeepMind says the feature will roll out in the Gemini app on Flash and Flash-Lite “soon,” and that in the coming months it will power YouTube’s Ask YouTube on the watch page. Treat those as company plans.

A dated DeepMind primary puts an agentic video loop on three Flash models — search the tape instead of paying for every frame. File the ship and the enable flag. Leave the token, cost, and accuracy percentages as DeepMind’s benches.

ONLINE

article thread

guidelines

warming…

warming…

Sources