Google announces agentic video understanding for Gemini API
Google says agentic video understanding is available through its Gemini API for three Flash models and can be enabled with an "agentic" processing setting. Google reports lower token use and costs and higher accuracy, and names LongVideoBench for one comparison, but does not provide the full benchmark methodology or conditions. It also says a Gemini app rollout will reach all users on Flash and Flash-Lite models soon, without a date or regional scope. [1]
This edition passed Imananq's enhanced publication checks. Some material claims remain explicitly attributed to official or company sources because no independent source is currently bound to this edition. The engine continues checking approved sources and will add corroboration only through a new edition that passes the full gate.

Google announced agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite on 1 September 2026. Google says the Gemini API option lets models selectively inspect video frames, audio, and transcripts rather than only processing video at a fixed rate. [1]
01
02
Models named in the launch
Google lists Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite.Google's stated maximum token reduction
This is a first-party maximum result. Google names LongVideoBench for one comparison but does not provide the full methodology or conditions.Google's stated maximum cost reduction
This is a first-party maximum result under conditions not fully detailed in the supplied announcement.Google's stated maximum accuracy gain
This is a first-party maximum result. LongVideoBench is named for one comparison, while full metrics, baselines, and test conditions are not supplied.All performance figures are Google's reported maximum results and are not independently validated in the supplied evidence. [1]
03
What Google announced
Google records the launch of agentic video understanding as an API capability for three Gemini Flash models. The announcement does not specify regional coverage, account-tier access, or quotas, so it does not establish access conditions beyond the named platforms and configuration. [1]
- Google says the feature is available through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
- The company names Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite as supported models.
- Google says it works with uploaded videos and YouTube videos and can be enabled by setting video processing to "agentic" in the API configuration.
04
How Google says it works
According to Google, the feature lets Gemini determine which parts of a video to inspect, at what speed, and through which modality. Google says it does this through an agentic loop that invokes an internal tool to load relevant video material. This is the provider's technical description; the supplied evidence includes no independent test of its behavior. [1]
- Google contrasts the feature with static processing at a fixed frames-per-second rate, which it says defaults to one frame per second and can be adjusted through the API.
- It says the model can dynamically search, scan, and inspect relevant segments across visual frames, audio, and transcripts.
- The announcement describes possible applications including brief-moment retrieval, long-video search, anomaly detection, and action or object counting.
05
Performance and cost claims need qualification
The stated improvements are Google's maximum reported results, not general guarantees. The announcement names LongVideoBench for one Gemini 3.7 Flash comparison, but does not disclose the complete benchmark set, full baselines, metrics, evaluation methodology, or conditions behind its overall maximum claims. It also does not state relevant token prices or say whether pricing changes by model, region, or processing conditions. [1]
- Google reports up to 88% lower token consumption.
- It reports up to 66% lower analysis costs.
- It reports up to 7% higher accuracy across what it calls standard video-analysis benchmarks.
- Google identifies LongVideoBench as a long-form video understanding benchmark used to compare Gemini 3.7 Flash with and without agentic video understanding.
- Google says standard Gemini API token pricing applies and that there is no additional feature fee.
06
What is not yet confirmed
These are future rollout statements from Google rather than completed availability confirmations. Google states an intended all-user scope for the Gemini app rollout, but supplies no specific date or regional scope. For Ask YouTube, it supplies neither a specific date nor regional scope. Developers considering the API should also treat duration limits, file-size limits, latency, quota rules, and data-handling terms as unresolved until documentation for their account and deployment addresses them. [1]
- Google says the feature will roll out soon to all users in the Gemini app across Flash and Flash-Lite models.
- It says agentic video understanding is expected to power YouTube's Ask YouTube feature on the video watch page in the coming months.
07
What developers can do
The announcement describes an API setting that developers can test where it is available to their account. Important operating details are not specified in the supplied material. [1]
- 01
Use Gemini 3.7 Flash, 3.6 Flash, or 3.5 Flash-Lite with uploaded videos or YouTube video inputs through the Gemini API in Google AI Studio or the Gemini Enterprise Agent Platform. [1]
- 02
Set video processing to "agentic" in the API configuration to enable the feature. [1]
- 03
Confirm applicable model token prices and account access before deployment. Google says standard Gemini API token pricing applies without a separate feature fee, but does not state token rates, quotas, regions, or account-tier terms. [1]
- 04
Test quality, latency, input limits, and costs on representative videos. Google names LongVideoBench in one comparison, but does not supply the full benchmark suite, methods, detailed conditions, or operating limits needed to forecast production results. [1]
08
Limits of this edition
The supplied evidence is a Google product announcement. It does not independently verify availability, technical behavior, benchmark outcomes, or pricing claims. [1]
Google identifies LongVideoBench as a long-form video benchmark used in a Gemini 3.7 Flash comparison, but does not provide the full benchmark suite, baselines, metrics, evaluation methodology, or detailed test conditions for its maximum results. [1]
The announcement does not state regions, account tiers, quotas, latency, supported video durations, file-size limits, data-retention conditions, or token prices. [1]
Google says the Gemini app feature is planned for all users on Flash and Flash-Lite models soon, but does not give a date or regional scope. Ask YouTube is described as a coming-months plan without stated timing or regional scope. [1]
SRC
Source desk
Direct links to the material behind this selection. Seeing the source matters as much as reading the synthesis.


