
Gemini Agentic Video
Analyzes lengthy video content by intelligently selecting specific audio, transcripts, frames, and playback speeds to locate information quickly. Useful for researchers and creators navigating extensive media archives.
The Blend
Google DeepMind announced a new system called agentic video understanding for its Gemini AI lineup, including models like Gemini 3.7 Flash and 3.5 Flash-Lite. Rather than processing entire recordings at a rigid frame rate, the model acts like a human researcher. It selectively skips through visual frames, audio clips, and text transcripts to pinpoint target information quickly.
This approach targets a major bottleneck in artificial intelligence: analyzing long media files without incurring massive computational expenses. Google reported on its official blog that this dynamic scanning method cuts token usage by as much as 88 percent and lowers processing costs by up to 66 percent. Simultaneously, the company claims video analysis accuracy increases by up to 7 percent. Developers can start using the feature immediately through Google AI Studio and the enterprise agent platform.
While these financial savings are promising for businesses handling media archives, real-world performance will depend on how accurately the agent decides where to look. If the model accidentally skips over critical visual context during rapid scanning, it could miss crucial moments. It remains unclear whether enterprise customers will comfortably trust these automated shortcuts for high-stakes tasks like security surveillance or forensic video review.
Written independently by AI News Smoothie from the reporting listed below. Facts belong to the original publishers. Follow the links for their full coverage.
Ingredients
- Introducing Agentic Video in Gemini
Google updated its Gemini AI models to dynamically scan video content, lowering processing costs by up to 66 percent.