gen‑ai.news
← Back
Video

Google Rolls Out Agentic Video Understanding Across Gemini Models

Google Rolls Out Agentic Video Understanding Across Gemini Models

Google has launched what it describes as agentic video understanding within the Gemini model family, a capability that lets the models reason over video content across multiple steps without requiring continuous human direction. Rather than treating video analysis as a single-shot task, the agentic approach allows Gemini to plan, retrieve relevant segments, and refine its understanding iteratively - closer to how a human analyst might work through a long recording.

The practical benefits Google is highlighting are accuracy and cost. By structuring video analysis as an agentic process, the models can focus attention on relevant portions of a video rather than processing entire sequences uniformly, which Google says reduces token consumption and improves result quality on complex queries.

The update applies across Google's latest Gemini models, though Google has not specified exactly which model versions are included or the precise scope of supported video lengths. The announcement covers both the consumer Gemini product and the underlying API, meaning developers building on Gemini can also access the capability.

Agentic video reasoning is increasingly relevant as video becomes a primary format for training data, enterprise knowledge management, and creative production pipelines. The ability to query and reason over video with greater autonomy - rather than requiring users to clip or annotate content manually - reduces one of the more persistent friction points in working with video at scale.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.