gen‑ai.news
← Back
Video

Introducing agentic video understanding with Gemini

Google DeepMind has announced the addition of agentic video understanding to Gemini, expanding the model's ability to work with video content in a more autonomous and structured way. Rather than simply answering questions about a video when asked, the agentic approach allows Gemini to proactively identify relevant segments, track events across time, and reason over what is happening without requiring step-by-step human direction.

The distinction between standard video comprehension and agentic video understanding is meaningful. Conventional multimodal models treat video as a passive input - a user asks a question, and the model responds. An agentic setup means the model can decompose a complex video-related task into sub-steps, decide what information to focus on, and work toward a goal over multiple internal reasoning stages. This makes it better suited to longer videos and more open-ended tasks, such as summarizing narrative arcs, identifying anomalies, or extracting structured data from recorded footage.

Gemini already supports long-context inputs, including extended video clips, which provides a foundation for this kind of temporal reasoning. The agentic layer builds on that capability by giving the model a more active role in deciding how to process and navigate that content. This is particularly relevant for use cases in fields like media production, security, research, and accessibility, where video content is dense and extracting specific information manually is time-consuming.

The move reflects a broader trend across the AI industry toward agentic systems - models that do not just respond, but plan and act across multi-step tasks. Applying that framing specifically to video is a natural extension, given how much information is encoded in motion, timing, and visual context that static image or text models cannot fully capture. Google DeepMind's continued investment in video understanding signals that video is increasingly treated as a first-class modality, not an add-on to language or image capabilities.

Also covered by

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.