gen‑ai.news
← Back
Video

See what 5 builders are making with Gemini Omni

See what 5 builders are making with Gemini Omni

Google has published a roundup highlighting five builders who are using Gemini Omni, the company's conversational multimodal model, to create and edit video. The piece is framed around the model's ability to handle video tasks through back-and-forth dialogue rather than through traditional timeline-based editing tools or complex prompt engineering. The five featured creators span a range of use cases, from rapid idea visualization to more structured video production work.

Gemini Omni sits within Google's broader Gemini model family and is distinguished by its ability to process and generate across text, image, audio, and video in a unified interface. The "conversation as editing" approach it enables is a meaningful departure from conventional video tools, where users typically need technical familiarity with software like Premiere Pro or DaVinci Resolve. By describing what they want in plain language, users can iterate on video content without a steep learning curve.

The builders featured in Google's post appear to be using Gemini Omni for tasks that would otherwise require either dedicated editing software or a separate AI pipeline - things like refining a rough cut, generating visual representations of concepts, or producing short-form content quickly. While Google's post functions partly as a showcase of its own product, the specific workflows described give a grounded sense of where conversational video AI is currently useful and where its limits might lie.

For those following the development of AI video tools, the roundup is a useful data point. It suggests that the most immediate value of models like Gemini Omni is not in fully autonomous video production, but in lowering the barrier for people who have ideas and need to move quickly - without relying on a dedicated editor or a technical background in post-production. As these models mature, the range and complexity of tasks handled through conversation is likely to expand.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.