gen‑ai.news
← Back
Video

See what 5 builders are making with Gemini Omni

See what 5 builders are making with Gemini Omni

Google has published a roundup highlighting five builders who are using Gemini Omni, the company's conversational multimodal model, to create and edit video. The piece is framed around the model's ability to handle video tasks through back-and-forth dialogue rather than through traditional timeline-based editing tools or complex prompt engineering. The five featured creators span a range of use cases, from rapid idea visualization to more structured video production work.

Gemini Omni sits within Google's broader Gemini model family and is distinguished by its ability to process and generate across text, image, audio, and video in a unified interface. The "conversation as editing" approach it enables is a meaningful departure from conventional video tools, where users typically need technical familiarity with software like Premiere Pro or DaVinci Resolve. By describing what they want in plain language, users can iterate on video content without a steep learning curve.

The builders featured in Google's post appear to be using Gemini Omni for tasks that would otherwise require either dedicated editing software or a separate AI pipeline - things like refining a rough cut, generating visual representations of concepts, or producing short-form content quickly. While Google's post functions partly as a showcase of its own product, the specific workflows described give a grounded sense of where conversational video AI is currently useful and where its limits might lie.

For those following the development of AI video tools, the roundup is a useful data point. It suggests that the most immediate value of models like Gemini Omni is not in fully autonomous video production, but in lowering the barrier for people who have ideas and need to move quickly - without relying on a dedicated editor or a technical background in post-production. As these models mature, the range and complexity of tasks handled through conversation is likely to expand.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Watching Roku’s AI channel is like eating from a trough
Video

Watching Roku’s AI channel is like eating from a trough

Roku has launched a 24/7 free ad-supported streaming channel dedicated entirely to AI-generated content, sourced from a startup called Fairground. The move marks one of the more visible attempts to bring generative video into mainstream living-room viewing. Whether audiences will warm to it is another question.

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Video

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0

Black Forest Labs has moved FLUX 3 Video out of early access and into general availability, offering Full HD video generation with clips up to 20 seconds long. The model includes native audio output and lip-synced dialogue across more than 14 languages. According to BFL's own Elo benchmark rankings, it outperforms both Gemini Omni Flash and Seedance 2.0.

China's MiniMax H3 is the first open model to top an AI video ranking
Video

China's MiniMax H3 is the first open model to top an AI video ranking

Chinese AI company MiniMax has released the weights for its H3 video generation model, marking the first time an open model has claimed the top spot on a major AI video benchmark ranking. The release is a notable moment for the open-source side of the generative video space, which has largely been outpaced by proprietary offerings from companies like OpenAI, Google, and Sora competitors. H3's rise to the top of the leaderboard signals that the gap between closed and open video models may be narr