Google’s New Gemini Omni AI Video Model Can Do Crazy Things
Google's new Gemini Omni artificial intelligence (AI) model can do some wild things. The model's key promise is to create anything from, well, anything. [Read More]
Google's new Gemini Omni artificial intelligence (AI) model can do some wild things. The model's key promise is to create anything from, well, anything. [Read More]
Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.
Free. Unsubscribe any time. No spam, ever.
Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.
Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.
Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.