gen‑ai.news
← Back
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway has detailed the architecture behind Gen World Models 2 (GWM 2), its real-time world simulation system, through a technical discussion published on Latent Space. At the center of the system is WorldPrompt - a mechanism that allows users and developers to steer a continuously running generative model using persistent context and timed actions, producing synchronized video and audio output as the world evolves.

The key distinction of GWM 2 is its treatment of generation as an ongoing process rather than a series of independent clips. Persistent context means the model retains information about the state of a scene across time, allowing it to behave more like a simulation engine than a standard video diffusion model. Timed actions let inputs be scheduled or triggered at specific moments, giving developers fine-grained control over how a generated world responds to interventions.

This has practical implications for interactive applications - particularly games, simulations, and training environments for robotics or autonomous systems. World models that can generate plausible, temporally consistent environments in real time have long been a research goal, since they could replace or supplement expensive physical simulators or hand-authored game engines. Runway's approach is notable for packaging this capability into a system that appears designed with developer integration in mind, not just research demonstration.

The audio component is also worth noting. Generating spatially and temporally coherent audio alongside video is a harder problem than it might seem - sound must respond to the same contextual state as the visuals, and desynchronization is immediately perceptible to users. Including audio within the world model's output pipeline, rather than treating it as a separate layer, suggests Runway is aiming for a more unified generative environment rather than a purely visual tool. How GWM 2 performs at scale and under varied prompting conditions will be the real test of whether WorldPrompt delivers the kind of controllability that applied use cases demand.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.

KPop Demon Hunters: How Sony Pictures Imageworks Used Adobe in Creating the Global Phenomenon
Video

KPop Demon Hunters: How Sony Pictures Imageworks Used Adobe in Creating the Global Phenomenon

Sony Pictures Imageworks texture and motion graphics artists detail how Adobe Substance 3D and After Effects formed the backbone of the production pipeline for KPop Demon Hunters, the 2026 Netflix animated film nominated for both Academy Award and Golden Globe honors. From generating over a thousand crowd character variations to crafting intricate Korean embroidery textures, the tools shaped both the scale and the fine detail of the film's distinctive look. Two key artists walk through the speci