gen‑ai.news
← Back
Video

xAI adds character references and 1080p to Imagine Video 1.5

xAI adds character references and 1080p to Imagine Video 1.5

xAI has rolled out Imagine Video 1.5, a notable update to its AI video generation model that introduces several new input and output capabilities. The headline additions are image-based character references and voice references, which allow users to anchor the appearance and sound of characters in generated video to specific source material rather than relying entirely on text descriptions.

The update also introduces a prompt-only mode, meaning users who do not have reference images or audio can still generate video from text alone - keeping the workflow accessible while the reference-based features serve those who need tighter creative control. Multi-reference support extends this further, letting users feed in more than one reference simultaneously to guide a single generation, which is useful for scenes involving multiple characters or blended visual styles.

On the output side, Imagine Video 1.5 now renders natively at 1080p. Previous generations of AI video tools have often topped out at lower resolutions, so native full-HD output is a meaningful step for users who want footage that holds up in production or on larger screens without upscaling artifacts.

xAI's Imagine suite - which also includes its image generation tool - sits within the broader Grok ecosystem the company has been building. Pushing Imagine Video toward more reference-driven, higher-resolution output positions it closer to competing tools from providers like Runway, Kling, and others that have made character consistency a key selling point. Whether the reference fidelity holds up under varied prompts and complex scenes will be the practical test for creators considering it as part of their workflow.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.

No image
Video

Introducing Gemini 3.8 Live with Live Avatar

Google DeepMind has introduced Gemini 3.8 Live, an updated multimodal model paired with a new Live Avatar feature that generates an animated, talking on-screen presence during real-time conversations. The combination allows users to interact with a responsive visual agent rather than a purely voice-based interface. The release marks another step in Google's effort to make AI interactions feel more immediate and embodied.