xAI adds character references and 1080p to Imagine Video 1.5

xAI has rolled out Imagine Video 1.5, a notable update to its AI video generation model that introduces several new input and output capabilities. The headline additions are image-based character references and voice references, which allow users to anchor the appearance and sound of characters in generated video to specific source material rather than relying entirely on text descriptions.
The update also introduces a prompt-only mode, meaning users who do not have reference images or audio can still generate video from text alone - keeping the workflow accessible while the reference-based features serve those who need tighter creative control. Multi-reference support extends this further, letting users feed in more than one reference simultaneously to guide a single generation, which is useful for scenes involving multiple characters or blended visual styles.
On the output side, Imagine Video 1.5 now renders natively at 1080p. Previous generations of AI video tools have often topped out at lower resolutions, so native full-HD output is a meaningful step for users who want footage that holds up in production or on larger screens without upscaling artifacts.
xAI's Imagine suite - which also includes its image generation tool - sits within the broader Grok ecosystem the company has been building. Pushing Imagine Video toward more reference-driven, higher-resolution output positions it closer to competing tools from providers like Runway, Kling, and others that have made character consistency a key selling point. Whether the reference fidelity holds up under varied prompts and complex scenes will be the practical test for creators considering it as part of their workflow.


