gen‑ai.news
← Back
Video

Microsoft Research's Mirage gives video generation a persistent spatial memory that doesn't forget what's around the corner

Microsoft Research's Mirage gives video generation a persistent spatial memory that doesn't forget what's around the corner

Most video generation models struggle with spatial continuity - pan the camera away from a scene and return to it, and details have often shifted or disappeared entirely. Mirage, a collaborative project from Microsoft Research and several universities, addresses this by giving the model a persistent memory of the space it has already generated, so previously seen areas remain coherent when revisited.

The core technical distinction in Mirage is where scene information is stored. Traditional approaches often rely on pixel-based point clouds - explicit 3D representations derived from rendered frames. Mirage instead encodes and retains scene data directly in latent space, the compressed internal representation that diffusion-based models already work within. This means the system does not need to reconstruct geometry from pixels every time it needs to reference what came before.

That design choice has practical consequences beyond consistency. Working in latent space rather than maintaining dense point cloud structures cuts both processing time and graphics memory consumption meaningfully, which matters for research scalability and any potential downstream deployment. The result is a model that can handle extended camera trajectories - moving through a corridor, circling a room - without the scene fragmenting or contradicting itself across segments.

The system is not without its current boundaries. Mirage handles static environments well but has not yet solved the harder problem of tracking moving objects reliably across video segments. A person or vehicle that exits the frame and re-enters may not be rendered consistently, which limits the model's usefulness for dynamic scene simulation. That gap points to the next natural area of development for world models of this kind - integrating persistent spatial memory with robust object-level tracking to handle scenes where not everything stays still.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0
Video

Black Forest Labs makes FLUX 3 Video generally available and claims it beats Seedance 2.0

Black Forest Labs has moved FLUX 3 Video out of early access and into general availability, offering Full HD video generation with clips up to 20 seconds long. The model includes native audio output and lip-synced dialogue across more than 14 languages. According to BFL's own Elo benchmark rankings, it outperforms both Gemini Omni Flash and Seedance 2.0.

China's MiniMax H3 is the first open model to top an AI video ranking
Video

China's MiniMax H3 is the first open model to top an AI video ranking

Chinese AI company MiniMax has released the weights for its H3 video generation model, marking the first time an open model has claimed the top spot on a major AI video benchmark ranking. The release is a notable moment for the open-source side of the generative video space, which has largely been outpaced by proprietary offerings from companies like OpenAI, Google, and Sora competitors. H3's rise to the top of the leaderboard signals that the gap between closed and open video models may be narr

Is paying artists enough to convince them to embrace AI?
Video

Is paying artists enough to convince them to embrace AI?

A new wave of AI startups is attempting to address longstanding concerns from the illustration community by compensating artists whose work is used in model training. Pippa is one such company, positioning itself as a more ethically grounded alternative to competitors that have trained on unlicensed work. Whether financial compensation alone is enough to shift artist sentiment remains an open question.