gen‑ai.news
← Back
Video

Xiaomi’s MiLM Plus Releases PROVE: Perception-Aligned Object Removal Metrics RC-S and RC-T With a Real-World Video Benchmark

Video object removal has matured considerably in recent years, with diffusion-based models now capable of plausibly reconstructing shadows, reflections, and structures that were partially hidden behind removed objects. The problem is that the metrics used to score these models have not kept pace. Widely used measures such as PSNR, SSIM, LPIPS, ReMOVE, and CFD frequently produce rankings that conflict with how humans actually perceive the quality of an edited video.

The core difficulty is structural. Object removal is an ill-posed, one-to-many problem - there is rarely a single correct way to fill the space left behind, meaning no definitive ground-truth frame exists to compare against. Pixel-level distance metrics assume such a ground truth, so they are working against the nature of the task from the start. This leads to situations where a visually convincing result scores worse than a blurry or artifact-laden one, simply because the convincing result chose a different but equally valid reconstruction.

To address this, Xiaomi's MiLM Plus team has released PROVE, which stands for Perception-aligned Object Removal Video Evaluation. The framework introduces two new metrics: RC-S, which measures spatial coherence - how well the reconstructed region fits the surrounding scene in a single frame - and RC-T, which measures temporal coherence, assessing whether the filled region remains consistent and stable across frames over time. Both metrics are designed to align more closely with human perceptual judgments rather than pixel-level ground-truth comparisons. The team also releases a real-world benchmark dataset to support standardized testing, which has been a secondary gap in the field since most prior evaluations relied on synthetic or semi-synthetic data.

The practical significance here is that better metrics can change which models get developed and prioritized. If researchers optimize for PSNR, they build different systems than if they optimize for perceptual coherence. PROVE represents an attempt to realign those incentives by giving the field evaluation tools that reward what actually matters visually. Whether RC-S and RC-T gain broad adoption will depend on how well they hold up across diverse removal scenarios, but the benchmark dataset at least provides a shared testing ground for that conversation to happen.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Google launches Gemini 3.8 Live with Live Avatar
Video

Google launches Gemini 3.8 Live with Live Avatar

Google has launched Gemini 3.8 Live with Live Avatar, a feature set that brings animated video personas to enterprise AI agents. The update also introduces background tool calls and multilingual speech support spanning 97 languages. Together, these additions push conversational AI agents closer to a more naturalistic, face-to-face interaction model for business use cases.

No image
Video

Runway’s WorldPrompt and the Engineering of Real-Time Worlds

Runway's WorldPrompt system powers its Gen World Models 2 (GWM 2), enabling real-time generation of video and audio through persistent context and timed actions. Rather than producing discrete clips, the model maintains a continuous understanding of an environment as it unfolds. The approach marks a notable shift in how world models can be steered interactively.

Gemini 3.8 Live with Live Avatar gives Google’s AI a face
Video

Gemini 3.8 Live with Live Avatar gives Google’s AI a face

Google has updated Gemini Live with an animated avatar that lip-syncs and displays facial expressions in real time during conversations. Called Live Avatar, the feature is currently limited to Gemini Enterprise customers and supports 97 languages without degrading video quality. It marks Google's latest step toward giving its AI assistant a visible, expressive presence.