gen‑ai.news
← Back
Video

World models that ignore human beliefs predict the wrong actions, new research shows

World models that ignore human beliefs predict the wrong actions, new research shows

World models - AI systems trained to simulate how environments unfold - have made significant progress in recent years, with projects like Sora and Genie demonstrating that models can learn surprisingly coherent physical dynamics from video. But a new line of research argues that these systems are missing something fundamental: the mental states of the people inside those environments. Without accounting for beliefs, desires, and intentions, the argument goes, a world model cannot reliably predict what a person will do next.

The new framework, called Mental World Modeling, proposes augmenting standard world models with explicit representations of mental variables alongside physical ones. The idea draws on longstanding concepts in cognitive science - particularly "theory of mind," the human capacity to attribute mental states to others and use them to anticipate behavior. Translating this into a computational framework means the model must track not just where objects are and how they move, but what an agent believes about those objects and what they are trying to achieve.

The research offers a striking empirical finding: language models that are relatively weak in terms of scale or general capability can outperform stronger models on action-prediction tasks when they are equipped with the Mental World Modeling approach. This suggests that the architecture and the information being modeled matter more than raw model size for this class of problems. It also points to a clear gap in how current world models are evaluated - benchmarks that focus on physical plausibility alone will not surface these shortcomings.

Perhaps the most useful insight from the work is its diagnosis of where the hardest problem lies. The biggest bottleneck is not representing mental states in isolation, but rather modeling how physical and mental states co-evolve - how a change in the environment updates someone's beliefs, and how those updated beliefs in turn shape their next action. That tight coupling between the physical and the mental is something current systems are poorly equipped to handle, and addressing it is likely to require both new training data and new modeling approaches tailored to that interaction.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

Vidu Q4 AI Video Model with Native Audio
Video

Vidu Q4 AI Video Model with Native Audio

Shengshu has released Vidu Q4, the latest version of its AI video generation model, adding native audio generation alongside improvements to motion quality and prompt adherence. The update positions Vidu Q4 as a more complete tool for creative video work, handling both visual and audio output within a single model. These changes reflect a broader trend in the field toward unified multimodal video generation.

Her AI-Generated Video Swayed the Judge. The Court Said it Carried 'Undue Emotional Weight'
Video

Her AI-Generated Video Swayed the Judge. The Court Said it Carried 'Undue Emotional Weight'

An AI-generated video of a murder victim was used as a victim impact statement in an Arizona court, with the deceased man's sister creating an avatar to speak in his place. While the judge was visibly moved by the presentation, an appeals court later found that the video carried "undue emotional weight" - raising serious questions about the role of generative AI in legal proceedings.

No image
Video

Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One

Reka has released Rho-1, a 19-billion-parameter model that handles text, images, video understanding, video generation, and robot action outputs within a single unified network. The architecture uses a shared key-value cache across all modalities, and a distilled variant can produce a 5.3-second video clip in roughly one second. The release is currently a research preview, with no public weights available.