gen‑ai.news
← Back
Multimodal

New benchmark confirms AI models still perform poorly at visual perception

New benchmark confirms AI models still perform poorly at visual perception

Moonshot AI has released PerceptionBench, a benchmark designed to measure how well multimodal AI models can genuinely interpret visual information - separate from their capacity for logical reasoning. The distinction matters because most existing evaluations bundle the two abilities together, making it difficult to identify where a model actually fails. PerceptionBench attempts to isolate the perception step, testing whether a model can correctly extract what is in an image before any higher-order thinking takes place.

The results across frontier models are consistently weak. No model in the evaluation reaches 60 percent accuracy on the benchmark, and the top performer, GPT-4o Sol, holds only a narrow lead over its competitors. That kind of ceiling, across models from multiple major labs, points to a shared and fundamental limitation rather than a gap that any one team has quietly solved.

One of the more significant findings is that errors researchers have long labeled as reasoning failures may be misattributed. If a model misreads key details in an image at the perception stage, any downstream reasoning built on that flawed input will also be wrong - but the mistake looks, from the outside, like a logic error. PerceptionBench's structure makes it possible to catch failures earlier in that chain, which could shift how researchers prioritize improvements to multimodal architectures.

The benchmark adds to a growing body of work questioning how deeply current vision-language models actually "see" versus pattern-match on statistical regularities in training data. For developers building applications that depend on accurate visual understanding - medical imaging tools, document analysis, autonomous systems - the results are a useful reminder that model confidence and model accuracy are not the same thing. Closing the perception gap may require architectural changes or new training approaches that go beyond simply scaling existing systems.

Enjoy this story? Get the next one in your inbox.

Twice a week: the most important stories in generative image and video AI, distilled into a 2-minute read.

Free. Unsubscribe any time. No spam, ever.

Your next read

You can now turn off Google Gemini’s visible watermarks
Multimodal

You can now turn off Google Gemini’s visible watermarks

Google has added a toggle in Gemini and its AI video tool Flow that lets users remove the visible "sparkle" watermark from AI-generated images, videos, and music. Even with the visible mark turned off, content will still carry invisible SynthID watermarks and C2PA metadata. The change affects content produced by Google's Nano Banana and Omni models.

Twitch streamers can now opt out from training Amazon’s AI
Multimodal

Twitch streamers can now opt out from training Amazon’s AI

Twitch has introduced an opt-out setting that lets streamers prevent their content - including streams, VODs, clips, chat logs, and channel images - from being used to train Amazon's generative AI models. The control applies to future training only and covers AI systems designed to generate or synthesize text, audio, images, or video. Non-generative AI features such as captions and safety tools are unaffected by the setting.

Claude will apply invisible watermarks to AI text and images
Multimodal

Claude will apply invisible watermarks to AI text and images

Anthropic has announced plans to embed invisible watermarks into text and images produced by Claude, making it easier for platforms and users to identify AI-generated content. The move is tied to European regulatory requirements around AI transparency. The changes are not yet live, but represent a formal commitment from the company.