Seeing Isn't Reasoning: New Benchmark Challenges Multimodal AI Models
Humans often solve complex problems by imagining them. We mentally rotate objects, visualize spatial relationships, or picture possible solutions before putting them into words. As AI systems become increasingly capable of generating both text and images, an important question emerges: can they do the same?
A new study by researchers at the ELLIS Institute explores exactly that question. A team around Principal Investigator Wieland Brendel spearheaded by Jana Zeller introduces MentisOculi, a benchmark designed to test whether frontier multimodal AI models can use intermediate visualizations as part of their reasoning process and whether these visual "thoughts" actually improve performance.
Unlike traditional vision-language benchmarks, MentisOculi consists of procedurally generated, multi-step reasoning tasks that are intuitive for humans to solve visually but difficult to describe purely in text. The benchmark also provides visual reasoning steps, allowing researchers to separate failures in reasoning from failures in image generation.
The team evaluated a wide range of state-of-the-art multimodal large language models, unified multimodal models capable of generating both text and images, video models, as well as models using latent visual representations. Across all approaches, visual reasoning strategies failed to consistently outperform text-only reasoning.
Their analysis revealed a key limitation. Although many models were individually capable of generating accurate images and solving the reasoning tasks in text, they struggled to combine these abilities effectively. Small errors accumulated over multiple reasoning steps, and remarkably, models often failed to benefit even when provided with correct intermediate visualizations.
These findings suggest that today's multimodal AI systems have not yet developed an equivalent of human mental imagery that can reliably support reasoning. Instead, the gap between generating visual content and using it effectively for problem-solving remains a significant challenge.
By providing a controlled and extensible framework for studying visual reasoning, MentisOculi establishes an important benchmark for future research. The authors hope it will help the community better understand how multimodal models reason and what advances will be needed before visual "thoughts" become a practical tool for AI.
This work by Jana Zeller, Thaddaus Wiedemer, Fanfei Li, Thomas Klein, Prasanna Mayilvahanan of the ELLIS Institute, Matthias Bethge (Tübingen AI Center), Felix Wichmann (Uni Tübingen), Ryan Cotterell (ETH), Wieland Brendel (ELLIS Institute Tübingen) was accepted to ICML 2026, the Forty-Third International Conference on Machine Learning.
Read the paper here.
Find out more about Wieland’s research group.
Are you smarter than AI? Play the game.