WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning
Abstract
Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.
Community
๐ฏIntroducing ๐ช๐ -๐ฉ๐๐ , a paradigm that lets the ๐ถ๐ป๐๐ฒ๐ฟ๐ป๐ฎ๐น ๐๐ผ๐ฟ๐น๐ฑ ๐บ๐ผ๐ฑ๐ฒ๐น generates the intermediate visual states that the ๐ฉ๐๐ uses to solve spatial reasoning tasks.
Some takeaways and thoughts:
๐ WM-VLM significantly improve over the SFTed VLM backbone, on tasks that require โ๐ถ๐บ๐ฎ๐ด๐ถ๐ป๐ฎ๐๐ถ๐ผ๐ปโ.
๐ชถ A ๐น๐ถ๐ด๐ต๐ ๐ ๐ผ๐ง ๐ฎ๐ฟ๐ฐ๐ต, with an understanding branch (for VLM) and a generation branch (for WM), is ๐ฒ๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ for generating visual thoughts and standard LLM verbal reasoning.
๐ฅช ๐๐ป๐๐ฒ๐ฟ๐น๐ฒ๐ฎ๐๐ฒ๐ฑ ๐๐ถ๐๐๐ฎ๐น-๐๐ฒ๐
๐๐๐ฎ๐น ๐ฐ๐ต๐ฎ๐ถ๐ป-๐ผ๐ณ-๐๐ต๐ผ๐๐ด๐ต๐ might be the next-gen format for VLMs perform reasoning, replacing the current LLM/VLMsโ verbal-only reasoning.
๐ญ Visual reasoning tasks (or maybe even embodied tasks) donโt require explicit pixel reconstruction, just generating the ๐น๐ฎ๐๐ฒ๐ป๐/๐ฒ๐บ๐ฏ๐ฒ๐ฑ๐ฑ๐ถ๐ป๐ด is enough. Pre-trained vision encoders (e.g., built-in model from Qwen2.5-VL) may already provide good representation for solving spatial reasoning tasks.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Distilling Visual Reasoning into Text Space (2026)
- Beyond Visual CoT: Internalized Visual Thinking for Proactive Video Reasoning (2026)
- GLaQ: Grounding Latent Queries in Visual Evidence for Multimodal Reasoning (2026)
- Scaffolding Minds: Optimizing Latent Visual Target Representations for Multimodal Reasoning (2026)
- Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence (2026)
- Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models (2026)
- Skip the Talk, Re-Focus on Vision: Latent Reasoning for Reasoning Segmentation in Multimodal Large Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.34826 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 2
yzha/WM-VLM-Tetris-2D
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper
