Papers
arxiv:2609.34826

WM-VLM: Probing Internal World Models for Interleaved Visual-Textual Reasoning

Published on Sep 28
ยท Submitted by
y_zha
on Oct 5
Authors:
,
,
,
,
,
,
,

Abstract

Humans often solve spatial problems by mentally simulating visual transformations. In contrast, conventional vision-language models (VLMs) reason primarily through language. We investigate whether VLMs can solve spatial problems by reasoning with both text and generated visual states. To this end, we introduce WM-VLM, which equips a pretrained VLM with a lightweight world model branch for generating intermediate visual states. Our two-stage training first teaches the model to generate the next visual state and then to use that state for reasoning. We programmatically construct spatial reasoning tasks with verifiable intermediate visual states. These tasks allow us to evaluate how well the model generates visual states and how much it relies on them to answer the question. On 2D and 3D mental rotation tasks, WM-VLM consistently outperforms the supervised fine-tuned backbone, with gains of up to 39.25 percentage points. Ablations suggest that these gains depend on the generated visual states, as removing or corrupting them sharply reduces performance. Together, these results suggest that internal world models offer a promising path toward VLMs that reason in both language and visual space.

Community

Paper submitter

๐ŸŽฏIntroducing ๐—ช๐— -๐—ฉ๐—Ÿ๐— , a paradigm that lets the ๐—ถ๐—ป๐˜๐—ฒ๐—ฟ๐—ป๐—ฎ๐—น ๐˜„๐—ผ๐—ฟ๐—น๐—ฑ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น generates the intermediate visual states that the ๐—ฉ๐—Ÿ๐—  uses to solve spatial reasoning tasks.

Some takeaways and thoughts:
๐ŸŒŸ WM-VLM significantly improve over the SFTed VLM backbone, on tasks that require โ€˜๐—ถ๐—บ๐—ฎ๐—ด๐—ถ๐—ป๐—ฎ๐˜๐—ถ๐—ผ๐—ปโ€™.
๐Ÿชถ A ๐—น๐—ถ๐—ด๐—ต๐˜ ๐— ๐—ผ๐—ง ๐—ฎ๐—ฟ๐—ฐ๐—ต, with an understanding branch (for VLM) and a generation branch (for WM), is ๐—ฒ๐—ณ๐—ณ๐—ถ๐—ฐ๐—ถ๐—ฒ๐—ป๐˜ for generating visual thoughts and standard LLM verbal reasoning.
๐Ÿฅช ๐—œ๐—ป๐˜๐—ฒ๐—ฟ๐—น๐—ฒ๐—ฎ๐˜ƒ๐—ฒ๐—ฑ ๐˜ƒ๐—ถ๐˜€๐˜‚๐—ฎ๐—น-๐˜๐—ฒ๐˜…๐˜๐˜‚๐—ฎ๐—น ๐—ฐ๐—ต๐—ฎ๐—ถ๐—ป-๐—ผ๐—ณ-๐˜๐—ต๐—ผ๐˜‚๐—ด๐—ต๐˜ might be the next-gen format for VLMs perform reasoning, replacing the current LLM/VLMsโ€™ verbal-only reasoning.
๐Ÿ’ญ Visual reasoning tasks (or maybe even embodied tasks) donโ€™t require explicit pixel reconstruction, just generating the ๐—น๐—ฎ๐˜๐—ฒ๐—ป๐˜/๐—ฒ๐—บ๐—ฏ๐—ฒ๐—ฑ๐—ฑ๐—ถ๐—ป๐—ด is enough. Pre-trained vision encoders (e.g., built-in model from Qwen2.5-VL) may already provide good representation for solving spatial reasoning tasks.

fig2

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.34826
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 2

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.34826 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.34826 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.