Papers
arxiv:2610.02521

Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory

Published on Oct 1
· Submitted by
ying yang
on Oct 5
Authors:
,
,
,
,
,
,

Abstract

Long-video generation and world models have shown strong potential for interactive entertainment and embodied simulation by predicting future observations conditioned on user actions and historical memory. However, as memory sequences grow longer and their structures become increasingly complex, managing long-range spatial context becomes increasingly challenging, calling for a more intelligent and systematic memory-management strategy. Building on the advancing spatial reasoning capabilities of multimodal large language models (MLLMs) and the broader vision of unified models, we propose Spatial Memory Intelligence (SMI), the first framework to systematically employ an understanding model for spatial-memory management in long-video world models. SMI introduces four coordinated atomic operations: spatial clustering, within-cluster sparsification, action-aware retrieval, and reliability-aware filtering. Extensive experiments across multiple baselines, benchmarks, and world-model backbones demonstrate the effectiveness and generalizability of SMI, achieving comprehensive improvements in memory sparsity, generation stability, and spatial consistency.

Community

Paper author Paper submitter

We introduce Spatial Memory Intelligence (SMI), the first unified framework for spatial-memory management in world models driven by a multimodal understanding model. Unlike existing work that focuses on individual objectives such as memory compression, consistency, or stability, SMI addresses these challenges systematically from a unified perspective. Through four coordinated atomic operations, SMI organizes, sparsifies, filters, and retrieves spatial memory, substantially reducing redundant memories while improving both generation stability and spatial consistency.

Our central motivation is that spatial-memory management involves spatial perception and reasoning, tracking semantic changes, and understanding spatial relationships across time and viewpoints. These demands call for intelligent management by multimodal understanding models equipped with these capabilities. The strong spatial perception and reasoning capabilities demonstrated by GPT-6-Astra further strengthen our interest in the potential for spatial-memory management to improve as data and model size scale. We hope SMI takes a step toward spatial-memory management within future unified models.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.02521
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2610.02521 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.02521 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.02521 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.