Papers
arxiv:2610.00544

Memorizon: Training World Models Beyond Their Context Window

Published on Sep 30
ยท Submitted by
Tingting Liao
on Oct 2
Authors:
,
,

Abstract

Streaming world models should render a place consistently across repeated visits. Directly supervising such revisits requires training samples that capture both visits, often spanning minutes. Yet dense attention over the full span incurs quadratic costs, making long-span supervision expensive. Memorizon breaks this coupling: long spans are needed for supervision, but not for attention, since the two visits can share a forward pass without including every intervening frame. A training sample covers a span of any length but is scored only on its last k chunks. Instead of tokenizing the history before them, each scored chunk retrieves its own top-K latents by camera co-visibility, and the union of these requests forms a shared bank. The bank is bounded by kK, so the sequence stays bounded however long the span; at the shortest span the recipe is exactly conventional training. Adding the bank raises the cost of a step once; beyond that, a longer span costs little, and going from 100 to 400 s adds 12% to the step time. Against a sliding-window baseline, retrieval raises revisit consistency on every split, and a span long enough to reach the first visit of each return adds a further 24% to 30%, at some cost in image quality; beyond that span, more length no longer helps. Filling the bank from another episode lowers revisit correlation by 83%, so the model uses what it retrieves. Project page: https://tingtingliao.github.io/memorizon

Community

Paper author Paper submitter

Memorizon trains a camera-controlled video world model on spans far longer than its context window. Each training sample packs up to 100-400s of history into a fixed-size sequence: every chunk retrieves the past frames whose camera frustums overlap its own. Training therefore costs the same as a 10 s window, yet the model learns to remember scenes it saw minutes earlier. Built on Wan2.2-TI2V-5B and distilled to 4 steps with Self-Forcing, it keeps returning views consistent over minute-long interactive rollouts.

๐ŸŒ Project page: https://tingtingliao.github.io/memorizon/
๐Ÿ’ป Code: https://github.com/TingtingLiao/memorizon
๐Ÿค— Weights: https://huggingface.co/Luffuly/memorizon

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2610.00544
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 1

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2610.00544 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2610.00544 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.