--- license: apache-2.0 library_name: diffusers base_model: Wan-AI/Wan2.2-TI2V-5B pipeline_tag: image-to-video tags: - world-model - video-generation - camera-control - long-video - linear-attention --- # LOCI LOCI is a hybrid spatial-memory video world model built on Wan2.2-TI2V-5B. When a camera revisits a previously observed region, it reproduces what was there before: in half of the transformer blocks, attention keeps a key-value cache of past observations; in the other half, attention is restricted to the current chunk and complemented by a recurrent linear-attention memory whose reads and writes are conditioned on projective camera geometry. On the public MIND memory benchmark and on held-out recorded trajectories, LOCI reproduces revisited content more faithfully than representative world models and a same-recipe full-softmax model; with a bounded bank of retained observations, it streams long videos at constant memory. - Paper: [LOCI: Spatial Linear Memory for Streaming World Models](https://arxiv.org/abs/2609.40222) - Code: https://github.com/xiaji2021/LOCI ## Models This repository holds two models with the same architecture, each in its own subfolder: | subfolder | sampler | steps | |---|---|---| | `LOCI/` | `flow_euler` | 50 | | `LOCI-4step/` | `few_step` (distilled) | 4 | Each subfolder contains `config.json`, the shard index, 3 bf16 shards (about 11 GB, transformer only) and `SHA256SUMS`. The sampler and the number of steps are read from `config.json`; `--steps N` overrides the number of steps. ## Download Download one model: ```bash hf download sum0214/LOCI --include "LOCI-4step/*" --local-dir ./weights # 4-step model hf download sum0214/LOCI --include "LOCI/*" --local-dir ./weights # 50-step model ``` The VAE, text encoder and tokenizer come from the Wan2.2 base release (`Wan-AI/Wan2.2-TI2V-5B-Diffusers`); the code loads them from the Hub. ## Run Install the code from https://github.com/xiaji2021/LOCI, then: ```bash python examples/make_trajectory.py --preset look_around --chunks 16 --output traj.json python scripts/generate.py --weights ./weights/LOCI-4step --wan Wan-AI/Wan2.2-TI2V-5B-Diffusers \ --image first_frame.png --prompt "A sunlit stone temple courtyard with red lanterns." \ --trajectory traj.json --height 512 --width 768 --output out.mp4 ``` ## Recommended use - Full history (default) is recommended for clips up to about one minute. - `--history sparse` selects the paper's bounded-memory setting: first frame + a view bank + the most recent frames, at constant memory. - Resolution 480x864 or 512x768, one GPU. An 80 GB-class GPU (H100 / H200) is recommended. ## Speed Measured on one NVIDIA H200 with FlashAttention-3: LOCI (50 steps), 480x864, bfloat16, a recorded 300 s camera trajectory. Time is seconds per second of generated video (one chunk = 20 frames = 1.25 s at 16 fps). Memory is the peak GPU memory allocated during generation (text encoding and VAE decoding not included). | history | s per video second | peak GPU memory | measured over | |---|---|---|---| | full (`dense`, default) | 7.0 | 30.0 GiB | the first 40 s (time and memory grow with length) | | bounded (`sparse`) | 5.0 | 16.2 GiB | 300 s (constant time per chunk and constant memory) | ## Citation ```bibtex @article{xia2026loci, title={LOCI: Spatial Linear Memory for Streaming World Models}, author={Xia, Ji and Liao, Tingting and Liang, Xuezhi and Li, Hao and Liu, Guangyi}, journal={arXiv preprint arXiv:2609.40222}, year={2026} } ``` ## Acknowledgements LOCI builds on [Wan2.2](https://github.com/Wan-Video/Wan2.2). We thank the authors of [ARL²](https://arxiv.org/abs/2605.16579) (Li, Shah and Shang, *Attend Locally, Remember Linearly*): LOCI adopts their hybrid layer layout and follows their sampling procedure for the recurrent memory. ## License Apache-2.0.