Image-to-Video
Safetensors
English
video-generation
text-to-video
memory
File size: 2,331 Bytes
f463e8c
016a461
 
 
f463e8c
 
 
2192296
f463e8c
 
 
 
 
 
b77184d
f463e8c
 
 
 
 
 
b77184d
f463e8c
03eb153
 
 
 
f463e8c
 
 
 
 
 
 
 
 
 
b77184d
f463e8c
 
03eb153
f463e8c
 
 
016a461
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
---
license: other
license_name: mosaichunk-backbone-licenses
license_link: https://github.com/mosaichunk/MosaiChunk#license
language:
- en
tags:
- arxiv:2610.02153
- video-generation
- text-to-video
- image-to-video
- memory
- safetensors
datasets:
- mosaichunk/RememBench
---

# MosaiChunk

Learned memory-router checkpoints for **MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation**.

[Project page](https://mosaichunk.github.io/) 路 [Paper](https://arxiv.org/abs/2610.02153) 路 [Code](https://github.com/mosaichunk/MosaiChunk) 路 [RememBench](https://huggingface.co/datasets/mosaichunk/RememBench)

| Checkpoint | Backbone |
|---|---|
| [`t2v/model.safetensors`](t2v/model.safetensors) | RAVEN-adapted MiniMax-H3 (H3-AR) |
| [`i2v/model.safetensors`](i2v/model.safetensors) | LingBot-World-Infinity |

Each folder contains router weights and `config.json`. The weights are exported without changing tensor values; optimizer and training state are excluded.

These checkpoints require their corresponding frozen video backbone. T2V also requires the pretrained RAVEN streaming adapter. Backbone and adapter weights are not bundled here.

## Download

```python
from huggingface_hub import snapshot_download

snapshot_download("mosaichunk/MosaiChunk", local_dir="checkpoints/MosaiChunk")
```

See the [code repository](https://github.com/mosaichunk/MosaiChunk) for training and inference.

## License

Each router checkpoint is subject to the license of its backbone:

- `i2v/`: LingBot-World-v2, [CC BY-NC-SA 4.0](https://creativecommons.org/licenses/by-nc-sa/4.0/)
- `t2v/`: MiniMax-H3 with the RAVEN streaming adapter, [MiniMax-H3 Community License Agreement](https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/LICENSE)

The checkpoints do not include backbone or adapter weights and grant no rights to them.

## Citation

```bibtex
@misc{zhang2026mosaichunkcompositingspatiotemporalmemory,
      title={MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation},
      author={Yiwen Zhang and Haocheng Xi and Michael Tian-Yue Liu and Alexei A. Efros and Hadar Averbuch-Elor and Qianqian Wang and Haiwen Feng},
      year={2026},
      eprint={2610.02153},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.02153},
}
```