File size: 5,473 Bytes
c8aa43a
be0f1ec
097d4b1
be0f1ec
 
 
 
 
 
c8aa43a
be0f1ec
 
 
 
 
bfa28d2
 
 
 
 
be0f1ec
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
24493bd
be0f1ec
 
 
 
 
 
 
 
 
 
 
 
 
 
c7ade81
be0f1ec
 
 
c7ade81
be0f1ec
 
 
24493bd
 
be0f1ec
 
 
c7ade81
 
be0f1ec
 
 
 
c7ade81
be0f1ec
 
 
 
 
 
 
24493bd
 
c7ade81
 
be0f1ec
 
 
 
 
 
 
c790363
 
 
 
 
be0f1ec
 
bfa28d2
 
 
 
 
 
 
 
097d4b1
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
---
license: cc-by-nc-4.0
pipeline_tag: image-to-video
tags:
- world-model
- diffusion
- retrieval
- navigation
- embodied-ai
---

# FAR checkpoints

Checkpoint bundles for **FAR**, a latent-diffusion world model with a learned,
action-conditioned retrieval memory, and its baselines, on the LoopNav,
SoundSpaces and AI2-THOR corpora.

- Paper: [Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models](https://arxiv.org/abs/2609.34677)
- Code: https://github.com/sony/far
- Project page: https://1202kbs.github.io/FAR-Project-Page/

## Layout

The repo mirrors the code's `models/` directory, so a download lands exactly
where the configs, launchers and notebooks look:

```
<corpus>/<arm>/config.yaml                  # rebuilds the model (Hydra)
<corpus>/<arm>/checkpoints/<step>.pth.tar   # EMA generator weights + memory state + step (inference only)
oasis_500m_vit_vae.pth                      # ViT-VAE tokenizer for LoopNav latents
agent_cell_detector.pt                      # latent agent-cell detector for the AI2-THOR-dyn probe
bundles.json                                # this listing with sizes and SHA-256
```

Bundles hold what inference needs (EMA generator, the trained or frozen memory
state, the training step) and cannot resume training. The SDXL VAE used for
the SoundSpaces and AI2-THOR corpora is fetched from `madebyollin/sdxl-vae-fp16-fix`
automatically.

## Download

```bash
python scripts/download_release.py                       # everything, ~13 GB
python scripts/download_release.py --corpus ai2thor_v3   # one corpus
python scripts/download_release.py --arm loopnav/far_multicue
```

or `huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models`.

## Bundles

| Corpus | Bundle | Arm | Step | Size |
|---|---|---|---|---|
| AI2-THOR-dyn | `ai2thor_dyn/far_meta` | FAR -- Meta | 500k | 0.56 GB |
| AI2-THOR-dyn | `ai2thor_dyn/far_multicue` | FAR -- Multi-Cue | 500k | 0.56 GB |
| AI2-THOR-dyn | `ai2thor_dyn/temporal` | Temporal | 500k | 0.49 GB |
| AI2-THOR-dyn | `ai2thor_dyn/worldmem` | WorldMem | 500k | 0.49 GB |
| AI2-THOR-dyn | `ai2thor_dyn/retriever_object` | retriever (object cue), init of the FAR arms | 12.5k | 0.06 GB |
| AI2-THOR v3 | `ai2thor_v3/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| AI2-THOR v3 | `ai2thor_v3/temporal` | Temporal | 700k | 0.49 GB |
| AI2-THOR v3 | `ai2thor_v3/worldmem` | WorldMem | 700k | 0.49 GB |
| AI2-THOR v3 | `ai2thor_v3/retriever_jepa` | retriever (visual), init of the FAR arm | 400k | 0.06 GB |
| LoopNav | `loopnav/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| LoopNav | `loopnav/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| LoopNav | `loopnav/far_visual` | FAR -- Visual | 700k | 0.55 GB |
| LoopNav | `loopnav/far_frozen_encoder` | ablation: the pre-trained visual retriever, frozen, no adapter | 700k | 0.55 GB |
| LoopNav | `loopnav/far_ablation_z` | ablation: FAR -- Multi-Cue with a fixed cue weight (0.5) | 700k | 0.55 GB |
| LoopNav | `loopnav/longlive_rag` | LongLive-RAG | 700k | 0.55 GB |
| LoopNav | `loopnav/temporal` | Temporal | 700k | 0.49 GB |
| LoopNav | `loopnav/worldmem` | WorldMem | 700k | 0.49 GB |
| LoopNav | `loopnav/retriever_jepa` | retriever (visual), init of the FAR arms | 400k | 0.06 GB |
| LoopNav | `loopnav/retriever_longlive` | retriever (content query), init of LongLive-RAG | 400k | 0.06 GB |
| SoundSpaces v1 | `soundspaces_v1/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v1 | `soundspaces_v1/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v1 | `soundspaces_v1/temporal` | Temporal | 700k | 0.49 GB |
| SoundSpaces v1 | `soundspaces_v1/worldmem` | WorldMem | 700k | 0.49 GB |
| SoundSpaces v1 | `soundspaces_v1/retriever_audio` | retriever (audio), init of the SoundSpaces v1 and v2 FAR arms | 400k | 0.06 GB |
| SoundSpaces v2 | `soundspaces_v2/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v2 | `soundspaces_v2/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v2 | `soundspaces_v2/temporal` | Temporal | 700k | 0.49 GB |
| SoundSpaces v2 | `soundspaces_v2/worldmem` | WorldMem | 700k | 0.49 GB |

"Temporal" and "WorldMem" are the recency and field-of-view retrieval
baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms
differ in the cues the retriever fuses (metadata, visual, multi-cue); the two LoopNav ablations
complete the paper's LoopNav ablation table. The retriever bundles are
the contrastively pre-trained encoders the FAR launchers start from; evaluation does not need
them (each FAR bundle already carries its trained retriever).

## License

CC BY-NC 4.0. The generator and diffusion code these weights belong to are
adapted from Navigation World Models and DiT (Meta Platforms, CC BY-NC 4.0);
see the code repository's THIRD_PARTY_NOTICES.md.

The SoundSpaces bundles (`soundspaces_v1/*`, `soundspaces_v2/*`) were trained on episodes
rendered from Matterport3D scenes, so their use is additionally subject to the
[Matterport3D Terms of Use](https://kaldir.vc.cit.tum.de/matterport/MP_TOS.pdf)
(non-commercial academic research).

## Citation

```bibtex
@article{kim2026far,
  title   = {Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models},
  author  = {Kim, Beomsu and Lai, Chieh-Hsin and Nguyen, Bac and Bar, Amir and Ye, Jong Chul and Mitsufuji, Yuki},
  journal = {arXiv preprint arXiv:2609.34677},
  year    = {2026},
  url     = {https://arxiv.org/abs/2609.34677}
}
```