FAR-Checkpoints / README.md
1202kbs's picture nielsr's picture
nielsr HF Staff
Add pipeline tag (#1)
097d4b1
|
Raw History Blame Contribute Delete
5.47 kB
---
license: cc-by-nc-4.0
pipeline_tag: image-to-video
tags:
- world-model
- diffusion
- retrieval
- navigation
- embodied-ai
---
# FAR checkpoints
Checkpoint bundles for **FAR**, a latent-diffusion world model with a learned,
action-conditioned retrieval memory, and its baselines, on the LoopNav,
SoundSpaces and AI2-THOR corpora.
- Paper: [Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models](https://arxiv.org/abs/2609.34677)
- Code: https://github.com/sony/far
- Project page: https://1202kbs.github.io/FAR-Project-Page/
## Layout
The repo mirrors the code's `models/` directory, so a download lands exactly
where the configs, launchers and notebooks look:
```
<corpus>/<arm>/config.yaml # rebuilds the model (Hydra)
<corpus>/<arm>/checkpoints/<step>.pth.tar # EMA generator weights + memory state + step (inference only)
oasis_500m_vit_vae.pth # ViT-VAE tokenizer for LoopNav latents
agent_cell_detector.pt # latent agent-cell detector for the AI2-THOR-dyn probe
bundles.json # this listing with sizes and SHA-256
```
Bundles hold what inference needs (EMA generator, the trained or frozen memory
state, the training step) and cannot resume training. The SDXL VAE used for
the SoundSpaces and AI2-THOR corpora is fetched from `madebyollin/sdxl-vae-fp16-fix`
automatically.
## Download
```bash
python scripts/download_release.py # everything, ~13 GB
python scripts/download_release.py --corpus ai2thor_v3 # one corpus
python scripts/download_release.py --arm loopnav/far_multicue
```
or `huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models`.
## Bundles
| Corpus | Bundle | Arm | Step | Size |
|---|---|---|---|---|
| AI2-THOR-dyn | `ai2thor_dyn/far_meta` | FAR -- Meta | 500k | 0.56 GB |
| AI2-THOR-dyn | `ai2thor_dyn/far_multicue` | FAR -- Multi-Cue | 500k | 0.56 GB |
| AI2-THOR-dyn | `ai2thor_dyn/temporal` | Temporal | 500k | 0.49 GB |
| AI2-THOR-dyn | `ai2thor_dyn/worldmem` | WorldMem | 500k | 0.49 GB |
| AI2-THOR-dyn | `ai2thor_dyn/retriever_object` | retriever (object cue), init of the FAR arms | 12.5k | 0.06 GB |
| AI2-THOR v3 | `ai2thor_v3/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| AI2-THOR v3 | `ai2thor_v3/temporal` | Temporal | 700k | 0.49 GB |
| AI2-THOR v3 | `ai2thor_v3/worldmem` | WorldMem | 700k | 0.49 GB |
| AI2-THOR v3 | `ai2thor_v3/retriever_jepa` | retriever (visual), init of the FAR arm | 400k | 0.06 GB |
| LoopNav | `loopnav/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| LoopNav | `loopnav/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| LoopNav | `loopnav/far_visual` | FAR -- Visual | 700k | 0.55 GB |
| LoopNav | `loopnav/far_frozen_encoder` | ablation: the pre-trained visual retriever, frozen, no adapter | 700k | 0.55 GB |
| LoopNav | `loopnav/far_ablation_z` | ablation: FAR -- Multi-Cue with a fixed cue weight (0.5) | 700k | 0.55 GB |
| LoopNav | `loopnav/longlive_rag` | LongLive-RAG | 700k | 0.55 GB |
| LoopNav | `loopnav/temporal` | Temporal | 700k | 0.49 GB |
| LoopNav | `loopnav/worldmem` | WorldMem | 700k | 0.49 GB |
| LoopNav | `loopnav/retriever_jepa` | retriever (visual), init of the FAR arms | 400k | 0.06 GB |
| LoopNav | `loopnav/retriever_longlive` | retriever (content query), init of LongLive-RAG | 400k | 0.06 GB |
| SoundSpaces v1 | `soundspaces_v1/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v1 | `soundspaces_v1/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v1 | `soundspaces_v1/temporal` | Temporal | 700k | 0.49 GB |
| SoundSpaces v1 | `soundspaces_v1/worldmem` | WorldMem | 700k | 0.49 GB |
| SoundSpaces v1 | `soundspaces_v1/retriever_audio` | retriever (audio), init of the SoundSpaces v1 and v2 FAR arms | 400k | 0.06 GB |
| SoundSpaces v2 | `soundspaces_v2/far_meta` | FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v2 | `soundspaces_v2/far_multicue` | FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v2 | `soundspaces_v2/temporal` | Temporal | 700k | 0.49 GB |
| SoundSpaces v2 | `soundspaces_v2/worldmem` | WorldMem | 700k | 0.49 GB |
"Temporal" and "WorldMem" are the recency and field-of-view retrieval
baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms
differ in the cues the retriever fuses (metadata, visual, multi-cue); the two LoopNav ablations
complete the paper's LoopNav ablation table. The retriever bundles are
the contrastively pre-trained encoders the FAR launchers start from; evaluation does not need
them (each FAR bundle already carries its trained retriever).
## License
CC BY-NC 4.0. The generator and diffusion code these weights belong to are
adapted from Navigation World Models and DiT (Meta Platforms, CC BY-NC 4.0);
see the code repository's THIRD_PARTY_NOTICES.md.
The SoundSpaces bundles (`soundspaces_v1/*`, `soundspaces_v2/*`) were trained on episodes
rendered from Matterport3D scenes, so their use is additionally subject to the
[Matterport3D Terms of Use](https://kaldir.vc.cit.tum.de/matterport/MP_TOS.pdf)
(non-commercial academic research).
## Citation
```bibtex
@article{kim2026far,
title = {Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models},
author = {Kim, Beomsu and Lai, Chieh-Hsin and Nguyen, Bac and Bar, Amir and Ye, Jong Chul and Mitsufuji, Yuki},
journal = {arXiv preprint arXiv:2609.34677},
year = {2026},
url = {https://arxiv.org/abs/2609.34677}
}
```