FAR checkpoints
Checkpoint bundles for FAR, a latent-diffusion world model with a learned, action-conditioned retrieval memory, and its baselines, on the LoopNav, SoundSpaces and AI2-THOR corpora.
- Paper: Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models
- Code: https://github.com/sony/far
- Project page: https://1202kbs.github.io/FAR-Project-Page/
Layout
The repo mirrors the code's models/ directory, so a download lands exactly
where the configs, launchers and notebooks look:
<corpus>/<arm>/config.yaml # rebuilds the model (Hydra)
<corpus>/<arm>/checkpoints/<step>.pth.tar # EMA generator weights + memory state + step (inference only)
oasis_500m_vit_vae.pth # ViT-VAE tokenizer for LoopNav latents
agent_cell_detector.pt # latent agent-cell detector for the AI2-THOR-dyn probe
bundles.json # this listing with sizes and SHA-256
Bundles hold what inference needs (EMA generator, the trained or frozen memory
state, the training step) and cannot resume training. The SDXL VAE used for
the SoundSpaces and AI2-THOR corpora is fetched from madebyollin/sdxl-vae-fp16-fix
automatically.
Download
python scripts/download_release.py # everything, ~12 GB
python scripts/download_release.py --corpus ai2thor_v3 # one corpus
python scripts/download_release.py --arm loopnav/far_multicue
or huggingface-cli download 1202kbs/FAR-Checkpoints --local-dir models.
Bundles
| Corpus | Bundle | Arm | Step | Size |
|---|---|---|---|---|
| AI2-THOR-dyn | ai2thor_dyn/far_meta |
FAR -- Meta | 500k | 0.56 GB |
| AI2-THOR-dyn | ai2thor_dyn/far_multicue |
FAR -- Multi-Cue | 500k | 0.56 GB |
| AI2-THOR-dyn | ai2thor_dyn/temporal |
Temporal | 500k | 0.49 GB |
| AI2-THOR-dyn | ai2thor_dyn/worldmem |
WorldMem | 500k | 0.49 GB |
| AI2-THOR v3 | ai2thor_v3/far_multicue |
FAR -- Multi-Cue | 700k | 0.55 GB |
| AI2-THOR v3 | ai2thor_v3/temporal |
Temporal | 700k | 0.49 GB |
| AI2-THOR v3 | ai2thor_v3/worldmem |
WorldMem | 700k | 0.49 GB |
| LoopNav | loopnav/far_meta |
FAR -- Meta | 700k | 0.55 GB |
| LoopNav | loopnav/far_multicue |
FAR -- Multi-Cue | 700k | 0.55 GB |
| LoopNav | loopnav/far_visual |
FAR -- Visual | 700k | 0.55 GB |
| LoopNav | loopnav/longlive_rag |
LongLive-RAG | 700k | 0.55 GB |
| LoopNav | loopnav/temporal |
Temporal | 700k | 0.49 GB |
| LoopNav | loopnav/worldmem |
WorldMem | 700k | 0.49 GB |
| SoundSpaces v1 | soundspaces_v1/far_meta |
FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v1 | soundspaces_v1/far_multicue |
FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v1 | soundspaces_v1/temporal |
Temporal | 700k | 0.49 GB |
| SoundSpaces v1 | soundspaces_v1/worldmem |
WorldMem | 700k | 0.49 GB |
| SoundSpaces v2 | soundspaces_v2/far_meta |
FAR -- Meta | 700k | 0.55 GB |
| SoundSpaces v2 | soundspaces_v2/far_multicue |
FAR -- Multi-Cue | 700k | 0.55 GB |
| SoundSpaces v2 | soundspaces_v2/temporal |
Temporal | 700k | 0.49 GB |
| SoundSpaces v2 | soundspaces_v2/worldmem |
WorldMem | 700k | 0.49 GB |
"Temporal" and "WorldMem" are the recency and field-of-view retrieval baselines, "LongLive-RAG" the content-query retrieval baseline; the FAR arms differ in the cues the retriever fuses (metadata, visual, multi-cue).
License
CC BY-NC 4.0. The generator and diffusion code these weights belong to are adapted from Navigation World Models and DiT (Meta Platforms, CC BY-NC 4.0); see the code repository's THIRD_PARTY_NOTICES.md.
Citation
@article{kim2026far,
title = {Learning What to Recall: Adaptive Multi-Cue Episodic Memory for World Models},
author = {Kim, Beomsu and Lai, Chieh-Hsin and Nguyen, Bac and Bar, Amir and Ye, Jong Chul and Mitsufuji, Yuki},
journal = {arXiv preprint arXiv:2609.34677},
year = {2026},
url = {https://arxiv.org/abs/2609.34677}
}