Instructions to use anhdao69/SimpleMemVLN-R2R-FullContext-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-FullContext-B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-FullContext-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from anhdao69/SimpleMemVLN-R2R-FullContext-B: direct link, hf CLI and curl.
- Browser
- Download file 2.19 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-FullContext-B/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-FullContext-B/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-FullContext-B/resolve/main/README.md
2.19 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: transformers | |
| tags: | |
| - vision-language-navigation | |
| - r2r | |
| - simplememvln | |
| # SimpleMemVLN R2R fullcontext — option B | |
| **Not yet evaluated in Habitat.** Training loss is not a navigation success metric. | |
| The observation-before-action alignment is declared but not independently verified | |
| from collector source or Habitat replay. | |
| ## Recipe | |
| Independently initialized from Qwen/Qwen3.5-4B revision | |
| `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, not from the other variant. | |
| Code commit: `d50174daf1fda692849cfaa3e69286a504339fb2`. | |
| Serializer: `vln_append_only_chat_v1`. Native LLM text-action head (option B), | |
| gold action history, per-action mean token CE including assistant terminator. | |
| Frozen vision/merger; trainable text backbone and LLM head. BF16, ZeRO-2, | |
| nonreentrant full-episode gradient checkpointing, no truncation or TBPTT. | |
| Four H100 GPUs, one episode/rank, GAS 2: global batch 8 episodes. | |
| R2R train: 10,819 unique episodes, 631,244 actions; distributed repetition | |
| adds one episode per epoch. Three epochs, 4,059 updates, 122 warmup updates. | |
| Peak LR 5e-6, cosine decay to 10% of peak, weight decay 0.01, seed 429. | |
| Memory mode: `fullcontext`. Window8 keeps prefix plus eight complete steps | |
| for full-attention KV; persistent GDN state and gradients are not reset. | |
| ## Snapshots | |
| `epoch-1`, `epoch-2`, `epoch-3` correspond to updates 1353, 2706, 4059. | |
| **Epochs 1–2 are mid-schedule snapshots of one 3-epoch cosine run**, not | |
| independently trained one-/two-epoch models. All optimizer states remain local. | |
| ## Loading and integrity | |
| These weights use the SimpleMemVLN navigation wrapper, not an unmodified | |
| `AutoModel` state dictionary. Download an epoch directory and load it with | |
| `qwen_vl.train.vln_runtime.load_checkpoint(path, model_path)` using the pinned | |
| base snapshot and the recorded environment. Navigation metadata contains the | |
| complete serializer and memory contract. | |
| `SHA256SUMS.json` records model/configuration/metric hashes. The uploader | |
| downloads every uploaded file at the immutable commit revision and verifies | |
| SHA256 before marking publication complete locally. Dataset images, credentials | |
| and optimizer states are not published. | |