Instructions to use anhdao69/SimpleMemVLN-R2R-Window8-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-Window8-B with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-Window8-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from anhdao69/SimpleMemVLN-R2R-Window8-B: direct link, hf CLI and curl.
- Browser
- Download file 2.65 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-Window8-B/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-Window8-B/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-Window8-B/resolve/main/README.md
2.65 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: transformers | |
| tags: | |
| - vision-language-navigation | |
| - r2r | |
| - simplememvln | |
| - intermediate-checkpoint | |
| # SimpleMemVLN R2R Window8+B — partial epoch-1 snapshot | |
| **This is NOT a completed three-epoch model.** Only `epoch-1/`, saved at | |
| optimizer step **1353**, is published. The run later stopped around step 2184 | |
| of 4059 (about epoch 1.61); those later unsaved updates are not in these weights. | |
| The failure cause has not been established. No epoch-2 or epoch-3 model exists | |
| in this publication. **Not yet evaluated in Habitat.** | |
| ## Training recipe and provenance | |
| Initialized independently from Qwen/Qwen3.5-4B revision | |
| `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, not from FullContext weights. | |
| Source commit: `d50174daf1fda692849cfaa3e69286a504339fb2`. Serializer: `vln_append_only_chat_v1`. | |
| Option B uses the native LLM text-action head with action-history tokens; | |
| token losses are averaged within each action, including assistant terminator, | |
| then normalized over valid actions in the distributed accumulation window. | |
| Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2, | |
| full-episode nonreentrant gradient checkpointing, no truncation or TBPTT. | |
| Global batch 8 episodes = four H100 GPUs × one episode × GAS2. | |
| Peak LR 5e-6, weight decay 0.01, seed 429, 122 warmup updates. | |
| The intended schedule was 3 epochs / 4059 updates with cosine decay to 10% | |
| of peak LR. **Epoch 1 is a mid-schedule snapshot of that three-epoch run, | |
| not a separately scheduled one-epoch training run.** | |
| R2R training corpus: 10,819 unique episodes and 631,244 actions, with one | |
| repeated episode per distributed epoch. Capture-before-action alignment | |
| remains declared but not independently verified from collector source/replay. | |
| ## Memory policy | |
| Full-attention layers retain the instruction prefix plus the current step | |
| and previous seven complete steps. Persistent GDN recurrent state is not | |
| reset when old KV tokens are evicted. The model is trained on full episodes. | |
| ## Loading and integrity | |
| Use the SimpleMemVLN wrapper loader | |
| `qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)` | |
| with the pinned base snapshot and recorded navigation configuration. These | |
| are wrapper weights, not a plain AutoModel state dictionary. | |
| Only model/processor/navigation artifacts and selected training metadata are | |
| published; optimizer state, dataset images and credentials remain local. | |
| `SHA256SUMS.json` records content hashes. Publication is marked verified | |
| locally only after every uploaded file is downloaded at the immutable Hub | |
| commit and its SHA256 matches, including the checksum manifest itself. | |