Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B: direct link, hf CLI and curl.
- Browser
- Download file 3.29 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B/resolve/main/README.md
3.29 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: transformers | |
| tags: | |
| - vision-language-navigation | |
| - r2r | |
| - rxr | |
| - simplememvln | |
| # SimpleMemVLN joint R2R + RxR_15deg Window8+B | |
| **Not yet evaluated in Habitat.** Training loss is not a navigation success metric. | |
| ## Published snapshots | |
| Job **4643 completed successfully**. `epoch-2/` contains the completed second | |
| epoch at optimizer update **7704**. The earlier checkpoint at update 3852 | |
| remains available under `epoch-1/`. | |
| Epoch-2 action-weighted training loss: **0.0907004696**. | |
| ## Recipe | |
| One joint model on 10,819 R2R and 19,996 English-guide RxR_15deg training episodes, | |
| 30,815 unique episodes total. Complete observation-before-action trajectories; | |
| no trajectory truncation or TBPTT. Initialized from Qwen/Qwen3.5-4B revision | |
| `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. | |
| Serializer: **vln_append_only_chat_v1**. Pretrained LLM text-action head (option B). | |
| Teacher-forced canonical action history during training, generated action history | |
| in streaming inference. Loss is token CE averaged within each action, then normalized | |
| over all actions across ranks and gradient accumulation. | |
| Window8 keeps the instruction prefix plus eight complete observation/action groups | |
| including the current observation. Old full-attention KV groups are evicted together; | |
| GDN recurrent memory persists, and logical/MRoPE positions continue advancing. | |
| The vision encoder/merger is frozen. Text backbone and tied LLM head are trainable. | |
| BF16, FlashAttention2 step attention, GDN, non-reentrant activation checkpointing, | |
| long-sequence activation offload, and DeepSpeed ZeRO2 with **CPU optimizer offload**. | |
| Global batch **8 = 2 H100 GPUs x 1 episode/rank x GAS4**. | |
| Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, | |
| cosine decay to 10% of peak, weight decay 0.01, seed 429. | |
| Four data-loader workers per rank, OMP_NUM_THREADS=2. | |
| ## Provenance | |
| Training source is the existing production Window8 text-policy snapshot with the | |
| two-rank reporting correction from commit **33a2735**, not a candidate-policy | |
| checkpoint. `provenance/source.sha256` identifies the exact staged source; | |
| `provenance/runtime.sha256` identifies launch/runtime artifacts, and | |
| `provenance/manifest.sha256` identifies the dataset manifest. | |
| The CPU optimizer configuration is included in `provenance/deepspeed.json`. | |
| Both epoch training reports are included under `reports/`. | |
| ## Loading and integrity | |
| Download this repository and pass its `epoch-2` directory to the SimpleMemVLN loader: | |
| ```python | |
| from qwen_vl.train.vln_runtime import load_checkpoint | |
| model, serializer = load_checkpoint("downloaded_repo/epoch-2", pinned_base_model_path) | |
| ``` | |
| Use the matching SimpleMemVLN source and validated Transformers 5.11.0 environment. | |
| These are navigation-wrapper weights, not a plain AutoModel state dictionary. | |
| The pinned base model/processor is required by the wrapper loader. | |
| Only inference weights, tokenizer/processor, navigation metadata, training-state | |
| summary, provenance and epoch report are published. Optimizer tensors, RNG state, | |
| dataset frames, episode manifest content and credentials remain local. | |
| `SHA256SUMS.json` records content hashes. Every uploaded file is downloaded from | |
| the immutable publication commit and SHA256-verified before success is reported. | |