Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B: direct link, hf CLI and curl.
- Browser
- Download file 2.41 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/resolve/main/README.md
2.41 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: transformers | |
| tags: | |
| - vision-language-navigation | |
| - r2r | |
| - rxr | |
| - simplememvln | |
| # SimpleMemVLN joint R2R + RxR_15deg FullContext+B | |
| **Not yet evaluated in Habitat.** Training loss is not a navigation success metric. | |
| ## Published snapshots | |
| `epoch-1/` is the mid-schedule snapshot at update 3852 of one two-epoch | |
| cosine run. `epoch-2/` is the completed second epoch at update 7704. | |
| These are the native LLM text-output policy (option B), not candidate logits. | |
| Both epochs are retained. This joint R2R + RxR_15deg model is not yet | |
| evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839. | |
| ## Recipe and provenance | |
| Joint R2R training data and English-guide RxR_15deg trajectories, full episodes. | |
| Initialized from Qwen/Qwen3.5-4B revision | |
| `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. | |
| Source base commit: `6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6`, with input-staging | |
| and checkpoint-recovery patches. The tested source is identified by | |
| `source.sha256`; the dataset manifest by `manifest.sha256`. | |
| Serializer: `vln_append_only_chat_v1`. Native LLM text-action head (option B), | |
| gold action history; per-action mean token cross-entropy including assistant | |
| terminator, normalized over actions in the distributed accumulation window. | |
| Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2, | |
| nonreentrant gradient checkpointing with long-sequence activation offload. | |
| Full-context causal attention, no trajectory truncation or TBPTT. | |
| Global batch 8 = four H100 GPUs × one episode/rank × GAS2. | |
| Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, | |
| cosine decay to 10% of peak, weight decay 0.01, seed 429. | |
| Observation-before-action alignment is declared but not independently | |
| verified against collector source or Habitat replay. | |
| ## Loading and integrity | |
| Use the SimpleMemVLN wrapper loader | |
| `qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)` | |
| with the pinned base snapshot. These are navigation-wrapper weights, not a | |
| plain AutoModel state dictionary. Navigation metadata records the recipe. | |
| Only model, processor, tokenizer, provenance and selected training metadata | |
| are published. Optimizer/RNG state, dataset images and credentials stay local. | |
| `SHA256SUMS.json` records content hashes. The uploader downloads every file | |
| at its immutable Hub commit and verifies SHA256 before reporting success. | |