Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits", device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits: direct link, hf CLI and curl.
- Browser
- Download file 2.19 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-CandidateLogits/resolve/main/README.md
2.19 kB
| base_model: Qwen/Qwen3.5-4B | |
| library_name: transformers | |
| tags: | |
| - vision-language-navigation | |
| - r2r | |
| - rxr | |
| - simplememvln | |
| - candidate-logits | |
| # SimpleMemVLN R2R + RxR_15deg FullContext candidate-logits | |
| **Not yet evaluated in Habitat.** Training loss is not navigation success. | |
| This is NOT the qwen_text FullContext-B checkpoint. | |
| ## Snapshot and recipe | |
| `epoch-1/` is the mid-schedule snapshot at step 3852. | |
| `epoch-2/` is the completed two-epoch model at step 7704. Both are retained. | |
| Joint R2R and English-guide RxR_15deg full episodes. Global batch 8 = 4 H100 | |
| GPUs x one episode per rank x gradient accumulation 2. LR 5e-6, warmup 232, | |
| cosine decay to 10% of peak, weight decay 0.01, seed 429. BF16, ZeRO-2, | |
| gradient checkpointing and long-sequence activation offload; vision frozen. | |
| Initialized from Qwen/Qwen3.5-4B revision | |
| `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. | |
| Training source commit: `8a6a28acabd43d378e610345b2ebbe24547c37f8` on | |
| SimpleMemVLN `streaming_logits`; exact source and dataset hashes are included. | |
| ## Navigation contract | |
| Serializer `vln_candidate_logits_v1`, full-context causal attention with | |
| persistent GDN state. Four selected pretrained LM-head rows, trainable at the | |
| backbone LR (`lm_rows_trainable`); no random classifier and no autoregressive | |
| action generation. Candidate labels A/B/C/D (IDs 32/33/34/35) map to | |
| MOVE_FORWARD/TURN_LEFT/TURN_RIGHT/STOP and Habitat IDs 1/2/3/0. | |
| Explicit candidate-token feedback; unweighted four-way cross entropy. | |
| Epoch-1 action-weighted loss: 0.3663298; epoch-2: 0.1044625. Not comparable numerically to text CE. | |
| ## Loading and integrity | |
| Use the SimpleMemVLN candidate-aware wrapper loader | |
| `qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)` | |
| from `streaming_logits`, with the pinned pretrained base snapshot. | |
| These are navigation-wrapper weights, not plain AutoModel weights. Do not load | |
| them using the old text-policy serializer. `navigation.json` is authoritative. | |
| Model/tokenizer/processor metadata only; optimizer states, RNG, images and | |
| credentials remain local. `SHA256SUMS.json` records file hashes. Publication | |
| is verified by downloading each uploaded file at its immutable Hub revision. | |