|
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-CandidateLogits: direct link, hf CLI and curl.
- Browser
- Download file 2.57 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-CandidateLogits/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-CandidateLogits/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-CandidateLogits/resolve/main/README.md
2.57 kB
| base_model: Qwen/Qwen3.5-4B | |
| tags: | |
| - vision-language-navigation | |
| - simplememvln | |
| - r2r | |
| - rxr | |
| # SimpleMemVLN Window8 Candidate Logits with Action History | |
| **Not yet evaluated in Habitat.** Training loss is not navigation success. | |
| Job **4649** completed successfully after two joint R2R + RxR epochs. | |
| `epoch-2/` contains the final epoch checkpoint at optimizer update **7704**. | |
| `epoch-1/` preserves the earlier mid-schedule checkpoint at update **3852**. | |
| Epoch-2 action-weighted training loss: **0.1146016239**; peak GPU reserved | |
| memory: **76.630859375 GiB**. These are training metrics, not evaluation results. | |
| ## Architecture | |
| Four-way candidate-logit readout from selected pretrained, trainable Qwen LM-head | |
| rows A/B/C/D, mapping to MOVE_FORWARD / TURN_LEFT / TURN_RIGHT / STOP. | |
| No autoregressive action generation or randomly initialized action classifier. | |
| Selected candidate tokens are explicitly appended to streaming action history; | |
| training uses teacher-forced candidate history. Serializer: **vln_candidate_logits_v1**. | |
| Window8 retains the fixed instruction prefix plus eight observation/action groups | |
| including the current observation. Action history is evicted with its step group. | |
| Persistent GDN recurrent state and advancing logical/MRoPE positions remain intact. | |
| Frozen vision encoder; trainable language backbone and tied LM-head rows. | |
| ## Recipe and provenance | |
| Joint 30,815 episodes: 10,819 R2R + 19,996 English RxR_15deg; 3,128,624 actions. | |
| Pretrained Qwen/Qwen3.5-4B revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. | |
| Global episode batch 8 = 4 H100 GPUs x 1 episode/rank x accumulation 2. | |
| Two epochs, 7704 updates, 232 warmup updates, backbone LR 5e-6, | |
| weight decay 0.01, seed 429, four-way CE, class weighting none. | |
| Source commit: `8a6a28acabd43d378e610345b2ebbe24547c37f8`. | |
| Staged-source hash inventory and submission recipe are included under provenance. | |
| ## Loading and integrity | |
| Use the matching SimpleMemVLN navigation-wrapper loader: | |
| ```python | |
| from qwen_vl.train.vln_runtime import load_checkpoint | |
| model, serializer = load_checkpoint("downloaded_repo/epoch-2", pinned_base_model_path) | |
| ``` | |
| These are navigation-wrapper weights, not a plain AutoModel state dictionary. | |
| The pinned base model/processor is required. This is neither the qwen_text policy | |
| nor the no-action-history candidate policy; honor navigation metadata on load. | |
| Optimizer tensors, RNG state, dataset content and credentials are excluded. | |
| SHA256SUMS.json records file hashes. Every published file is downloaded from | |
| the immutable publication commit and SHA256-verified. | |