--- base_model: Qwen/Qwen3.5-4B tags: - vision-language-navigation - simplememvln - r2r - rxr --- # SimpleMemVLN Window8 Candidate Logits with Action History **Not yet evaluated in Habitat.** Training loss is not navigation success. Job **4649** completed successfully after two joint R2R + RxR epochs. `epoch-2/` contains the final epoch checkpoint at optimizer update **7704**. `epoch-1/` preserves the earlier mid-schedule checkpoint at update **3852**. Epoch-2 action-weighted training loss: **0.1146016239**; peak GPU reserved memory: **76.630859375 GiB**. These are training metrics, not evaluation results. ## Architecture Four-way candidate-logit readout from selected pretrained, trainable Qwen LM-head rows A/B/C/D, mapping to MOVE_FORWARD / TURN_LEFT / TURN_RIGHT / STOP. No autoregressive action generation or randomly initialized action classifier. Selected candidate tokens are explicitly appended to streaming action history; training uses teacher-forced candidate history. Serializer: **vln_candidate_logits_v1**. Window8 retains the fixed instruction prefix plus eight observation/action groups including the current observation. Action history is evicted with its step group. Persistent GDN recurrent state and advancing logical/MRoPE positions remain intact. Frozen vision encoder; trainable language backbone and tied LM-head rows. ## Recipe and provenance Joint 30,815 episodes: 10,819 R2R + 19,996 English RxR_15deg; 3,128,624 actions. Pretrained Qwen/Qwen3.5-4B revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. Global episode batch 8 = 4 H100 GPUs x 1 episode/rank x accumulation 2. Two epochs, 7704 updates, 232 warmup updates, backbone LR 5e-6, weight decay 0.01, seed 429, four-way CE, class weighting none. Source commit: `8a6a28acabd43d378e610345b2ebbe24547c37f8`. Staged-source hash inventory and submission recipe are included under provenance. ## Loading and integrity Use the matching SimpleMemVLN navigation-wrapper loader: ```python from qwen_vl.train.vln_runtime import load_checkpoint model, serializer = load_checkpoint("downloaded_repo/epoch-2", pinned_base_model_path) ``` These are navigation-wrapper weights, not a plain AutoModel state dictionary. The pinned base model/processor is required. This is neither the qwen_text policy nor the no-action-history candidate policy; honor navigation metadata on load. Optimizer tensors, RNG state, dataset content and credentials are excluded. SHA256SUMS.json records file hashes. Every published file is downloaded from the immutable publication commit and SHA256-verified.