--- base_model: Qwen/Qwen3.5-4B library_name: transformers tags: - vision-language-navigation - r2r - simplememvln - intermediate-checkpoint --- # SimpleMemVLN R2R Window8+B — partial epoch-1 snapshot **This is NOT a completed three-epoch model.** Only `epoch-1/`, saved at optimizer step **1353**, is published. The run later stopped around step 2184 of 4059 (about epoch 1.61); those later unsaved updates are not in these weights. The failure cause has not been established. No epoch-2 or epoch-3 model exists in this publication. **Not yet evaluated in Habitat.** ## Training recipe and provenance Initialized independently from Qwen/Qwen3.5-4B revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, not from FullContext weights. Source commit: `d50174daf1fda692849cfaa3e69286a504339fb2`. Serializer: `vln_append_only_chat_v1`. Option B uses the native LLM text-action head with action-history tokens; token losses are averaged within each action, including assistant terminator, then normalized over valid actions in the distributed accumulation window. Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2, full-episode nonreentrant gradient checkpointing, no truncation or TBPTT. Global batch 8 episodes = four H100 GPUs × one episode × GAS2. Peak LR 5e-6, weight decay 0.01, seed 429, 122 warmup updates. The intended schedule was 3 epochs / 4059 updates with cosine decay to 10% of peak LR. **Epoch 1 is a mid-schedule snapshot of that three-epoch run, not a separately scheduled one-epoch training run.** R2R training corpus: 10,819 unique episodes and 631,244 actions, with one repeated episode per distributed epoch. Capture-before-action alignment remains declared but not independently verified from collector source/replay. ## Memory policy Full-attention layers retain the instruction prefix plus the current step and previous seven complete steps. Persistent GDN recurrent state is not reset when old KV tokens are evicted. The model is trained on full episodes. ## Loading and integrity Use the SimpleMemVLN wrapper loader `qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)` with the pinned base snapshot and recorded navigation configuration. These are wrapper weights, not a plain AutoModel state dictionary. Only model/processor/navigation artifacts and selected training metadata are published; optimizer state, dataset images and credentials remain local. `SHA256SUMS.json` records content hashes. Publication is marked verified locally only after every uploaded file is downloaded at the immutable Hub commit and its SHA256 matches, including the checksum manifest itself.