--- base_model: Qwen/Qwen3.5-4B library_name: transformers tags: - vision-language-navigation - r2r - rxr - simplememvln --- # SimpleMemVLN joint R2R + RxR_15deg Window8+B **Not yet evaluated in Habitat.** Training loss is not a navigation success metric. ## Published snapshots Job **4643 completed successfully**. `epoch-2/` contains the completed second epoch at optimizer update **7704**. The earlier checkpoint at update 3852 remains available under `epoch-1/`. Epoch-2 action-weighted training loss: **0.0907004696**. ## Recipe One joint model on 10,819 R2R and 19,996 English-guide RxR_15deg training episodes, 30,815 unique episodes total. Complete observation-before-action trajectories; no trajectory truncation or TBPTT. Initialized from Qwen/Qwen3.5-4B revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. Serializer: **vln_append_only_chat_v1**. Pretrained LLM text-action head (option B). Teacher-forced canonical action history during training, generated action history in streaming inference. Loss is token CE averaged within each action, then normalized over all actions across ranks and gradient accumulation. Window8 keeps the instruction prefix plus eight complete observation/action groups including the current observation. Old full-attention KV groups are evicted together; GDN recurrent memory persists, and logical/MRoPE positions continue advancing. The vision encoder/merger is frozen. Text backbone and tied LLM head are trainable. BF16, FlashAttention2 step attention, GDN, non-reentrant activation checkpointing, long-sequence activation offload, and DeepSpeed ZeRO2 with **CPU optimizer offload**. Global batch **8 = 2 H100 GPUs x 1 episode/rank x GAS4**. Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, cosine decay to 10% of peak, weight decay 0.01, seed 429. Four data-loader workers per rank, OMP_NUM_THREADS=2. ## Provenance Training source is the existing production Window8 text-policy snapshot with the two-rank reporting correction from commit **33a2735**, not a candidate-policy checkpoint. `provenance/source.sha256` identifies the exact staged source; `provenance/runtime.sha256` identifies launch/runtime artifacts, and `provenance/manifest.sha256` identifies the dataset manifest. The CPU optimizer configuration is included in `provenance/deepspeed.json`. Both epoch training reports are included under `reports/`. ## Loading and integrity Download this repository and pass its `epoch-2` directory to the SimpleMemVLN loader: ```python from qwen_vl.train.vln_runtime import load_checkpoint model, serializer = load_checkpoint("downloaded_repo/epoch-2", pinned_base_model_path) ``` Use the matching SimpleMemVLN source and validated Transformers 5.11.0 environment. These are navigation-wrapper weights, not a plain AutoModel state dictionary. The pinned base model/processor is required by the wrapper loader. Only inference weights, tokenizer/processor, navigation metadata, training-state summary, provenance and epoch report are published. Optimizer tensors, RNG state, dataset frames, episode manifest content and credentials remain local. `SHA256SUMS.json` records content hashes. Every uploaded file is downloaded from the immutable publication commit and SHA256-verified before success is reported.