--- base_model: Qwen/Qwen3.5-4B library_name: transformers tags: - vision-language-navigation - r2r - rxr - simplememvln --- # SimpleMemVLN joint R2R + RxR_15deg FullContext+B **Not yet evaluated in Habitat.** Training loss is not a navigation success metric. ## Published snapshots `epoch-1/` is the mid-schedule snapshot at update 3852 of one two-epoch cosine run. `epoch-2/` is the completed second epoch at update 7704. These are the native LLM text-output policy (option B), not candidate logits. Both epochs are retained. This joint R2R + RxR_15deg model is not yet evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839. ## Recipe and provenance Joint R2R training data and English-guide RxR_15deg trajectories, full episodes. Initialized from Qwen/Qwen3.5-4B revision `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`. Source base commit: `6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6`, with input-staging and checkpoint-recovery patches. The tested source is identified by `source.sha256`; the dataset manifest by `manifest.sha256`. Serializer: `vln_append_only_chat_v1`. Native LLM text-action head (option B), gold action history; per-action mean token cross-entropy including assistant terminator, normalized over actions in the distributed accumulation window. Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2, nonreentrant gradient checkpointing with long-sequence activation offload. Full-context causal attention, no trajectory truncation or TBPTT. Global batch 8 = four H100 GPUs × one episode/rank × GAS2. Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, cosine decay to 10% of peak, weight decay 0.01, seed 429. Observation-before-action alignment is declared but not independently verified against collector source or Habitat replay. ## Loading and integrity Use the SimpleMemVLN wrapper loader `qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)` with the pinned base snapshot. These are navigation-wrapper weights, not a plain AutoModel state dictionary. Navigation metadata records the recipe. Only model, processor, tokenizer, provenance and selected training metadata are published. Optimizer/RNG state, dataset images and credentials stay local. `SHA256SUMS.json` records content hashes. The uploader downloads every file at its immutable Hub commit and verifies SHA256 before reporting success.