SimpleMemVLN joint R2R + RxR_15deg Window8+B

Not yet evaluated in Habitat. Training loss is not a navigation success metric.

Published snapshot

epoch-1/ is the completed first epoch at optimizer update 3852, from two-GPU training job 4643. This is the qwen_text policy, not candidate logits. This is a mid-schedule snapshot of one two-epoch cosine run, not a separately scheduled one-epoch model. Epoch 2 continues separately and is not included here.

Recipe

One joint model on 10,819 R2R and 19,996 English-guide RxR_15deg training episodes, 30,815 unique episodes total. Complete observation-before-action trajectories; no trajectory truncation or TBPTT. Initialized from Qwen/Qwen3.5-4B revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.

Serializer: vln_append_only_chat_v1. Pretrained LLM text-action head (option B). Teacher-forced canonical action history during training, generated action history in streaming inference. Loss is token CE averaged within each action, then normalized over all actions across ranks and gradient accumulation.

Window8 keeps the instruction prefix plus eight complete observation/action groups including the current observation. Old full-attention KV groups are evicted together; GDN recurrent memory persists, and logical/MRoPE positions continue advancing. The vision encoder/merger is frozen. Text backbone and tied LLM head are trainable. BF16, FlashAttention2 step attention, GDN, non-reentrant activation checkpointing, long-sequence activation offload, and DeepSpeed ZeRO2 with CPU optimizer offload.

Global batch 8 = 2 H100 GPUs x 1 episode/rank x GAS4. Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, cosine decay to 10% of peak, weight decay 0.01, seed 429. Four data-loader workers per rank, OMP_NUM_THREADS=2.

Provenance

Training source is the existing production Window8 text-policy snapshot with the two-rank reporting correction from commit 33a2735, not a candidate-policy checkpoint. provenance/source.sha256 identifies the exact staged source; provenance/runtime.sha256 identifies launch/runtime artifacts, and provenance/manifest.sha256 identifies the dataset manifest. The CPU optimizer configuration is included in provenance/deepspeed.json. The epoch-1 training report is included in reports/epoch-1.json.

Loading and integrity

Download this repository and pass its epoch-1 directory to the SimpleMemVLN loader:

from qwen_vl.train.vln_runtime import load_checkpoint
model, serializer = load_checkpoint("downloaded_repo/epoch-1", pinned_base_model_path)

Use the matching SimpleMemVLN source and validated Transformers 5.11.0 environment. These are navigation-wrapper weights, not a plain AutoModel state dictionary. The pinned base model/processor is required by the wrapper loader.

Only inference weights, tokenizer/processor, navigation metadata, training-state summary, provenance and epoch report are published. Optimizer tensors, RNG state, dataset frames, episode manifest content and credentials remain local. SHA256SUMS.json records content hashes. Every uploaded file is downloaded from the immutable publication commit and SHA256-verified before success is reported.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for anhdao69/SimpleMemVLN-R2R-RxR15deg-Window8-B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(848)
this model