File size: 2,406 Bytes
caee002
 
 
 
 
 
 
 
 
 
 
 
 
a683b5d
 
 
 
 
 
caee002
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
---
base_model: Qwen/Qwen3.5-4B
library_name: transformers
tags:
- vision-language-navigation
- r2r
- rxr
- simplememvln
---
# SimpleMemVLN joint R2R + RxR_15deg FullContext+B

**Not yet evaluated in Habitat.** Training loss is not a navigation success metric.

## Published snapshots
`epoch-1/` is the mid-schedule snapshot at update 3852 of one two-epoch
cosine run. `epoch-2/` is the completed second epoch at update 7704.
These are the native LLM text-output policy (option B), not candidate logits.
Both epochs are retained. This joint R2R + RxR_15deg model is not yet
evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839.

## Recipe and provenance
Joint R2R training data and English-guide RxR_15deg trajectories, full episodes.
Initialized from Qwen/Qwen3.5-4B revision
`851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`.
Source base commit: `6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6`, with input-staging
and checkpoint-recovery patches. The tested source is identified by
`source.sha256`; the dataset manifest by `manifest.sha256`.
Serializer: `vln_append_only_chat_v1`. Native LLM text-action head (option B),
gold action history; per-action mean token cross-entropy including assistant
terminator, normalized over actions in the distributed accumulation window.
Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2,
nonreentrant gradient checkpointing with long-sequence activation offload.
Full-context causal attention, no trajectory truncation or TBPTT.
Global batch 8 = four H100 GPUs × one episode/rank × GAS2.
Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6,
cosine decay to 10% of peak, weight decay 0.01, seed 429.
Observation-before-action alignment is declared but not independently
verified against collector source or Habitat replay.

## Loading and integrity
Use the SimpleMemVLN wrapper loader
`qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)`
with the pinned base snapshot. These are navigation-wrapper weights, not a
plain AutoModel state dictionary. Navigation metadata records the recipe.
Only model, processor, tokenizer, provenance and selected training metadata
are published. Optimizer/RNG state, dataset images and credentials stay local.
`SHA256SUMS.json` records content hashes. The uploader downloads every file
at its immutable Hub commit and verifies SHA256 before reporting success.