anhdao69's picture
Publish partial Window8+B epoch-1 snapshot (1353 steps; interrupted 3-epoch run)
cd17ac0 verified
|
Raw History Blame Contribute Delete
2.65 kB
---
base_model: Qwen/Qwen3.5-4B
library_name: transformers
tags:
- vision-language-navigation
- r2r
- simplememvln
- intermediate-checkpoint
---
# SimpleMemVLN R2R Window8+B — partial epoch-1 snapshot
**This is NOT a completed three-epoch model.** Only `epoch-1/`, saved at
optimizer step **1353**, is published. The run later stopped around step 2184
of 4059 (about epoch 1.61); those later unsaved updates are not in these weights.
The failure cause has not been established. No epoch-2 or epoch-3 model exists
in this publication. **Not yet evaluated in Habitat.**
## Training recipe and provenance
Initialized independently from Qwen/Qwen3.5-4B revision
`851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`, not from FullContext weights.
Source commit: `d50174daf1fda692849cfaa3e69286a504339fb2`. Serializer: `vln_append_only_chat_v1`.
Option B uses the native LLM text-action head with action-history tokens;
token losses are averaged within each action, including assistant terminator,
then normalized over valid actions in the distributed accumulation window.
Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2,
full-episode nonreentrant gradient checkpointing, no truncation or TBPTT.
Global batch 8 episodes = four H100 GPUs × one episode × GAS2.
Peak LR 5e-6, weight decay 0.01, seed 429, 122 warmup updates.
The intended schedule was 3 epochs / 4059 updates with cosine decay to 10%
of peak LR. **Epoch 1 is a mid-schedule snapshot of that three-epoch run,
not a separately scheduled one-epoch training run.**
R2R training corpus: 10,819 unique episodes and 631,244 actions, with one
repeated episode per distributed epoch. Capture-before-action alignment
remains declared but not independently verified from collector source/replay.
## Memory policy
Full-attention layers retain the instruction prefix plus the current step
and previous seven complete steps. Persistent GDN recurrent state is not
reset when old KV tokens are evicted. The model is trained on full episodes.
## Loading and integrity
Use the SimpleMemVLN wrapper loader
`qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)`
with the pinned base snapshot and recorded navigation configuration. These
are wrapper weights, not a plain AutoModel state dictionary.
Only model/processor/navigation artifacts and selected training metadata are
published; optimizer state, dataset images and credentials remain local.
`SHA256SUMS.json` records content hashes. Publication is marked verified
locally only after every uploaded file is downloaded at the immutable Hub
commit and its SHA256 matches, including the checksum manifest itself.