anhdao69's picture
Publish completed joint R2R + RxR_15deg FullContext text epoch 2 (step 7704)
a683b5d verified
|
Raw History Blame Contribute Delete
2.41 kB
metadata
base_model: Qwen/Qwen3.5-4B
library_name: transformers
tags:
  - vision-language-navigation
  - r2r
  - rxr
  - simplememvln

SimpleMemVLN joint R2R + RxR_15deg FullContext+B

Not yet evaluated in Habitat. Training loss is not a navigation success metric.

Published snapshots

epoch-1/ is the mid-schedule snapshot at update 3852 of one two-epoch cosine run. epoch-2/ is the completed second epoch at update 7704. These are the native LLM text-output policy (option B), not candidate logits. Both epochs are retained. This joint R2R + RxR_15deg model is not yet evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839.

Recipe and provenance

Joint R2R training data and English-guide RxR_15deg trajectories, full episodes. Initialized from Qwen/Qwen3.5-4B revision 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a. Source base commit: 6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6, with input-staging and checkpoint-recovery patches. The tested source is identified by source.sha256; the dataset manifest by manifest.sha256. Serializer: vln_append_only_chat_v1. Native LLM text-action head (option B), gold action history; per-action mean token cross-entropy including assistant terminator, normalized over actions in the distributed accumulation window. Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2, nonreentrant gradient checkpointing with long-sequence activation offload. Full-context causal attention, no trajectory truncation or TBPTT. Global batch 8 = four H100 GPUs × one episode/rank × GAS2. Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6, cosine decay to 10% of peak, weight decay 0.01, seed 429. Observation-before-action alignment is declared but not independently verified against collector source or Habitat replay.

Loading and integrity

Use the SimpleMemVLN wrapper loader qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path) with the pinned base snapshot. These are navigation-wrapper weights, not a plain AutoModel state dictionary. Navigation metadata records the recipe. Only model, processor, tokenizer, provenance and selected training metadata are published. Optimizer/RNG state, dataset images and credentials stay local. SHA256SUMS.json records content hashes. The uploader downloads every file at its immutable Hub commit and verifies SHA256 before reporting success.