Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 2,406 Bytes
caee002 a683b5d caee002 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 | ---
base_model: Qwen/Qwen3.5-4B
library_name: transformers
tags:
- vision-language-navigation
- r2r
- rxr
- simplememvln
---
# SimpleMemVLN joint R2R + RxR_15deg FullContext+B
**Not yet evaluated in Habitat.** Training loss is not a navigation success metric.
## Published snapshots
`epoch-1/` is the mid-schedule snapshot at update 3852 of one two-epoch
cosine run. `epoch-2/` is the completed second epoch at update 7704.
These are the native LLM text-output policy (option B), not candidate logits.
Both epochs are retained. This joint R2R + RxR_15deg model is not yet
evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839.
## Recipe and provenance
Joint R2R training data and English-guide RxR_15deg trajectories, full episodes.
Initialized from Qwen/Qwen3.5-4B revision
`851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a`.
Source base commit: `6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6`, with input-staging
and checkpoint-recovery patches. The tested source is identified by
`source.sha256`; the dataset manifest by `manifest.sha256`.
Serializer: `vln_append_only_chat_v1`. Native LLM text-action head (option B),
gold action history; per-action mean token cross-entropy including assistant
terminator, normalized over actions in the distributed accumulation window.
Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2,
nonreentrant gradient checkpointing with long-sequence activation offload.
Full-context causal attention, no trajectory truncation or TBPTT.
Global batch 8 = four H100 GPUs × one episode/rank × GAS2.
Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6,
cosine decay to 10% of peak, weight decay 0.01, seed 429.
Observation-before-action alignment is declared but not independently
verified against collector source or Habitat replay.
## Loading and integrity
Use the SimpleMemVLN wrapper loader
`qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)`
with the pinned base snapshot. These are navigation-wrapper weights, not a
plain AutoModel state dictionary. Navigation metadata records the recipe.
Only model, processor, tokenizer, provenance and selected training metadata
are published. Optimizer/RNG state, dataset images and credentials stay local.
`SHA256SUMS.json` records content hashes. The uploader downloads every file
at its immutable Hub commit and verifies SHA256 before reporting success.
|