Instructions to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B with Transformers:
# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Download README.md from anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B: direct link, hf CLI and curl.
- Browser
- Download file 2.41 kB
-
https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/resolve/main/README.md
- Command line
-
hf download hf://anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/README.md
-
curl -L -o README.md https://huggingface.co/anhdao69/SimpleMemVLN-R2R-RxR15deg-FullContext-B/resolve/main/README.md
base_model: Qwen/Qwen3.5-4B
library_name: transformers
tags:
- vision-language-navigation
- r2r
- rxr
- simplememvln
SimpleMemVLN joint R2R + RxR_15deg FullContext+B
Not yet evaluated in Habitat. Training loss is not a navigation success metric.
Published snapshots
epoch-1/ is the mid-schedule snapshot at update 3852 of one two-epoch
cosine run. epoch-2/ is the completed second epoch at update 7704.
These are the native LLM text-output policy (option B), not candidate logits.
Both epochs are retained. This joint R2R + RxR_15deg model is not yet
evaluated in Habitat. Epoch-2 action-weighted training loss: 0.0410839.
Recipe and provenance
Joint R2R training data and English-guide RxR_15deg trajectories, full episodes.
Initialized from Qwen/Qwen3.5-4B revision
851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a.
Source base commit: 6bb2b10a2fe21cb6e25a265afc5a2af39b7de4f6, with input-staging
and checkpoint-recovery patches. The tested source is identified by
source.sha256; the dataset manifest by manifest.sha256.
Serializer: vln_append_only_chat_v1. Native LLM text-action head (option B),
gold action history; per-action mean token cross-entropy including assistant
terminator, normalized over actions in the distributed accumulation window.
Frozen vision/merger, trainable text backbone and LLM head; BF16, ZeRO-2,
nonreentrant gradient checkpointing with long-sequence activation offload.
Full-context causal attention, no trajectory truncation or TBPTT.
Global batch 8 = four H100 GPUs × one episode/rank × GAS2.
Two epochs, 7704 optimization updates, 232 warmup updates, peak LR 5e-6,
cosine decay to 10% of peak, weight decay 0.01, seed 429.
Observation-before-action alignment is declared but not independently
verified against collector source or Habitat replay.
Loading and integrity
Use the SimpleMemVLN wrapper loader
qwen_vl.train.vln_runtime.load_checkpoint(epoch_directory, base_model_path)
with the pinned base snapshot. These are navigation-wrapper weights, not a
plain AutoModel state dictionary. Navigation metadata records the recipe.
Only model, processor, tokenizer, provenance and selected training metadata
are published. Optimizer/RNG state, dataset images and credentials stay local.
SHA256SUMS.json records content hashes. The uploader downloads every file
at its immutable Hub commit and verifies SHA256 before reporting success.