PanoVLN

PanoVLN is a vision-and-language navigation policy that follows instructions from 360° RGB observations. It combines a Qwen3.5-4B backbone with PanoVGGT geometry features and predicts 18-action sequences. Confidence-guided execution selects how far to move before observing and planning again.

This model is presented in the paper PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation.

For usage, training, evaluation, and deployment instructions, please refer to the GitHub repository.

Downloads last month
8
Safetensors
Model size
6B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for wangzhen-w/PanoVLN