Visual Navigation Model Checkpoints

Pretrained navigation policies trained with VisNavKit. Each folder holds one model: the Lightning checkpoint, its ONNX export, the export metadata (shapes, anchor times, SHA-256 of both files, PyTorch/ONNX parity) and a sample input batch.

Folder VisNavKit config Model Params Val top-1 ADE@1/2/4 s (m) Val top-1 FDE (m)
flowpilot-dst-small experiment=flowpilot_dst_clips1k FastViT-T12 on frame pairs -> anchored flow DiT (256-d) 21.4M 0.093 / 0.193 / 0.422 0.897
flowpilot-dst-dune experiment=flowpilot_dune_dst_clips1k frozen DUNE ViT-B/14 -> anchored flow DiT (1024-d) 208.2M 0.080 / 0.168 / 0.364 0.769

Both are FlowPilot-DST policies trained on clips1k (20 Hz, 4 s horizon, point goal with 50% dropout, frozen route VAE, 64 k-means anchors, 4 flow steps). Metrics are on the VisNavKit clips1k validation split, decoded from zero noise. ONNX: fp32, opset 17, batch 1, top-6 plans; PyTorch/ONNX parity is within 1e-5.

Files

<folder>/flowpilot_dst_<encoder>.ckpt           # Lightning checkpoint (weights + optimizer state)
<folder>/flowpilot_dst_<encoder>.onnx           # deployment graph
<folder>/flowpilot_dst_<encoder>.metadata.json  # shapes, anchor times, hashes, parity
<folder>/flowpilot_dst_<encoder>.inputs.npz     # traced sample inputs for a smoke run

Usage

ONNX, no VisNavKit needed:

import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download

repo, stem = "UCLA-VAIL/Visual-Navigation-Model-Checkpoints", "flowpilot-dst-small/flowpilot_dst_fastvit_t12"
session = ort.InferenceSession(hf_hub_download(repo, f"{stem}.onnx"), providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
feeds = dict(np.load(hf_hub_download(repo, f"{stem}.inputs.npz")))  # replace with live data
modes, probs, speed = session.run(["modes", "probs", "speed"], feeds)
x, y, yaw, v, w = modes[0, 0].T  # best plan: 80 steps at 0.05 s, ego frame (x forward, y left)

Inputs are the last 20 frames (1, 20, 3, 216, 384) RGB in [0, 1], route patches, goal, ego speed/yaw rate and action bounds. The full contract, including frame preparation, is in FlowPilot-DST ONNX IO.

Re-export from the checkpoint with VisNavKit:

uv run visnavkit-export-dst checkpoint=flowpilot_dst_fastvit_t12.ckpt output=flowpilot_dst_fastvit_t12.onnx

Citation

@Misc{visnavkit2026,
  author       = {Honglin He and Bolei Zhou},
  title        = {{VisNavKit}: a composable toolkit for visual navigation policies},
  howpublished = {\url{https://github.com/VAIL-UCLA/visnavkit}},
  year         = {2026},
}

FlowPilot: arXiv:2606.12603. The DUNE variant uses the DUNE ViT-B/14 encoder.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Paper for UCLA-VAIL/Visual-Navigation-Model-Checkpoints