FlowPilot-DST streaming ONNX (one frame per call, feature buffer) for small and dune; README
4dae8a5 verified |
Download README.md from UCLA-VAIL/Visual-Navigation-Model-Checkpoints: direct link, hf CLI and curl.
- Browser
- Download file 5.3 kB
-
https://huggingface.co/UCLA-VAIL/Visual-Navigation-Model-Checkpoints/resolve/main/README.md
- Command line
-
hf download hf://UCLA-VAIL/Visual-Navigation-Model-Checkpoints/README.md
-
curl -L -o README.md https://huggingface.co/UCLA-VAIL/Visual-Navigation-Model-Checkpoints/resolve/main/README.md
5.3 kB
| license: apache-2.0 | |
| library_name: visnavkit | |
| pipeline_tag: robotics | |
| tags: | |
| - visual-navigation | |
| - flow-matching | |
| - onnx | |
| - robotics | |
| # Visual Navigation Model Checkpoints | |
| Pretrained navigation policies trained with [VisNavKit](https://github.com/VAIL-UCLA/visnavkit). | |
| Each folder holds one model: the Lightning checkpoint, its two ONNX exports (window and streaming), | |
| the export metadata (shapes, anchor times, SHA-256 of both files, PyTorch/ONNX parity) and a sample | |
| input batch. | |
| | Folder | VisNavKit config | Model | Params | Val top-1 ADE@1/2/4 s (m) | Val top-1 FDE (m) | | |
| | --- | --- | --- | --- | --- | --- | | |
| | [`flowpilot-dst-small`](flowpilot-dst-small) | `experiment=flowpilot_dst_clips1k` | FastViT-T12 on frame pairs -> anchored flow DiT (256-d) | 21.4M | 0.093 / 0.193 / 0.422 | 0.897 | | |
| | [`flowpilot-dst-dune`](flowpilot-dst-dune) | `experiment=flowpilot_dune_dst_clips1k` | frozen DUNE ViT-B/14 -> anchored flow DiT (1024-d) | 208.2M | 0.080 / 0.168 / 0.364 | 0.769 | | |
| Both are FlowPilot-DST policies trained on clips1k (20 Hz, 4 s horizon, point goal with 50% | |
| dropout, frozen route VAE, 64 k-means anchors, 4 flow steps). Metrics are on the VisNavKit | |
| clips1k validation split, decoded from zero noise. ONNX: fp32, opset 17, batch 1, top-6 plans; | |
| PyTorch/ONNX parity is within 1e-5. | |
| ## Files | |
| ``` | |
| <folder>/flowpilot_dst_<encoder>.ckpt # Lightning checkpoint (weights + optimizer state) | |
| <folder>/flowpilot_dst_<encoder>.onnx # deployment graph | |
| <folder>/flowpilot_dst_<encoder>.metadata.json # shapes, anchor times, hashes, parity | |
| <folder>/flowpilot_dst_<encoder>.inputs.npz # traced sample inputs for a smoke run | |
| <folder>/flowpilot_dst_<encoder>_streaming.{onnx,metadata.json,inputs.npz} # streaming graph, same weights | |
| ``` | |
| ## Usage | |
| ONNX, no VisNavKit needed: | |
| ```python | |
| import numpy as np, onnxruntime as ort | |
| from huggingface_hub import hf_hub_download | |
| repo, stem = "UCLA-VAIL/Visual-Navigation-Model-Checkpoints", "flowpilot-dst-small/flowpilot_dst_fastvit_t12" | |
| session = ort.InferenceSession(hf_hub_download(repo, f"{stem}.onnx"), providers=["CUDAExecutionProvider", "CPUExecutionProvider"]) | |
| feeds = dict(np.load(hf_hub_download(repo, f"{stem}.inputs.npz"))) # replace with live data | |
| modes, probs, speed = session.run(["modes", "probs", "speed"], feeds) | |
| x, y, yaw, v, w = modes[0, 0].T # best plan: 80 steps at 0.05 s, ego frame (x forward, y left) | |
| ``` | |
| Inputs are the last 20 frames (1, 20, 3, 216, 384) RGB in [0, 1], route patches, goal, ego | |
| speed/yaw rate and action bounds. The full contract, including frame preparation, is in | |
| [FlowPilot-DST ONNX IO](https://github.com/VAIL-UCLA/visnavkit/blob/dev/docs/flowpilot_dst_onnx.md). | |
| ### Streaming ONNX | |
| `*_streaming.onnx` runs one frame per call. It caches each past frame's features (`[global image | |
| feature | route latent]`) in a buffer, encodes only the current frame, and decodes only the current | |
| plan. After 20 frames its decision equals the window graph's (checked at export: DST exact, DUNE | |
| to 5e-7 in ONNX Runtime). | |
| | input | shape | meaning | | |
| | --- | --- | --- | | |
| | `vision` | small (1, 2, 3, 216, 384): [previous, current]; dune (1, 1, 3, 216, 384) | the current frame; for small, the previous frame is all zeros on the first call | | |
| | `route_patch`, `goal`, `ego` | (1, 1, 80, 80), (1, 1, 3), (1, 1, 2) | the current frame's route, goal, `[v, w]` | | |
| | `action_bounds` | (1, 2, 5) | as in the window graph | | |
| | `feat_buffer` | (1, 19, F): F = 768 small, 2048 dune | the last call's `feat_buffer_out`; zeros at startup | | |
| | `buffer_mask` | (1, 19) | the last call's `buffer_mask_out`; zeros at startup | | |
| Outputs: `modes`, `probs`, `speed` as in the window graph, plus `feat_buffer_out`, `buffer_mask_out`. | |
| ```python | |
| stem = "flowpilot-dst-dune/flowpilot_dst_dune_vitb14_streaming" | |
| session = ort.InferenceSession(hf_hub_download(repo, f"{stem}.onnx")) | |
| buf, mask = np.zeros((1, 19, 2048), np.float32), np.zeros((1, 19), np.float32) | |
| for frame, route, goal, ego in stream: # 20 Hz, frames exactly 50 ms apart | |
| modes, probs, speed, buf, mask = session.run(None, dict( | |
| vision=frame[None, None], route_patch=route[None, None], goal=goal[None, None], | |
| ego=ego[None, None], action_bounds=bounds, feat_buffer=buf, buffer_mask=mask)) | |
| plan = modes[0, 0] # (80, 5) | |
| ``` | |
| - Feed `feat_buffer_out` and `buffer_mask_out` back unchanged. | |
| - Call once per frame at 20 Hz. Don't skip frames. Reset the buffer (and, for small, the previous | |
| frame) to zeros after a gap or restart. | |
| - Keep one buffer per camera stream. | |
| - CPU latency per call (ONNX Runtime, loaded machine): small 1248 ms window vs 381 ms streaming; | |
| dune 7537 ms vs 2021 ms. | |
| Re-export from the checkpoint with VisNavKit (add `streaming=true` for the streaming graph): | |
| ```bash | |
| uv run visnavkit-export-dst checkpoint=flowpilot_dst_fastvit_t12.ckpt output=flowpilot_dst_fastvit_t12.onnx | |
| ``` | |
| ## Citation | |
| ```bibtex | |
| @Misc{visnavkit2026, | |
| author = {Honglin He and Bolei Zhou}, | |
| title = {{VisNavKit}: a composable toolkit for visual navigation policies}, | |
| howpublished = {\url{https://github.com/VAIL-UCLA/visnavkit}}, | |
| year = {2026}, | |
| } | |
| ``` | |
| FlowPilot: [arXiv:2606.12603](https://arxiv.org/abs/2606.12603). The DUNE variant uses the | |
| [DUNE](https://github.com/naver/dune) ViT-B/14 encoder. | |