File size: 5,297 Bytes
ae72411
 
bc1094d
 
 
 
 
 
 
ae72411
bc1094d
 
 
 
4dae8a5
 
 
bc1094d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4dae8a5
bc1094d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4dae8a5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
bc1094d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
---
license: apache-2.0
library_name: visnavkit
pipeline_tag: robotics
tags:
  - visual-navigation
  - flow-matching
  - onnx
  - robotics
---

# Visual Navigation Model Checkpoints

Pretrained navigation policies trained with [VisNavKit](https://github.com/VAIL-UCLA/visnavkit).
Each folder holds one model: the Lightning checkpoint, its two ONNX exports (window and streaming),
the export metadata (shapes, anchor times, SHA-256 of both files, PyTorch/ONNX parity) and a sample
input batch.

| Folder | VisNavKit config | Model | Params | Val top-1 ADE@1/2/4 s (m) | Val top-1 FDE (m) |
| --- | --- | --- | --- | --- | --- |
| [`flowpilot-dst-small`](flowpilot-dst-small) | `experiment=flowpilot_dst_clips1k` | FastViT-T12 on frame pairs -> anchored flow DiT (256-d) | 21.4M | 0.093 / 0.193 / 0.422 | 0.897 |
| [`flowpilot-dst-dune`](flowpilot-dst-dune) | `experiment=flowpilot_dune_dst_clips1k` | frozen DUNE ViT-B/14 -> anchored flow DiT (1024-d) | 208.2M | 0.080 / 0.168 / 0.364 | 0.769 |

Both are FlowPilot-DST policies trained on clips1k (20 Hz, 4 s horizon, point goal with 50%
dropout, frozen route VAE, 64 k-means anchors, 4 flow steps). Metrics are on the VisNavKit
clips1k validation split, decoded from zero noise. ONNX: fp32, opset 17, batch 1, top-6 plans;
PyTorch/ONNX parity is within 1e-5.

## Files

```
<folder>/flowpilot_dst_<encoder>.ckpt           # Lightning checkpoint (weights + optimizer state)
<folder>/flowpilot_dst_<encoder>.onnx           # deployment graph
<folder>/flowpilot_dst_<encoder>.metadata.json  # shapes, anchor times, hashes, parity
<folder>/flowpilot_dst_<encoder>.inputs.npz     # traced sample inputs for a smoke run
<folder>/flowpilot_dst_<encoder>_streaming.{onnx,metadata.json,inputs.npz}  # streaming graph, same weights
```

## Usage

ONNX, no VisNavKit needed:

```python
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download

repo, stem = "UCLA-VAIL/Visual-Navigation-Model-Checkpoints", "flowpilot-dst-small/flowpilot_dst_fastvit_t12"
session = ort.InferenceSession(hf_hub_download(repo, f"{stem}.onnx"), providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
feeds = dict(np.load(hf_hub_download(repo, f"{stem}.inputs.npz")))  # replace with live data
modes, probs, speed = session.run(["modes", "probs", "speed"], feeds)
x, y, yaw, v, w = modes[0, 0].T  # best plan: 80 steps at 0.05 s, ego frame (x forward, y left)
```

Inputs are the last 20 frames (1, 20, 3, 216, 384) RGB in [0, 1], route patches, goal, ego
speed/yaw rate and action bounds. The full contract, including frame preparation, is in
[FlowPilot-DST ONNX IO](https://github.com/VAIL-UCLA/visnavkit/blob/dev/docs/flowpilot_dst_onnx.md).

### Streaming ONNX

`*_streaming.onnx` runs one frame per call. It caches each past frame's features (`[global image
feature | route latent]`) in a buffer, encodes only the current frame, and decodes only the current
plan. After 20 frames its decision equals the window graph's (checked at export: DST exact, DUNE
to 5e-7 in ONNX Runtime).

| input | shape | meaning |
| --- | --- | --- |
| `vision` | small (1, 2, 3, 216, 384): [previous, current]; dune (1, 1, 3, 216, 384) | the current frame; for small, the previous frame is all zeros on the first call |
| `route_patch`, `goal`, `ego` | (1, 1, 80, 80), (1, 1, 3), (1, 1, 2) | the current frame's route, goal, `[v, w]` |
| `action_bounds` | (1, 2, 5) | as in the window graph |
| `feat_buffer` | (1, 19, F): F = 768 small, 2048 dune | the last call's `feat_buffer_out`; zeros at startup |
| `buffer_mask` | (1, 19) | the last call's `buffer_mask_out`; zeros at startup |

Outputs: `modes`, `probs`, `speed` as in the window graph, plus `feat_buffer_out`, `buffer_mask_out`.

```python
stem = "flowpilot-dst-dune/flowpilot_dst_dune_vitb14_streaming"
session = ort.InferenceSession(hf_hub_download(repo, f"{stem}.onnx"))
buf, mask = np.zeros((1, 19, 2048), np.float32), np.zeros((1, 19), np.float32)
for frame, route, goal, ego in stream:  # 20 Hz, frames exactly 50 ms apart
    modes, probs, speed, buf, mask = session.run(None, dict(
        vision=frame[None, None], route_patch=route[None, None], goal=goal[None, None],
        ego=ego[None, None], action_bounds=bounds, feat_buffer=buf, buffer_mask=mask))
    plan = modes[0, 0]  # (80, 5)
```

- Feed `feat_buffer_out` and `buffer_mask_out` back unchanged.
- Call once per frame at 20 Hz. Don't skip frames. Reset the buffer (and, for small, the previous
  frame) to zeros after a gap or restart.
- Keep one buffer per camera stream.
- CPU latency per call (ONNX Runtime, loaded machine): small 1248 ms window vs 381 ms streaming;
  dune 7537 ms vs 2021 ms.

Re-export from the checkpoint with VisNavKit (add `streaming=true` for the streaming graph):

```bash
uv run visnavkit-export-dst checkpoint=flowpilot_dst_fastvit_t12.ckpt output=flowpilot_dst_fastvit_t12.onnx
```

## Citation

```bibtex
@Misc{visnavkit2026,
  author       = {Honglin He and Bolei Zhou},
  title        = {{VisNavKit}: a composable toolkit for visual navigation policies},
  howpublished = {\url{https://github.com/VAIL-UCLA/visnavkit}},
  year         = {2026},
}
```

FlowPilot: [arXiv:2606.12603](https://arxiv.org/abs/2606.12603). The DUNE variant uses the
[DUNE](https://github.com/naver/dune) ViT-B/14 encoder.