Add ONNX exports (dynamic batch; static 480x736 and dynamic H/W)
#2
by vpraveen-nv - opened
This PR adds two ONNX exports of model_best_bp2_serialize.pth under onnx/. Both have a dynamic batch axis.
| File | Input shape | Size (bytes) | sha256 |
|---|---|---|---|
onnx/c_fast_foundationstereo_bp2_dynbatch_480x736.onnx |
[B, 3, 480, 736] (H/W fixed) |
103,228,520 | 794647dc718c9564d5ccb56d15527b70a9267e8ddfbd5ecb1e784b323fbd60ed |
onnx/c_fast_foundationstereo_bp2_dynbatch_dynhw.onnx |
[B, 3, H, W] (H/W dynamic) |
103,228,473 | 2c79bbb274dc0aef687a8732b987077e9e3aa26ff9dd653141cda48c57dad5d3 |
Inputs and outputs
| Tensor | Type | Shape | Dynamic axes |
|---|---|---|---|
left_image |
float32 | [B, 3, H, W] |
batch_size, plus height/width in the dyn-H/W file |
right_image |
float32 | [B, 3, H, W] |
same as left_image |
disparity |
float32 | [B, 1, H, W] |
follows the inputs |
- The inputs are a rectified stereo pair in RGB, NCHW layout. Both inputs must have the same shape.
- Normalization is not part of the graph. Before inference, scale pixels to
[0, 1], then normalize with the ImageNet mean(0.485, 0.456, 0.406)and std(0.229, 0.224, 0.225). TAOdepth_netuses this same preprocessing for training and inference. - The output is the disparity of the left image, in pixels, at the input resolution.
- H and W must be multiples of 32. Other sizes fail at a
Concatin the feature decoder. Pad the pair up to the next multiple of 32 and crop the output back to the original size. The static file accepts only 480x736.
Provenance
- Both files were exported from this repo's
model_best_bp2_serialize.pth(sha2567aee85948373da62b0503c2542507129a3e7cab9d97d10e6790d89512a7db214). - Export settings: TAO Toolkit
depth_net export(torch.onnx.export), opset 17, traced at 1x3x480x736,valid_iters: 8. Each graph has 22,792 nodes and no custom ops. - The export source includes two fixes:
- NVIDIA-TAO/tao-pytorch#152 keeps the batch axis truly dynamic.
- NVIDIA-TAO/tao-pytorch#153 replaces the correlation-pyramid AveragePool. Without it, batch >= 3 fails to build. PyTorch (eager) outputs are bitwise identical with and without the change.
Validation
- TensorRT 10.16 on an A100 80GB PCIe:
- Both files build and run at 480x736 in fp16 and fp32 (TF32) for batch 1, 2, 4, 8, 16 and 32.
- Each engine used a single shape,
min = opt = max. - fp16 throughput: 49.3 img/s at batch 1 and 54.2 img/s at batch 32 (static file); 48.5 and 54.4 img/s (dyn-H/W file).
- ONNX Runtime CUDA EP:
- Batch 3, 4 and 8 run at 480x736.
- The dyn-H/W file also runs at 256x320 with batch 39 and 48.
- At batch 4 and 8, changing other images in the batch did not change any image's output.
- ONNX Runtime CPU:
onnx.checkerpasses for both files.- Batch-2 smoke runs complete, with finite outputs of the expected shape: 480x736 for both files, and 256x320 for the dyn-H/W file.
- With a synthetic pair shifted by 16 px, the dyn-H/W file predicts 16.0 px of disparity.
- TensorRT was benchmarked at 480x736 only. Other resolutions were checked only with ONNX Runtime.
TensorRT tips
trtexec --onnx=onnx/c_fast_foundationstereo_bp2_dynbatch_dynhw.onnx --fp16 \
--minShapes=left_image:4x3x480x736,right_image:4x3x480x736 \
--optShapes=left_image:4x3x480x736,right_image:4x3x480x736 \
--maxShapes=left_image:4x3x480x736,right_image:4x3x480x736 \
--memPoolSize=workspace:24576 --saveEngine=c_fast_foundationstereo_fp16_b4.engine
- Give both
left_imageandright_imagein every shape flag. - Builder workspace:
- 8 GiB is enough for fp16 up to batch 8 and fp32 up to batch 2.
- fp16 at batch 16 and 32 needs about 11 GB and 22 GB;
--memPoolSize=workspace:24576(24 GiB) covers fp16 batch 1-32. - fp32 needs about 17 GB at batch 4-8 and 34 GB at batch 16. At batch 32 it needs a 64 GiB workspace, and its execution context uses about 78 GB, which is the limit of an 80 GB GPU.
- If the workspace is too small, the build fails with
Could not find any implementation for node {ForeignNode[...Concat_387]}. ASkipping tactic ... insufficient memorywarning on its own is harmless if the build succeeds.
Verify the download
sha256sum onnx/*.onnx
# 794647dc718c9564d5ccb56d15527b70a9267e8ddfbd5ecb1e784b323fbd60ed onnx/c_fast_foundationstereo_bp2_dynbatch_480x736.onnx
# 2c79bbb274dc0aef687a8732b987077e9e3aa26ff9dd653141cda48c57dad5d3 onnx/c_fast_foundationstereo_bp2_dynbatch_dynhw.onnx
vpraveen-nv changed pull request status to open
vpraveen-nv changed pull request status to merged