VisionHOPE ONNX

ONNX exports of PSRben/VisionHOPE (paper, code) for onnxruntime / onnxruntime-web (WebGPU). They power the in-browser demo at mrfakename/visionhope-webgpu.

Files (onnx/)

file task input outputs
visionhope_{tiny,small,base}_cls.onnx ImageNet-1k classification pixel_values float32 [1,3,224,224] logits [1,1000]
visionhope_{tiny,small,base}_ade20k_512.onnx ADE20K semantic segmentation (UPerNet) pixel_values float32 [1,3,512,512] logits [1,150,128,128], labels uint8 [1,512,512]

*_parity.json next to each model holds the parity numbers below.

Static shapes, opset 17, fp32. Each graph was exported from a pure-PyTorch port of the official model (the SRNL scan is unrolled over row/column chunks), then pre-optimised with onnxruntime's EP-independent basic level (constant folding), so sessions start in seconds. Detection (Mask R-CNN) is not exported: its NMS and dynamic RoI stages do not map well to a static WebGPU graph.

Preprocessing

  • Classification: RGB, resize short side to 224 (bicubic in the reference), center-crop 224, scale to [0,1], normalize with mean (0.485, 0.456, 0.406), std (0.229, 0.224, 0.225), NCHW.
  • Segmentation: RGB resized to 512×512 (aspect ratio not kept), normalize in 0–255 space with mean (123.675, 116.28, 103.53), std (58.395, 57.12, 57.375), NCHW. labels is the argmax after bilinear upsampling to 512×512; resize it back to the image size with nearest neighbour. This differs from the official evaluation (keep-ratio, short side 512, sliding/whole inference), so scores will not match the paper exactly.

Parity (ONNX Runtime CPU vs. PyTorch)

model nodes size images agreement max abs diff (logits) ORT CPU s/img
visionhope_tiny_cls.onnx 52,889 129 MB 6 top-1 6/6, top-5 6/6 1.9e-05 0.9
visionhope_small_cls.onnx 80,510 248 MB 6 top-1 6/6, top-5 6/6 2.8e-05 1.7
visionhope_base_cls.onnx 80,510 407 MB 6 top-1 6/6, top-5 6/6 3.7e-05 2.9
visionhope_tiny_ade20k_512.onnx 128,179 278 MB 4 pixels ≥ 99.9996% 4.0e-05 4.2
visionhope_small_ade20k_512.onnx 195,050 420 MB 4 pixels ≥ 99.9996% 4.8e-05 6.1
visionhope_base_ade20k_512.onnx 195,050 589 MB 4 pixels 100% 4.8e-05 7.4

Reference = the PyTorch port in fp32 on the same inputs (COCO val2017 images; 6 for classification, 4 for segmentation), onnxruntime 1.20.1 CPU on 16 vCPUs. Segmentation agreement is the share of identical argmax labels on the 512×512 map; the abs diff is on the 128×128 logits. All six models were also run in headless Chrome with onnxruntime-web on WebGPU (see the Space).

The PyTorch port itself was validated against the official implementation (mmcv/mmseg/timm, CUDA kernels): classification logits within 1e-5, ADE20K pixel agreement ≥ 0.99998.

Usage (Python)

import onnxruntime as ort, numpy as np
s = ort.InferenceSession("visionhope_tiny_cls.onnx")
logits = s.run(None, {"pixel_values": x.astype(np.float32)})[0]
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mrfakename/VisionHOPE-ONNX

Quantized
(1)
this model

Space using mrfakename/VisionHOPE-ONNX 1

Paper for mrfakename/VisionHOPE-ONNX