File size: 6,389 Bytes
249048c c451aec 249048c c451aec 249048c 1024868 249048c | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 | ---
license: apache-2.0
base_model:
- onnxmodelzoo/retinanet-9
pipeline_tag: object-detection
tags:
- object-detection
- computer-vision
- renesas
- x5h
- onnx
- retinanet
- resnet101
- detection
---
# RetinaNet-R101 (ONNX) – Renesas X5H
## Introduction
This repository hosts **RetinaNet** in ONNX FP32 format, targeting the **Renesas R-Car X5H** platform for object detection inference on the NPX6 NPU.
- **Model Architecture:** RetinaNet with ResNet101 backbone and Feature Pyramid Network (FPN)
- **Source Model:** [onnxmodelzoo/retinanet-9](https://huggingface.co/onnxmodelzoo/retinanet-9) — ONNX Model Zoo [`retinanet-9`](https://github.com/onnx/models)
- **Task:** Object Detection
- **Dataset:** COCO
- **Accuracy:** mAP = 0.376
- **Backbone:** ResNet101
## Deployment Flow
The repository provides the model in **FP32 ONNX** format. Both supported runtimes automatically cast the FP32 model to **INT8** at load time for optimised NPU execution — no separate quantization step is required.
```text
retinanet-9.onnx (FP32)
│
├─▶ ONNX Runtime (Custom NPU EP) ──▶ INT8 auto-cast ──▶ NPX6 NPU
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
```
## Provided Artifacts
| Artifact | Status | Notes |
|----------|---------|---------|
| **FP32 (ONNX)** | ✅ Provided | Reference model from ONNX Model Zoo |
> INT8 execution is handled automatically by the NPU runtime — no additional quantized model file is needed.
## Performance
All HIL results were measured on **Renesas R-Car X5H** physical hardware.
The FP32 ONNX model is auto-cast to INT8 by the runtime before NPU execution.
PPA Estimator results are software estimates based on model characteristics and hardware configuration.
> **Benchmark configuration:** Single NPU · Single AI Core · Input: 3 × 480 × 640 · Batch size: 1
### Inference Latency & Throughput
| Runtime | Precision | Device | Latency (ms) | Throughput (fps) | Type |
|----------|----------|----------|----------|----------|----------|
| ORT Custom NPU EP | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| PPA Estimator | INT8 | X5H · 1× NPU · 1 Core · 1066 MHz | TBD | — | Estimated |
### Accuracy (COCO Validation Set)
| Runtime / Precision | mAP (IoU=0.50:0.95) | Notes |
|----------|----------|----------|
| FP32 Reference | 0.376 | ONNX Model Zoo reference |
| ORT Custom NPU EP (INT8) | TBD | NPU execution |
| MWMX Runtime (INT8) | TBD | NPU execution |
---
## Runtime Details
### ONNX Runtime – Custom NPU Execution Provider
- **Engine:** ONNX Runtime with Renesas Custom NPU Execution Provider
- **Input format:** FP32 ONNX (`.onnx`)
- **NPU execution precision:** INT8 (auto-cast at load time)
- **Execution target:** NPX6-48K NPU on R-Car X5H
### MWMX Runtime
- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
- **Input format:** FP32 ONNX (ingested and compiled by the MWMX toolchain)
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
- **Execution target:** NPX6-48K NPU on R-Car X5H
### PPA Estimator
- **Engine:** Renesas PPA Estimator
- **Input format:** FP32 ONNX
- **NPU execution precision:** INT8
- **Type:** Software performance estimate — not measured on physical silicon
---
## Model Input
### Input Tensor
- Shape: `(N, 3, H, W)`
- Format: RGB
- Data Type: FP32
- Pixel Range: `[0, 1]`
### Preprocessing
```python
from torchvision import transforms
preprocess = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
```
---
## Model Outputs
The model produces **10 output tensors** corresponding to RetinaNet's multi-scale detection heads.
### Classification Heads
Five tensors corresponding to object classification on feature pyramid levels P3–P7.
Example shapes for an input image of size `1 × 3 × 480 × 640`:
```text
[1, 720, 60, 80]
[1, 720, 30, 40]
[1, 720, 15, 20]
[1, 720, 8, 10]
[1, 720, 4, 5]
```
### Bounding Box Regression Heads
Five tensors corresponding to anchor-box regression outputs.
```text
[1, 36, 60, 80]
[1, 36, 30, 40]
[1, 36, 15, 20]
[1, 36, 8, 10]
[1, 36, 4, 5]
```
### Postprocessing
RetinaNet requires the following postprocessing steps:
1. Anchor generation
2. Bounding box decoding
3. Confidence threshold filtering
4. Non-Maximum Suppression (NMS)
These steps produce the final object detections:
- Bounding boxes
- Confidence scores
- Class labels
---
## Prerequisites
To run inference on Renesas R-Car X5H, you need:
1. **Renesas R-Car X5H board** with NPX6 NPU
2. **ONNX Runtime** with Renesas NPU Custom Execution Provider, or the **Renesas MWMX Runtime**
3. **Hugging Face CLI** to download the model
## Download
```bash
hf download Renesas/RetinaNet-R101-ONNX --repo-type=model --include "fp32/*"
```
## Inference
### ONNX Runtime (Custom NPU Execution Provider)
```python
import onnxruntime as ort
import numpy as np
providers = [
("RenesasNPUExecutionProvider", {}),
"CPUExecutionProvider"
]
sess = ort.InferenceSession(
"fp32/retinanet-9.onnx",
providers=providers
)
input_data = np.random.rand(
1, 3, 480, 640
).astype(np.float32)
outputs = sess.run(
None,
{"images": input_data}
)
# outputs[0:5] -> classification heads
# outputs[5:10] -> box regression heads
```
### MWMX Runtime
Refer to the Renesas MWMX Runtime documentation for compilation and inference scripts targeting the NPX6 NPU on R-Car X5H. The MWMX toolchain ingests the FP32 ONNX model and automatically compiles it for INT8 NPU execution.
---
## Benchmark Methodology
- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon; single NPU, single AI core, 850 MHz NPU clock
- **Estimation:** PPA Estimator software estimate; single NPU, single AI core, 1066 MHz NPU clock
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by runtime)
- **Latency:** Median over 1000 consecutive inference runs with warm cache
- **Throughput:** Computed as `1000 / latency_ms`
- **Accuracy:** Evaluated using the COCO validation dataset
- **Postprocessing:** Includes anchor generation, bounding-box decoding, confidence filtering, and NMS |