File size: 6,389 Bytes
249048c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c451aec
249048c
 
 
 
 
 
c451aec
249048c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1024868
249048c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
---
license: apache-2.0
base_model:
- onnxmodelzoo/retinanet-9
pipeline_tag: object-detection
tags:
- object-detection
- computer-vision
- renesas
- x5h
- onnx
- retinanet
- resnet101
- detection
---

# RetinaNet-R101 (ONNX) – Renesas X5H

## Introduction

This repository hosts **RetinaNet** in ONNX FP32 format, targeting the **Renesas R-Car X5H** platform for object detection inference on the NPX6 NPU.

- **Model Architecture:** RetinaNet with ResNet101 backbone and Feature Pyramid Network (FPN)
- **Source Model:** [onnxmodelzoo/retinanet-9](https://huggingface.co/onnxmodelzoo/retinanet-9) — ONNX Model Zoo [`retinanet-9`](https://github.com/onnx/models)
- **Task:** Object Detection
- **Dataset:** COCO
- **Accuracy:** mAP = 0.376
- **Backbone:** ResNet101

## Deployment Flow

The repository provides the model in **FP32 ONNX** format. Both supported runtimes automatically cast the FP32 model to **INT8** at load time for optimised NPU execution — no separate quantization step is required.

```text
retinanet-9.onnx (FP32)
        │
        ├─▶  ONNX Runtime (Custom NPU EP)  ──▶  INT8 auto-cast  ──▶  NPX6 NPU
        │
        └─▶  MWMX Runtime                  ──▶  INT8 auto-cast  ──▶  NPX6 NPU
```

## Provided Artifacts

| Artifact | Status | Notes |
|----------|---------|---------|
| **FP32 (ONNX)** | ✅ Provided | Reference model from ONNX Model Zoo |

> INT8 execution is handled automatically by the NPU runtime — no additional quantized model file is needed.

## Performance

All HIL results were measured on **Renesas R-Car X5H** physical hardware.  
The FP32 ONNX model is auto-cast to INT8 by the runtime before NPU execution.  
PPA Estimator results are software estimates based on model characteristics and hardware configuration.

> **Benchmark configuration:** Single NPU · Single AI Core · Input: 3 × 480 × 640 · Batch size: 1

### Inference Latency & Throughput

| Runtime | Precision | Device | Latency (ms) | Throughput (fps) | Type |
|----------|----------|----------|----------|----------|----------|
| ORT Custom NPU EP | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| PPA Estimator | INT8 | X5H · 1× NPU · 1 Core · 1066 MHz | TBD | — | Estimated |

### Accuracy (COCO Validation Set)

| Runtime / Precision | mAP (IoU=0.50:0.95) | Notes |
|----------|----------|----------|
| FP32 Reference | 0.376 | ONNX Model Zoo reference |
| ORT Custom NPU EP (INT8) | TBD | NPU execution |
| MWMX Runtime (INT8) | TBD | NPU execution |

---

## Runtime Details

### ONNX Runtime – Custom NPU Execution Provider

- **Engine:** ONNX Runtime with Renesas Custom NPU Execution Provider
- **Input format:** FP32 ONNX (`.onnx`)
- **NPU execution precision:** INT8 (auto-cast at load time)
- **Execution target:** NPX6-48K NPU on R-Car X5H

### MWMX Runtime

- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
- **Input format:** FP32 ONNX (ingested and compiled by the MWMX toolchain)
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
- **Execution target:** NPX6-48K NPU on R-Car X5H

### PPA Estimator

- **Engine:** Renesas PPA Estimator
- **Input format:** FP32 ONNX
- **NPU execution precision:** INT8
- **Type:** Software performance estimate — not measured on physical silicon

---

## Model Input

### Input Tensor

- Shape: `(N, 3, H, W)`
- Format: RGB
- Data Type: FP32
- Pixel Range: `[0, 1]`

### Preprocessing

```python
from torchvision import transforms

preprocess = transforms.Compose([
    transforms.ToTensor(),
    transforms.Normalize(
        mean=[0.485, 0.456, 0.406],
        std=[0.229, 0.224, 0.225]
    ),
])
```

---

## Model Outputs

The model produces **10 output tensors** corresponding to RetinaNet's multi-scale detection heads.

### Classification Heads

Five tensors corresponding to object classification on feature pyramid levels P3–P7.

Example shapes for an input image of size `1 × 3 × 480 × 640`:

```text
[1, 720, 60, 80]
[1, 720, 30, 40]
[1, 720, 15, 20]
[1, 720, 8, 10]
[1, 720, 4, 5]
```

### Bounding Box Regression Heads

Five tensors corresponding to anchor-box regression outputs.

```text
[1, 36, 60, 80]
[1, 36, 30, 40]
[1, 36, 15, 20]
[1, 36, 8, 10]
[1, 36, 4, 5]
```

### Postprocessing

RetinaNet requires the following postprocessing steps:

1. Anchor generation
2. Bounding box decoding
3. Confidence threshold filtering
4. Non-Maximum Suppression (NMS)

These steps produce the final object detections:

- Bounding boxes
- Confidence scores
- Class labels

---

## Prerequisites

To run inference on Renesas R-Car X5H, you need:

1. **Renesas R-Car X5H board** with NPX6 NPU
2. **ONNX Runtime** with Renesas NPU Custom Execution Provider, or the **Renesas MWMX Runtime**
3. **Hugging Face CLI** to download the model

## Download

```bash
hf download Renesas/RetinaNet-R101-ONNX --repo-type=model --include "fp32/*"
```

## Inference

### ONNX Runtime (Custom NPU Execution Provider)

```python
import onnxruntime as ort
import numpy as np

providers = [
    ("RenesasNPUExecutionProvider", {}),
    "CPUExecutionProvider"
]

sess = ort.InferenceSession(
    "fp32/retinanet-9.onnx",
    providers=providers
)

input_data = np.random.rand(
    1, 3, 480, 640
).astype(np.float32)

outputs = sess.run(
    None,
    {"images": input_data}
)

# outputs[0:5] -> classification heads
# outputs[5:10] -> box regression heads
```

### MWMX Runtime

Refer to the Renesas MWMX Runtime documentation for compilation and inference scripts targeting the NPX6 NPU on R-Car X5H. The MWMX toolchain ingests the FP32 ONNX model and automatically compiles it for INT8 NPU execution.

---

## Benchmark Methodology

- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon; single NPU, single AI core, 850 MHz NPU clock
- **Estimation:** PPA Estimator software estimate; single NPU, single AI core, 1066 MHz NPU clock
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by runtime)
- **Latency:** Median over 1000 consecutive inference runs with warm cache
- **Throughput:** Computed as `1000 / latency_ms`
- **Accuracy:** Evaluated using the COCO validation dataset
- **Postprocessing:** Includes anchor generation, bounding-box decoding, confidence filtering, and NMS