RetinaNet-R101-ONNX / README.md
Artem Plastinkin
Update
1024868
|
Raw History Blame Contribute Delete
6.39 kB
---
license: apache-2.0
base_model:
- onnxmodelzoo/retinanet-9
pipeline_tag: object-detection
tags:
- object-detection
- computer-vision
- renesas
- x5h
- onnx
- retinanet
- resnet101
- detection
---
# RetinaNet-R101 (ONNX) – Renesas X5H
## Introduction
This repository hosts **RetinaNet** in ONNX FP32 format, targeting the **Renesas R-Car X5H** platform for object detection inference on the NPX6 NPU.
- **Model Architecture:** RetinaNet with ResNet101 backbone and Feature Pyramid Network (FPN)
- **Source Model:** [onnxmodelzoo/retinanet-9](https://huggingface.co/onnxmodelzoo/retinanet-9) — ONNX Model Zoo [`retinanet-9`](https://github.com/onnx/models)
- **Task:** Object Detection
- **Dataset:** COCO
- **Accuracy:** mAP = 0.376
- **Backbone:** ResNet101
## Deployment Flow
The repository provides the model in **FP32 ONNX** format. Both supported runtimes automatically cast the FP32 model to **INT8** at load time for optimised NPU execution — no separate quantization step is required.
```text
retinanet-9.onnx (FP32)
│
├─▶ ONNX Runtime (Custom NPU EP) ──▶ INT8 auto-cast ──▶ NPX6 NPU
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
```
## Provided Artifacts
| Artifact | Status | Notes |
|----------|---------|---------|
| **FP32 (ONNX)** | ✅ Provided | Reference model from ONNX Model Zoo |
> INT8 execution is handled automatically by the NPU runtime — no additional quantized model file is needed.
## Performance
All HIL results were measured on **Renesas R-Car X5H** physical hardware.
The FP32 ONNX model is auto-cast to INT8 by the runtime before NPU execution.
PPA Estimator results are software estimates based on model characteristics and hardware configuration.
> **Benchmark configuration:** Single NPU · Single AI Core · Input: 3 × 480 × 640 · Batch size: 1
### Inference Latency & Throughput
| Runtime | Precision | Device | Latency (ms) | Throughput (fps) | Type |
|----------|----------|----------|----------|----------|----------|
| ORT Custom NPU EP | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | TBD | TBD | Measured |
| PPA Estimator | INT8 | X5H · 1× NPU · 1 Core · 1066 MHz | TBD | — | Estimated |
### Accuracy (COCO Validation Set)
| Runtime / Precision | mAP (IoU=0.50:0.95) | Notes |
|----------|----------|----------|
| FP32 Reference | 0.376 | ONNX Model Zoo reference |
| ORT Custom NPU EP (INT8) | TBD | NPU execution |
| MWMX Runtime (INT8) | TBD | NPU execution |
---
## Runtime Details
### ONNX Runtime – Custom NPU Execution Provider
- **Engine:** ONNX Runtime with Renesas Custom NPU Execution Provider
- **Input format:** FP32 ONNX (`.onnx`)
- **NPU execution precision:** INT8 (auto-cast at load time)
- **Execution target:** NPX6-48K NPU on R-Car X5H
### MWMX Runtime
- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
- **Input format:** FP32 ONNX (ingested and compiled by the MWMX toolchain)
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
- **Execution target:** NPX6-48K NPU on R-Car X5H
### PPA Estimator
- **Engine:** Renesas PPA Estimator
- **Input format:** FP32 ONNX
- **NPU execution precision:** INT8
- **Type:** Software performance estimate — not measured on physical silicon
---
## Model Input
### Input Tensor
- Shape: `(N, 3, H, W)`
- Format: RGB
- Data Type: FP32
- Pixel Range: `[0, 1]`
### Preprocessing
```python
from torchvision import transforms
preprocess = transforms.Compose([
transforms.ToTensor(),
transforms.Normalize(
mean=[0.485, 0.456, 0.406],
std=[0.229, 0.224, 0.225]
),
])
```
---
## Model Outputs
The model produces **10 output tensors** corresponding to RetinaNet's multi-scale detection heads.
### Classification Heads
Five tensors corresponding to object classification on feature pyramid levels P3–P7.
Example shapes for an input image of size `1 × 3 × 480 × 640`:
```text
[1, 720, 60, 80]
[1, 720, 30, 40]
[1, 720, 15, 20]
[1, 720, 8, 10]
[1, 720, 4, 5]
```
### Bounding Box Regression Heads
Five tensors corresponding to anchor-box regression outputs.
```text
[1, 36, 60, 80]
[1, 36, 30, 40]
[1, 36, 15, 20]
[1, 36, 8, 10]
[1, 36, 4, 5]
```
### Postprocessing
RetinaNet requires the following postprocessing steps:
1. Anchor generation
2. Bounding box decoding
3. Confidence threshold filtering
4. Non-Maximum Suppression (NMS)
These steps produce the final object detections:
- Bounding boxes
- Confidence scores
- Class labels
---
## Prerequisites
To run inference on Renesas R-Car X5H, you need:
1. **Renesas R-Car X5H board** with NPX6 NPU
2. **ONNX Runtime** with Renesas NPU Custom Execution Provider, or the **Renesas MWMX Runtime**
3. **Hugging Face CLI** to download the model
## Download
```bash
hf download Renesas/RetinaNet-R101-ONNX --repo-type=model --include "fp32/*"
```
## Inference
### ONNX Runtime (Custom NPU Execution Provider)
```python
import onnxruntime as ort
import numpy as np
providers = [
("RenesasNPUExecutionProvider", {}),
"CPUExecutionProvider"
]
sess = ort.InferenceSession(
"fp32/retinanet-9.onnx",
providers=providers
)
input_data = np.random.rand(
1, 3, 480, 640
).astype(np.float32)
outputs = sess.run(
None,
{"images": input_data}
)
# outputs[0:5] -> classification heads
# outputs[5:10] -> box regression heads
```
### MWMX Runtime
Refer to the Renesas MWMX Runtime documentation for compilation and inference scripts targeting the NPX6 NPU on R-Car X5H. The MWMX toolchain ingests the FP32 ONNX model and automatically compiles it for INT8 NPU execution.
---
## Benchmark Methodology
- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon; single NPU, single AI core, 850 MHz NPU clock
- **Estimation:** PPA Estimator software estimate; single NPU, single AI core, 1066 MHz NPU clock
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by runtime)
- **Latency:** Median over 1000 consecutive inference runs with warm cache
- **Throughput:** Computed as `1000 / latency_ms`
- **Accuracy:** Evaluated using the COCO validation dataset
- **Postprocessing:** Includes anchor generation, bounding-box decoding, confidence filtering, and NMS