Initial upload of CenterNet-R18-ONNX
Browse files- .metadata.yaml +38 -0
- README.md +97 -0
- int8/.metadata.yaml +14 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml +42 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml +42 -0
.metadata.yaml
ADDED
|
@@ -0,0 +1,38 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# PARAMETERS DELIBERATELY ABSENT. The previous ~14.4M was an estimate built by
|
| 2 |
+
# adding a guessed 2-3M of keypoint heads to the ResNet18 backbone's 11.7M; no
|
| 3 |
+
# published figure exists for this mmdetection config. Count it from the graph.
|
| 4 |
+
|
| 5 |
+
model:
|
| 6 |
+
name: centernet-r18
|
| 7 |
+
display_name: CenterNet-R18
|
| 8 |
+
# upstream: intentionally absent — these weights have no HuggingFace repo.
|
| 9 |
+
# See the `source` block below. Never write "TBD" here: the
|
| 10 |
+
# generator copies it into base_model and renders it as the
|
| 11 |
+
# model's architecture label in the catalog.
|
| 12 |
+
source:
|
| 13 |
+
kind: openmmlab
|
| 14 |
+
id: centernet_resnet18_140e_coco
|
| 15 |
+
url: https://github.com/open-mmlab/mmdetection/blob/main/configs/centernet/metafile.yml
|
| 16 |
+
|
| 17 |
+
architecture:
|
| 18 |
+
family: centernet # lineage only — no version, no size
|
| 19 |
+
backbone: resnet18
|
| 20 |
+
dataset: coco
|
| 21 |
+
num_classes: 80
|
| 22 |
+
modality:
|
| 23 |
+
- vision
|
| 24 |
+
# parameters: intentionally absent — see the note above. Fill it with the
|
| 25 |
+
# exact count instead of an estimate:
|
| 26 |
+
# import onnx; from onnx import numpy_helper
|
| 27 |
+
# m = onnx.load('model.onnx')
|
| 28 |
+
# sum(numpy_helper.to_array(t).size for t in m.graph.initializer)
|
| 29 |
+
parameters_source: unknown
|
| 30 |
+
|
| 31 |
+
format:
|
| 32 |
+
type: onnx
|
| 33 |
+
# opset: read it off the graph rather than guessing —
|
| 34 |
+
# onnx.load('model.onnx').opset_import[0].version
|
| 35 |
+
# NOTE: `version` is a GGUF-only field and must not be used for ONNX.
|
| 36 |
+
|
| 37 |
+
tasks:
|
| 38 |
+
- object-detection
|
README.md
ADDED
|
@@ -0,0 +1,97 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: []
|
| 4 |
+
pipeline_tag: object-detection
|
| 5 |
+
tags:
|
| 6 |
+
- object-detection
|
| 7 |
+
- computer-vision
|
| 8 |
+
- renesas
|
| 9 |
+
- x5h
|
| 10 |
+
- onnx
|
| 11 |
+
- centernet
|
| 12 |
+
- resnet18
|
| 13 |
+
- detection
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# CenterNet-R18 (ONNX) – Renesas X5H
|
| 17 |
+
|
| 18 |
+
> ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
|
| 19 |
+
|
| 20 |
+
## Introduction
|
| 21 |
+
|
| 22 |
+
This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
|
| 23 |
+
R-Car X5H** platform for object detection inference on the NPX6 NPU.
|
| 24 |
+
|
| 25 |
+
- **Model Architecture:** CenterNet — keypoint-based, anchor-free object detector, ResNet18 backbone
|
| 26 |
+
- **Source Model:** OpenMMLab config [`centernet_resnet18_140e_coco`](https://github.com/open-mmlab/mmdetection/blob/main/configs/centernet/metafile.yml)
|
| 27 |
+
*(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*
|
| 28 |
+
- **Task:** Object Detection
|
| 29 |
+
- **Dataset:** COCO (inferred from checkpoint name)
|
| 30 |
+
- **Input Resolution:** 512 × 512 (inferred from `crop512` in the checkpoint name)
|
| 31 |
+
- **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)
|
| 32 |
+
|
| 33 |
+
## Deployment Flow
|
| 34 |
+
|
| 35 |
+
The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no
|
| 36 |
+
separate quantization step is required.
|
| 37 |
+
|
| 38 |
+
```
|
| 39 |
+
centernet_r18_..._optimized.onnx (FP32)
|
| 40 |
+
│
|
| 41 |
+
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
|
| 42 |
+
```
|
| 43 |
+
|
| 44 |
+
## Provided Artifacts
|
| 45 |
+
|
| 46 |
+
| Artifact | Status | Notes |
|
| 47 |
+
|----------|--------|-------|
|
| 48 |
+
| **FP32 (ONNX)** | ⏳ Not yet uploaded | Benchmark numbers below exist; the model file has not been published to this repo yet |
|
| 49 |
+
|
| 50 |
+
## Performance
|
| 51 |
+
|
| 52 |
+
Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).
|
| 53 |
+
|
| 54 |
+
> **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 512 (inferred)
|
| 55 |
+
|
| 56 |
+
| Runtime | Precision | Device | Latency (ms) | Type |
|
| 57 |
+
|---------|-----------|--------|---------------|------|
|
| 58 |
+
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
|
| 59 |
+
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
|
| 60 |
+
|
| 61 |
+
### Accuracy
|
| 62 |
+
|
| 63 |
+
TBD — not yet measured/published for this repo.
|
| 64 |
+
|
| 65 |
+
---
|
| 66 |
+
|
| 67 |
+
## Runtime Details
|
| 68 |
+
|
| 69 |
+
### MWMX Runtime
|
| 70 |
+
|
| 71 |
+
- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
|
| 72 |
+
- **Input format:** FP32 ONNX (compiled by the MWMX toolchain)
|
| 73 |
+
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
|
| 74 |
+
- **Execution target:** NPX6-48K NPU on R-Car X5H
|
| 75 |
+
|
| 76 |
+
---
|
| 77 |
+
|
| 78 |
+
## Prerequisites
|
| 79 |
+
|
| 80 |
+
To run inference on Renesas R-Car X5H, you need:
|
| 81 |
+
|
| 82 |
+
1. **Renesas R-Car X5H board** with NPX6 NPU
|
| 83 |
+
2. **Renesas MWMX Runtime**
|
| 84 |
+
3. **Hugging Face CLI** to download the model (once the model file is published)
|
| 85 |
+
|
| 86 |
+
## Download
|
| 87 |
+
|
| 88 |
+
TBD — model file not yet published to this repository.
|
| 89 |
+
|
| 90 |
+
---
|
| 91 |
+
|
| 92 |
+
## Benchmark Methodology
|
| 93 |
+
|
| 94 |
+
- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
|
| 95 |
+
runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
|
| 96 |
+
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
|
| 97 |
+
- **Slices:** results reported for both 1 AI core and 12 AI cores per NPU instance
|
int8/.metadata.yaml
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
variant:
|
| 2 |
+
id: onnx_int8
|
| 3 |
+
format: onnx
|
| 4 |
+
precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime
|
| 5 |
+
method: onnx_export # base model exported to ONNX; quantization applied by runtime
|
| 6 |
+
|
| 7 |
+
quantization:
|
| 8 |
+
datatype: int8
|
| 9 |
+
scope:
|
| 10 |
+
- weights
|
| 11 |
+
- activations
|
| 12 |
+
granularity: per-tensor # typical for runtime-cast INT8
|
| 13 |
+
calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain
|
| 14 |
+
toolchain: onnx
|
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 12
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 3.467288 # APM80 pipeline run (metawaremx_runtime CI)
|
| 33 |
+
|
| 34 |
+
metrics:
|
| 35 |
+
accuracy: null
|
| 36 |
+
top5_accuracy: null
|
| 37 |
+
|
| 38 |
+
memory:
|
| 39 |
+
peak_mb: null
|
| 40 |
+
|
| 41 |
+
power:
|
| 42 |
+
avg_w: null
|
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml
ADDED
|
@@ -0,0 +1,42 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 1
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 11.978266 # APM80 pipeline run (metawaremx_runtime CI)
|
| 33 |
+
|
| 34 |
+
metrics:
|
| 35 |
+
accuracy: null
|
| 36 |
+
top5_accuracy: null
|
| 37 |
+
|
| 38 |
+
memory:
|
| 39 |
+
peak_mb: null
|
| 40 |
+
|
| 41 |
+
power:
|
| 42 |
+
avg_w: null
|