artemPlastinkin commited on
Commit
7a845cb
·
verified ·
1 Parent(s): b83bff5

Initial upload of CenterNet-R18-ONNX

Browse files
.metadata.yaml ADDED
@@ -0,0 +1,38 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # PARAMETERS DELIBERATELY ABSENT. The previous ~14.4M was an estimate built by
2
+ # adding a guessed 2-3M of keypoint heads to the ResNet18 backbone's 11.7M; no
3
+ # published figure exists for this mmdetection config. Count it from the graph.
4
+
5
+ model:
6
+ name: centernet-r18
7
+ display_name: CenterNet-R18
8
+ # upstream: intentionally absent — these weights have no HuggingFace repo.
9
+ # See the `source` block below. Never write "TBD" here: the
10
+ # generator copies it into base_model and renders it as the
11
+ # model's architecture label in the catalog.
12
+ source:
13
+ kind: openmmlab
14
+ id: centernet_resnet18_140e_coco
15
+ url: https://github.com/open-mmlab/mmdetection/blob/main/configs/centernet/metafile.yml
16
+
17
+ architecture:
18
+ family: centernet # lineage only — no version, no size
19
+ backbone: resnet18
20
+ dataset: coco
21
+ num_classes: 80
22
+ modality:
23
+ - vision
24
+ # parameters: intentionally absent — see the note above. Fill it with the
25
+ # exact count instead of an estimate:
26
+ # import onnx; from onnx import numpy_helper
27
+ # m = onnx.load('model.onnx')
28
+ # sum(numpy_helper.to_array(t).size for t in m.graph.initializer)
29
+ parameters_source: unknown
30
+
31
+ format:
32
+ type: onnx
33
+ # opset: read it off the graph rather than guessing —
34
+ # onnx.load('model.onnx').opset_import[0].version
35
+ # NOTE: `version` is a GGUF-only field and must not be used for ONNX.
36
+
37
+ tasks:
38
+ - object-detection
README.md ADDED
@@ -0,0 +1,97 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: []
4
+ pipeline_tag: object-detection
5
+ tags:
6
+ - object-detection
7
+ - computer-vision
8
+ - renesas
9
+ - x5h
10
+ - onnx
11
+ - centernet
12
+ - resnet18
13
+ - detection
14
+ ---
15
+
16
+ # CenterNet-R18 (ONNX) – Renesas X5H
17
+
18
+ > ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
19
+
20
+ ## Introduction
21
+
22
+ This repository hosts **CenterNet** with a **ResNet18** backbone, targeting the **Renesas
23
+ R-Car X5H** platform for object detection inference on the NPX6 NPU.
24
+
25
+ - **Model Architecture:** CenterNet — keypoint-based, anchor-free object detector, ResNet18 backbone
26
+ - **Source Model:** OpenMMLab config [`centernet_resnet18_140e_coco`](https://github.com/open-mmlab/mmdetection/blob/main/configs/centernet/metafile.yml)
27
+ *(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*
28
+ - **Task:** Object Detection
29
+ - **Dataset:** COCO (inferred from checkpoint name)
30
+ - **Input Resolution:** 512 × 512 (inferred from `crop512` in the checkpoint name)
31
+ - **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)
32
+
33
+ ## Deployment Flow
34
+
35
+ The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no
36
+ separate quantization step is required.
37
+
38
+ ```
39
+ centernet_r18_..._optimized.onnx (FP32)
40
+ │
41
+ └─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
42
+ ```
43
+
44
+ ## Provided Artifacts
45
+
46
+ | Artifact | Status | Notes |
47
+ |----------|--------|-------|
48
+ | **FP32 (ONNX)** | ⏳ Not yet uploaded | Benchmark numbers below exist; the model file has not been published to this repo yet |
49
+
50
+ ## Performance
51
+
52
+ Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).
53
+
54
+ > **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 512 (inferred)
55
+
56
+ | Runtime | Precision | Device | Latency (ms) | Type |
57
+ |---------|-----------|--------|---------------|------|
58
+ | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 11.978 | Measured |
59
+ | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 3.467 | Measured |
60
+
61
+ ### Accuracy
62
+
63
+ TBD — not yet measured/published for this repo.
64
+
65
+ ---
66
+
67
+ ## Runtime Details
68
+
69
+ ### MWMX Runtime
70
+
71
+ - **Engine:** Renesas MWMX (Middleware MX) native inference runtime
72
+ - **Input format:** FP32 ONNX (compiled by the MWMX toolchain)
73
+ - **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
74
+ - **Execution target:** NPX6-48K NPU on R-Car X5H
75
+
76
+ ---
77
+
78
+ ## Prerequisites
79
+
80
+ To run inference on Renesas R-Car X5H, you need:
81
+
82
+ 1. **Renesas R-Car X5H board** with NPX6 NPU
83
+ 2. **Renesas MWMX Runtime**
84
+ 3. **Hugging Face CLI** to download the model (once the model file is published)
85
+
86
+ ## Download
87
+
88
+ TBD — model file not yet published to this repository.
89
+
90
+ ---
91
+
92
+ ## Benchmark Methodology
93
+
94
+ - **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
95
+ runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
96
+ - **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
97
+ - **Slices:** results reported for both 1 AI core and 12 AI cores per NPU instance
int8/.metadata.yaml ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ variant:
2
+ id: onnx_int8
3
+ format: onnx
4
+ precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime
5
+ method: onnx_export # base model exported to ONNX; quantization applied by runtime
6
+
7
+ quantization:
8
+ datatype: int8
9
+ scope:
10
+ - weights
11
+ - activations
12
+ granularity: per-tensor # typical for runtime-cast INT8
13
+ calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain
14
+ toolchain: onnx
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 12
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 3.467288 # APM80 pipeline run (metawaremx_runtime CI)
33
+
34
+ metrics:
35
+ accuracy: null
36
+ top5_accuracy: null
37
+
38
+ memory:
39
+ peak_mb: null
40
+
41
+ power:
42
+ avg_w: null
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml ADDED
@@ -0,0 +1,42 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 1
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 512] # inferred from "crop512" in the source checkpoint name
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 11.978266 # APM80 pipeline run (metawaremx_runtime CI)
33
+
34
+ metrics:
35
+ accuracy: null
36
+ top5_accuracy: null
37
+
38
+ memory:
39
+ peak_mb: null
40
+
41
+ power:
42
+ avg_w: null