artemPlastinkin commited on
Commit
aeea7fa
·
verified ·
1 Parent(s): c87e6ff

Initial upload of DeepLabV3Plus-R50-ONNX

Browse files
.metadata.yaml ADDED
@@ -0,0 +1,46 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # family is `deeplab` with version `v3+`, not `deeplabv3plus`: app.js familyOf()
2
+ # rendered the old slug as "Deeplabv3plus".
3
+ # PARTIAL MODEL. The source checkpoint name ends in
4
+ # `custom_seg_split_4_split_2` — this artifact is ONE segment of a 4-way split
5
+ # network, not the full end-to-end model. Latency below reflects only that segment;
6
+ # do not quote it as whole-model latency.
7
+ # PARAMETERS DELIBERATELY ABSENT for exactly that reason: the commonly-cited ~41M
8
+ # is the FULL DeepLabV3+/ResNet50 model, so attaching it to one of four segments
9
+ # would overstate this artifact several-fold. Count this segment from the graph.
10
+
11
+ model:
12
+ name: deeplabv3plus-r50
13
+ display_name: DeepLabV3Plus-R50
14
+ # upstream: intentionally absent — these weights have no HuggingFace repo.
15
+ # See the `source` block below. Never write "TBD" here: the
16
+ # generator copies it into base_model and renders it as the
17
+ # model's architecture label in the catalog.
18
+ source:
19
+ kind: openmmlab
20
+ id: deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024
21
+ url: https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml
22
+
23
+ architecture:
24
+ family: deeplab # lineage only — no version, no size
25
+ version: "v3+"
26
+ backbone: resnet50-d8
27
+ dataset: cityscapes
28
+ input_resolution: 512x1024
29
+ num_classes: 19
30
+ modality:
31
+ - vision
32
+ # parameters: intentionally absent — see the note above. Fill it with the
33
+ # exact count instead of an estimate:
34
+ # import onnx; from onnx import numpy_helper
35
+ # m = onnx.load('model.onnx')
36
+ # sum(numpy_helper.to_array(t).size for t in m.graph.initializer)
37
+ parameters_source: unknown
38
+
39
+ format:
40
+ type: onnx
41
+ # opset: read it off the graph rather than guessing —
42
+ # onnx.load('model.onnx').opset_import[0].version
43
+ # NOTE: `version` is a GGUF-only field and must not be used for ONNX.
44
+
45
+ tasks:
46
+ - semantic-segmentation
README.md ADDED
@@ -0,0 +1,105 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ base_model: []
4
+ pipeline_tag: image-segmentation
5
+ tags:
6
+ - semantic-segmentation
7
+ - computer-vision
8
+ - renesas
9
+ - x5h
10
+ - onnx
11
+ - deeplabv3plus
12
+ - resnet50
13
+ - cityscapes
14
+ ---
15
+
16
+ # DeepLabV3Plus-R50 (ONNX) – Renesas X5H
17
+
18
+ > ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
19
+
20
+ > **⚠️ Partial-model caveat.** The source checkpoint name ends in
21
+ > `custom_seg_split_4_split_2` — this artifact is **one segment of a 4-way split network**, not
22
+ > the full end-to-end DeepLabV3+ model. The latency below reflects only that segment; do not
23
+ > quote it as whole-model latency until the other splits are accounted for.
24
+
25
+ ## Introduction
26
+
27
+ This repository hosts **DeepLabV3+ (ResNet50-D8 backbone)** targeting the **Renesas R-Car X5H**
28
+ platform for semantic segmentation inference on the NPX6 NPU.
29
+
30
+ - **Model Architecture:** DeepLabV3+ with ResNet50-D8 backbone
31
+ - **Source Model:** OpenMMLab config [`deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024`](https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml)
32
+ *(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*
33
+ - **Task:** Semantic Segmentation
34
+ - **Dataset:** Cityscapes (inferred from checkpoint name)
35
+ - **Input Resolution:** 512 × 1024 (explicit in checkpoint name)
36
+ - **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)
37
+
38
+ ## Deployment Flow
39
+
40
+ The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no
41
+ separate quantization step is required.
42
+
43
+ ```
44
+ deeplabv3plus_..._split_2.onnx (FP32)
45
+ │
46
+ └─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
47
+ ```
48
+
49
+ ## Provided Artifacts
50
+
51
+ | Artifact | Status | Notes |
52
+ |----------|--------|-------|
53
+ | **FP32 (ONNX)** | ⏳ Not yet uploaded | Benchmark numbers below exist; the model file has not been published to this repo yet |
54
+
55
+ ## Performance
56
+
57
+ Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).
58
+
59
+ > **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 1024
60
+ >
61
+ > The 1-AI-core slice failed to compile in the source CI pipeline, so only the 12-core result
62
+ > is available. Latency is for the `split_2` segment only (see caveat above).
63
+
64
+ | Runtime | Precision | Device | Latency (ms) | Type |
65
+ |---------|-----------|--------|---------------|------|
66
+ | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 38.104 | Measured |
67
+
68
+ ### Accuracy (mIoU)
69
+
70
+ TBD — not yet measured/published for this repo.
71
+
72
+ ---
73
+
74
+ ## Runtime Details
75
+
76
+ ### MWMX Runtime
77
+
78
+ - **Engine:** Renesas MWMX (Middleware MX) native inference runtime
79
+ - **Input format:** FP32 ONNX (compiled by the MWMX toolchain)
80
+ - **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
81
+ - **Execution target:** NPX6-48K NPU on R-Car X5H
82
+
83
+ ---
84
+
85
+ ## Prerequisites
86
+
87
+ To run inference on Renesas R-Car X5H, you need:
88
+
89
+ 1. **Renesas R-Car X5H board** with NPX6 NPU
90
+ 2. **Renesas MWMX Runtime**
91
+ 3. **Hugging Face CLI** to download the model (once the model file is published)
92
+
93
+ ## Download
94
+
95
+ TBD — model file not yet published to this repository.
96
+
97
+ ---
98
+
99
+ ## Benchmark Methodology
100
+
101
+ - **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
102
+ runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
103
+ - **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
104
+ - **Slices:** only the 12-AI-core result is available (1-core compile failed)
105
+ - **Scope:** this artifact is a single segment (`split_2` of 4) of the full segmentation pipeline
int8/.metadata.yaml ADDED
@@ -0,0 +1,14 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ variant:
2
+ id: onnx_int8
3
+ format: onnx
4
+ precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime
5
+ method: onnx_export # base model exported to ONNX; quantization applied by runtime
6
+
7
+ quantization:
8
+ datatype: int8
9
+ scope:
10
+ - weights
11
+ - activations
12
+ granularity: per-tensor # typical for runtime-cast INT8
13
+ calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain
14
+ toolchain: onnx
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 12
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 38.103941 # APM80 pipeline run (metawaremx_runtime CI); this artifact is ONE segment
33
+ # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
+ # 1-core run for this artifact failed to compile — no figure available for that slice.
35
+
36
+ metrics:
37
+ accuracy: null
38
+ top5_accuracy: null
39
+
40
+ memory:
41
+ peak_mb: null
42
+
43
+ power:
44
+ avg_w: null