Initial upload of DeepLabV3Plus-R50-ONNX
Browse files- .metadata.yaml +46 -0
- README.md +105 -0
- int8/.metadata.yaml +14 -0
- int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml +44 -0
.metadata.yaml
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# family is `deeplab` with version `v3+`, not `deeplabv3plus`: app.js familyOf()
|
| 2 |
+
# rendered the old slug as "Deeplabv3plus".
|
| 3 |
+
# PARTIAL MODEL. The source checkpoint name ends in
|
| 4 |
+
# `custom_seg_split_4_split_2` — this artifact is ONE segment of a 4-way split
|
| 5 |
+
# network, not the full end-to-end model. Latency below reflects only that segment;
|
| 6 |
+
# do not quote it as whole-model latency.
|
| 7 |
+
# PARAMETERS DELIBERATELY ABSENT for exactly that reason: the commonly-cited ~41M
|
| 8 |
+
# is the FULL DeepLabV3+/ResNet50 model, so attaching it to one of four segments
|
| 9 |
+
# would overstate this artifact several-fold. Count this segment from the graph.
|
| 10 |
+
|
| 11 |
+
model:
|
| 12 |
+
name: deeplabv3plus-r50
|
| 13 |
+
display_name: DeepLabV3Plus-R50
|
| 14 |
+
# upstream: intentionally absent — these weights have no HuggingFace repo.
|
| 15 |
+
# See the `source` block below. Never write "TBD" here: the
|
| 16 |
+
# generator copies it into base_model and renders it as the
|
| 17 |
+
# model's architecture label in the catalog.
|
| 18 |
+
source:
|
| 19 |
+
kind: openmmlab
|
| 20 |
+
id: deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024
|
| 21 |
+
url: https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml
|
| 22 |
+
|
| 23 |
+
architecture:
|
| 24 |
+
family: deeplab # lineage only — no version, no size
|
| 25 |
+
version: "v3+"
|
| 26 |
+
backbone: resnet50-d8
|
| 27 |
+
dataset: cityscapes
|
| 28 |
+
input_resolution: 512x1024
|
| 29 |
+
num_classes: 19
|
| 30 |
+
modality:
|
| 31 |
+
- vision
|
| 32 |
+
# parameters: intentionally absent — see the note above. Fill it with the
|
| 33 |
+
# exact count instead of an estimate:
|
| 34 |
+
# import onnx; from onnx import numpy_helper
|
| 35 |
+
# m = onnx.load('model.onnx')
|
| 36 |
+
# sum(numpy_helper.to_array(t).size for t in m.graph.initializer)
|
| 37 |
+
parameters_source: unknown
|
| 38 |
+
|
| 39 |
+
format:
|
| 40 |
+
type: onnx
|
| 41 |
+
# opset: read it off the graph rather than guessing —
|
| 42 |
+
# onnx.load('model.onnx').opset_import[0].version
|
| 43 |
+
# NOTE: `version` is a GGUF-only field and must not be used for ONNX.
|
| 44 |
+
|
| 45 |
+
tasks:
|
| 46 |
+
- semantic-segmentation
|
README.md
ADDED
|
@@ -0,0 +1,105 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
base_model: []
|
| 4 |
+
pipeline_tag: image-segmentation
|
| 5 |
+
tags:
|
| 6 |
+
- semantic-segmentation
|
| 7 |
+
- computer-vision
|
| 8 |
+
- renesas
|
| 9 |
+
- x5h
|
| 10 |
+
- onnx
|
| 11 |
+
- deeplabv3plus
|
| 12 |
+
- resnet50
|
| 13 |
+
- cityscapes
|
| 14 |
+
---
|
| 15 |
+
|
| 16 |
+
# DeepLabV3Plus-R50 (ONNX) – Renesas X5H
|
| 17 |
+
|
| 18 |
+
> ⏳ **Model file not yet uploaded.** Benchmark results on this page were published ahead of the model weights — see **Provided Artifacts** below. Download/deployment steps will not work until the file is added to this repository.
|
| 19 |
+
|
| 20 |
+
> **⚠️ Partial-model caveat.** The source checkpoint name ends in
|
| 21 |
+
> `custom_seg_split_4_split_2` — this artifact is **one segment of a 4-way split network**, not
|
| 22 |
+
> the full end-to-end DeepLabV3+ model. The latency below reflects only that segment; do not
|
| 23 |
+
> quote it as whole-model latency until the other splits are accounted for.
|
| 24 |
+
|
| 25 |
+
## Introduction
|
| 26 |
+
|
| 27 |
+
This repository hosts **DeepLabV3+ (ResNet50-D8 backbone)** targeting the **Renesas R-Car X5H**
|
| 28 |
+
platform for semantic segmentation inference on the NPX6 NPU.
|
| 29 |
+
|
| 30 |
+
- **Model Architecture:** DeepLabV3+ with ResNet50-D8 backbone
|
| 31 |
+
- **Source Model:** OpenMMLab config [`deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024`](https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml)
|
| 32 |
+
*(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*
|
| 33 |
+
- **Task:** Semantic Segmentation
|
| 34 |
+
- **Dataset:** Cityscapes (inferred from checkpoint name)
|
| 35 |
+
- **Input Resolution:** 512 × 1024 (explicit in checkpoint name)
|
| 36 |
+
- **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)
|
| 37 |
+
|
| 38 |
+
## Deployment Flow
|
| 39 |
+
|
| 40 |
+
The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no
|
| 41 |
+
separate quantization step is required.
|
| 42 |
+
|
| 43 |
+
```
|
| 44 |
+
deeplabv3plus_..._split_2.onnx (FP32)
|
| 45 |
+
│
|
| 46 |
+
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
|
| 47 |
+
```
|
| 48 |
+
|
| 49 |
+
## Provided Artifacts
|
| 50 |
+
|
| 51 |
+
| Artifact | Status | Notes |
|
| 52 |
+
|----------|--------|-------|
|
| 53 |
+
| **FP32 (ONNX)** | ⏳ Not yet uploaded | Benchmark numbers below exist; the model file has not been published to this repo yet |
|
| 54 |
+
|
| 55 |
+
## Performance
|
| 56 |
+
|
| 57 |
+
Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).
|
| 58 |
+
|
| 59 |
+
> **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 1024
|
| 60 |
+
>
|
| 61 |
+
> The 1-AI-core slice failed to compile in the source CI pipeline, so only the 12-core result
|
| 62 |
+
> is available. Latency is for the `split_2` segment only (see caveat above).
|
| 63 |
+
|
| 64 |
+
| Runtime | Precision | Device | Latency (ms) | Type |
|
| 65 |
+
|---------|-----------|--------|---------------|------|
|
| 66 |
+
| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 38.104 | Measured |
|
| 67 |
+
|
| 68 |
+
### Accuracy (mIoU)
|
| 69 |
+
|
| 70 |
+
TBD — not yet measured/published for this repo.
|
| 71 |
+
|
| 72 |
+
---
|
| 73 |
+
|
| 74 |
+
## Runtime Details
|
| 75 |
+
|
| 76 |
+
### MWMX Runtime
|
| 77 |
+
|
| 78 |
+
- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
|
| 79 |
+
- **Input format:** FP32 ONNX (compiled by the MWMX toolchain)
|
| 80 |
+
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
|
| 81 |
+
- **Execution target:** NPX6-48K NPU on R-Car X5H
|
| 82 |
+
|
| 83 |
+
---
|
| 84 |
+
|
| 85 |
+
## Prerequisites
|
| 86 |
+
|
| 87 |
+
To run inference on Renesas R-Car X5H, you need:
|
| 88 |
+
|
| 89 |
+
1. **Renesas R-Car X5H board** with NPX6 NPU
|
| 90 |
+
2. **Renesas MWMX Runtime**
|
| 91 |
+
3. **Hugging Face CLI** to download the model (once the model file is published)
|
| 92 |
+
|
| 93 |
+
## Download
|
| 94 |
+
|
| 95 |
+
TBD — model file not yet published to this repository.
|
| 96 |
+
|
| 97 |
+
---
|
| 98 |
+
|
| 99 |
+
## Benchmark Methodology
|
| 100 |
+
|
| 101 |
+
- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
|
| 102 |
+
runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
|
| 103 |
+
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
|
| 104 |
+
- **Slices:** only the 12-AI-core result is available (1-core compile failed)
|
| 105 |
+
- **Scope:** this artifact is a single segment (`split_2` of 4) of the full segmentation pipeline
|
int8/.metadata.yaml
ADDED
|
@@ -0,0 +1,14 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
variant:
|
| 2 |
+
id: onnx_int8
|
| 3 |
+
format: onnx
|
| 4 |
+
precision: int8 # INT8 quantized — auto-cast by NPU execution provider at runtime
|
| 5 |
+
method: onnx_export # base model exported to ONNX; quantization applied by runtime
|
| 6 |
+
|
| 7 |
+
quantization:
|
| 8 |
+
datatype: int8
|
| 9 |
+
scope:
|
| 10 |
+
- weights
|
| 11 |
+
- activations
|
| 12 |
+
granularity: per-tensor # typical for runtime-cast INT8
|
| 13 |
+
calibration: ptq # post-training quantization applied by the NPU EP / MWMX toolchain
|
| 14 |
+
toolchain: onnx
|
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 12
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 38.103941 # APM80 pipeline run (metawaremx_runtime CI); this artifact is ONE segment
|
| 33 |
+
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
+
# 1-core run for this artifact failed to compile — no figure available for that slice.
|
| 35 |
+
|
| 36 |
+
metrics:
|
| 37 |
+
accuracy: null
|
| 38 |
+
top5_accuracy: null
|
| 39 |
+
|
| 40 |
+
memory:
|
| 41 |
+
peak_mb: null
|
| 42 |
+
|
| 43 |
+
power:
|
| 44 |
+
avg_w: null
|