File size: 3,873 Bytes
aeea7fa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41edfa0
aeea7fa
 
 
 
 
 
 
 
41edfa0
aeea7fa
 
 
 
 
 
 
 
 
 
 
 
 
 
41edfa0
 
aeea7fa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
41edfa0
aeea7fa
 
 
41edfa0
 
 
aeea7fa
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
---

license: apache-2.0
base_model: []
pipeline_tag: image-segmentation
tags:
- semantic-segmentation
- computer-vision
- renesas
- x5h
- onnx
- deeplabv3plus
- resnet50
- cityscapes
---


# DeepLabV3Plus-R50 (ONNX) – Renesas X5H

> **⚠️ Partial-model caveat.** The source checkpoint name ends in
> `custom_seg_split_4_split_2` — this artifact is **one segment of a 4-way split network**, not

> the full end-to-end DeepLabV3+ model. The latency below reflects only that segment; do not

> quote it as whole-model latency until the other splits are accounted for.



## Introduction



This repository hosts **DeepLabV3+ (ResNet50-D8 backbone)** targeting the **Renesas R-Car X5H**

platform for semantic segmentation inference on the NPX6 NPU.



- **Model Architecture:** DeepLabV3+ with ResNet50-D8 backbone

- **Source Model:** OpenMMLab config [`deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024`](https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml)

  *(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*

- **Task:** Semantic Segmentation

- **Dataset:** Cityscapes (inferred from checkpoint name)

- **Input Resolution:** 512 × 1024 (explicit in checkpoint name)

- **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)



## Deployment Flow



The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no

separate quantization step is required.



```

deeplabv3plus_r50_oss_sim_inf.onnx (FP32, split_2 segment)
        │

        └─▶  MWMX Runtime  ──▶  INT8 auto-cast  ──▶  NPX6 NPU

```


## Provided Artifacts

| Artifact | Status | Notes |
|----------|--------|-------|
| **FP32 (ONNX)** | ✅ Published | `fp32/deeplabv3plus_r50_oss_sim_inf.onnx` — segment `split_2` of 4; auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped |

## Performance

Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).

> **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 1024
>

> The 1-AI-core slice failed to compile in the source CI pipeline, so only the 12-core result
> is available. Latency is for the `split_2` segment only (see caveat above).



| Runtime | Precision | Device | Latency (ms) | Type |

|---------|-----------|--------|---------------|------|

| MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 38.104 | Measured |



> Reconfirmed: 1-core compile still fails as of the 2026-09-16 benchmark run.



### Accuracy (mIoU)



TBD — not yet measured/published for this repo.



---



## Runtime Details



### MWMX Runtime



- **Engine:** Renesas MWMX (Middleware MX) native inference runtime

- **Input format:** FP32 ONNX (compiled by the MWMX toolchain)

- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)

- **Execution target:** NPX6-48K NPU on R-Car X5H



---



## Prerequisites



To run inference on Renesas R-Car X5H, you need:



1. **Renesas R-Car X5H board** with NPX6 NPU

2. **Renesas MWMX Runtime**

3. **Hugging Face CLI** to download the model



## Download



```bash

hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type=model --include "fp32/*"

```



---



## Benchmark Methodology



- **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX

  runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- **Slices:** only the 12-AI-core result is available (1-core compile failed)
- **Scope:** this artifact is a single segment (`split_2` of 4) of the full segmentation pipeline