Artem Plastinkin
Add model
41edfa0
|
Raw History Blame Contribute Delete
3.87 kB
---
license: apache-2.0
base_model: []
pipeline_tag: image-segmentation
tags:
- semantic-segmentation
- computer-vision
- renesas
- x5h
- onnx
- deeplabv3plus
- resnet50
- cityscapes
---
# DeepLabV3Plus-R50 (ONNX) – Renesas X5H
> **⚠️ Partial-model caveat.** The source checkpoint name ends in
> `custom_seg_split_4_split_2` β€” this artifact is **one segment of a 4-way split network**, not
> the full end-to-end DeepLabV3+ model. The latency below reflects only that segment; do not
> quote it as whole-model latency until the other splits are accounted for.
## Introduction
This repository hosts **DeepLabV3+ (ResNet50-D8 backbone)** targeting the **Renesas R-Car X5H**
platform for semantic segmentation inference on the NPX6 NPU.
- **Model Architecture:** DeepLabV3+ with ResNet50-D8 backbone
- **Source Model:** OpenMMLab config [`deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024`](https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml)
*(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)*
- **Task:** Semantic Segmentation
- **Dataset:** Cityscapes (inferred from checkpoint name)
- **Input Resolution:** 512 Γ— 1024 (explicit in checkpoint name)
- **Parameters:** not published β€” count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`)
## Deployment Flow
The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time β€” no
separate quantization step is required.
```
deeplabv3plus_r50_oss_sim_inf.onnx (FP32, split_2 segment)
β”‚
└─▢ MWMX Runtime ──▢ INT8 auto-cast ──▢ NPX6 NPU
```
## Provided Artifacts
| Artifact | Status | Notes |
|----------|--------|-------|
| **FP32 (ONNX)** | βœ… Published | `fp32/deeplabv3plus_r50_oss_sim_inf.onnx` β€” segment `split_2` of 4; auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped |
## Performance
Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline).
> **Benchmark configuration:** Single NPU Β· Batch size: 1 Β· Input: 3 Γ— 512 Γ— 1024
>
> The 1-AI-core slice failed to compile in the source CI pipeline, so only the 12-core result
> is available. Latency is for the `split_2` segment only (see caveat above).
| Runtime | Precision | Device | Latency (ms) | Type |
|---------|-----------|--------|---------------|------|
| MWMX Runtime | INT8 (auto) | X5H Β· 1Γ— NPU Β· 12 Cores Β· 850 MHz | 38.104 | Measured |
> Reconfirmed: 1-core compile still fails as of the 2026-09-16 benchmark run.
### Accuracy (mIoU)
TBD β€” not yet measured/published for this repo.
---
## Runtime Details
### MWMX Runtime
- **Engine:** Renesas MWMX (Middleware MX) native inference runtime
- **Input format:** FP32 ONNX (compiled by the MWMX toolchain)
- **NPU execution precision:** INT8 (auto-cast by MWMX toolchain)
- **Execution target:** NPX6-48K NPU on R-Car X5H
---
## Prerequisites
To run inference on Renesas R-Car X5H, you need:
1. **Renesas R-Car X5H board** with NPX6 NPU
2. **Renesas MWMX Runtime**
3. **Hugging Face CLI** to download the model
## Download
```bash
hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type=model --include "fp32/*"
```
---
## Benchmark Methodology
- **HIL runs:** Hardware-in-the-loop β€” measured on physical R-Car X5H silicon via the MWMX
runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target)
- **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- **Slices:** only the 12-AI-core result is available (1-core compile failed)
- **Scope:** this artifact is a single segment (`split_2` of 4) of the full segmentation pipeline