--- license: apache-2.0 base_model: [] pipeline_tag: image-segmentation tags: - semantic-segmentation - computer-vision - renesas - x5h - onnx - deeplabv3plus - resnet50 - cityscapes --- # DeepLabV3Plus-R50 (ONNX) – Renesas X5H > **⚠️ Partial-model caveat.** The source checkpoint name ends in > `custom_seg_split_4_split_2` — this artifact is **one segment of a 4-way split network**, not > the full end-to-end DeepLabV3+ model. The latency below reflects only that segment; do not > quote it as whole-model latency until the other splits are accounted for. ## Introduction This repository hosts **DeepLabV3+ (ResNet50-D8 backbone)** targeting the **Renesas R-Car X5H** platform for semantic segmentation inference on the NPX6 NPU. - **Model Architecture:** DeepLabV3+ with ResNet50-D8 backbone - **Source Model:** OpenMMLab config [`deeplabv3plus_r50_d8_4xb2_80k_cityscapes_512x1024`](https://github.com/open-mmlab/mmsegmentation/blob/main/configs/deeplabv3plus/metafile.yaml) *(no HuggingFace mirror of these weights; see `model.source` in `.metadata.yaml`)* - **Task:** Semantic Segmentation - **Dataset:** Cityscapes (inferred from checkpoint name) - **Input Resolution:** 512 × 1024 (explicit in checkpoint name) - **Parameters:** not published — count them from the ONNX graph (`sum(numpy_helper.to_array(t).size for t in model.graph.initializer)`) ## Deployment Flow The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no separate quantization step is required. ``` deeplabv3plus_r50_oss_sim_inf.onnx (FP32, split_2 segment) │ └─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU ``` ## Provided Artifacts | Artifact | Status | Notes | |----------|--------|-------| | **FP32 (ONNX)** | ✅ Published | `fp32/deeplabv3plus_r50_oss_sim_inf.onnx` — segment `split_2` of 4; auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file is shipped | ## Performance Measured on **Renesas R-Car X5H** via the MWMX runtime (APM80 ship-performance CI pipeline). > **Benchmark configuration:** Single NPU · Batch size: 1 · Input: 3 × 512 × 1024 > > The 1-AI-core slice failed to compile in the source CI pipeline, so only the 12-core result > is available. Latency is for the `split_2` segment only (see caveat above). | Runtime | Precision | Device | Latency (ms) | Type | |---------|-----------|--------|---------------|------| | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 38.104 | Measured | > Reconfirmed: 1-core compile still fails as of the 2026-09-16 benchmark run. ### Accuracy (mIoU) TBD — not yet measured/published for this repo. --- ## Runtime Details ### MWMX Runtime - **Engine:** Renesas MWMX (Middleware MX) native inference runtime - **Input format:** FP32 ONNX (compiled by the MWMX toolchain) - **NPU execution precision:** INT8 (auto-cast by MWMX toolchain) - **Execution target:** NPX6-48K NPU on R-Car X5H --- ## Prerequisites To run inference on Renesas R-Car X5H, you need: 1. **Renesas R-Car X5H board** with NPX6 NPU 2. **Renesas MWMX Runtime** 3. **Hugging Face CLI** to download the model ## Download ```bash hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type=model --include "fp32/*" ``` --- ## Benchmark Methodology - **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX runtime (`metawaremx_runtime` CI pipeline, "APM80" ship-performance target) - **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX) - **Slices:** only the 12-AI-core result is available (1-core compile failed) - **Scope:** this artifact is a single segment (`split_2` of 4) of the full segmentation pipeline