ViT-Base-P16 (ONNX) – Renesas X5H
Introduction
This repository hosts ViT-Base/16 (Vision Transformer, 16×16 patches), targeting the Renesas R-Car X5H platform for image classification inference on the NPX6 NPU.
- Model Architecture: ViT-Base/16 — a Vision Transformer backbone, pretrained with MAE (Masked Autoencoder) self-supervision and then fine-tuned for ImageNet-1k classification.
- Source Model: OpenMMLab config
vit_base_p16_32xb128_mae_in1k(no HuggingFace mirror of these weights; seemodel.sourcein.metadata.yaml) - Task: Image Classification (ImageNet-1k, 1000 classes)
- Parameters: 86M
Note: A
vit_tinyvariant also exists in the source benchmark data but produced zero passing compile/execute runs in the APM50 CI pipeline (both slices failed). It is intentionally not included in this repository — only the ViT-Base/16 checkpoint, which has passing benchmark numbers, is published here.
Deployment Flow
The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.
vit_base_p16_..._optimized.onnx (FP32)
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
Provided Artifacts
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | ✅ | fp32/vit-base-p16_32xb128-mae_in1k.onnx — FP32 ONNX export |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance CI pipeline).
Benchmark configuration: Single NPU · Batch size: 1 · Input resolution: not available from source data (TBD)
| Parameters | Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|---|
| 86M | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 25.796629 | Measured |
| 86M | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 25.923109 | Measured (2026-09-16) |
| 86M | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 6.873014 | Measured |
Accuracy
TBD — not yet measured/published for this repo.
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the model
Download
hf download Renesas/ViT-Base-P16-ONNX --repo-type=model --include "fp32/*"
Benchmark Methodology
- HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
runtime (
metawaremx_runtimeCI pipeline, "APM50" ship-performance target) - Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- Slices: results reported for both 1 AI core and 12 AI cores per NPU instance
- Excluded: the
vit_tinyvariant from the same source data batch had 0 passing runs across both slices and is not represented in this repository