LeNet-5 (ONNX) β Renesas X5H
Introduction
This repository hosts LeNet-5, targeting the Renesas R-Car X5H platform for handwritten digit classification inference on the NPX6 NPU.
Source not fully confirmed: the compile-artifact filename in the source benchmark export (
LeNet5_MNIST_quantization_config.json) confirms this is the classic MNIST digit LeNet-5, with a single-channel 28Γ28 input. No specific upstream GitHub/Hugging Face repository could be confidently identified as the exact source of this ONNX file β many generic "LeNet-5 on MNIST" tutorial implementations exist publicly, and none could be confirmed as authoritative for this checkpoint. The Source Model link below is therefore TBD; only the original architecture paper is cited with confidence.
- Model Architecture: LeNet-5 β one of the earliest convolutional neural networks (two convolutional layers with subsampling, followed by fully-connected layers), historically developed for handwritten digit/document recognition
- Source Model: TBD β not confirmed. Architecture reference: LeCun, Bottou, Bengio & Haffner, "Gradient-Based Learning Applied to Document Recognition," Proceedings of the IEEE, 86(11), 2278β2324, 1998.
- Task: Image Classification (MNIST handwritten digits, 10 classes)
- Parameters: ~61.7K (commonly cited figure for the modern ReLU/max-pooling LeNet-5-on-MNIST variant with direct 28Γ28 input; approximate, not an official published figure for this exact checkpoint)
- Note: LeCun's original 1998 paper used a 32Γ32 padded input; this ONNX export instead uses the common modern convention of a direct 28Γ28 MNIST input (consistent with the GF benchmark's recorded resolution of 1Γ1Γ28Γ28).
Deployment Flow
The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time β no separate quantization step is required.
LeNet5_MNIST_..._optimized.onnx (FP32)
β
βββΆ MWMX Runtime βββΆ INT8 auto-cast βββΆ NPX6 NPU
Provided Artifacts
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | β | fp32/LeNet5_MNIST.onnx β FP32 ONNX export |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance CI pipeline).
Benchmark configuration: Single NPU Β· Single AI Core Β· Batch size: 1 Β· Input: 1 Γ 28 Γ 28 (single-channel, MNIST-sized)
| Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|
| MWMX Runtime | INT8 (auto) | X5H Β· 1Γ NPU Β· 1 Core Β· 850 MHz | 0.073368 | Measured |
This is the lowest-latency model in this benchmark sweep by a wide margin, consistent with LeNet-5's small size and single-channel low-resolution input. Only the 1-AI-core slice was run in the source benchmark export β the 12-core slice was skipped, so no 12-core row is reported here.
Accuracy
TBD β not yet measured/published for this repo.
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the model
Download
hf download Renesas/LeNet5-ONNX --repo-type=model --include "fp32/*"
Benchmark Methodology
- HIL runs: Hardware-in-the-loop β measured on physical R-Car X5H silicon via the MWMX
runtime (
metawaremx_runtimeCI pipeline, "APM50" ship-performance target) - Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- Slices: only the 1 AI core slice was run for this model; the 12-core slice was skipped in the source export