--- license: apache-2.0 base_model: - facebook/detr-resnet-50 pipeline_tag: object-detection tags: - object-detection - computer-vision - renesas - x5h - onnx - detr - transformer - detection --- # DETR-R50 (ONNX) – Renesas X5H ## Introduction This repository hosts **DETR (DEtection TRansformer)**, targeting the **Renesas R-Car X5H** platform for object detection inference on the NPX6 NPU. - **Model Architecture:** DETR — transformer-based, end-to-end object detector using set prediction (bipartite matching), CNN backbone + transformer encoder/decoder - **Source Model:** [facebook/detr-resnet-50](https://huggingface.co/facebook/detr-resnet-50) — checkpoint [`sim_mod_detr`](https://huggingface.co/facebook/detr-resnet-50) - **Task:** Object Detection - **Dataset:** COCO (inferred from checkpoint name) - **Input Resolution:** TBD ## Deployment Flow The FP32 ONNX model is auto-cast to **INT8** by the Renesas MWMX toolchain at compile time — no separate quantization step is required. ``` sim_mod_detr_..._optimized.onnx (FP32) │ └─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU ``` ## Provided Artifacts | Artifact | Status | Notes | |----------|--------|-------| | **FP32 (ONNX)** | ✅ | `fp32/sim_mod_detr.onnx` — FP32 ONNX export | ## Performance Measured on **Renesas R-Car X5H** via the MWMX runtime (APM50 ship-performance CI pipeline). > **Benchmark configuration:** Single NPU · Batch size: 1 · Input resolution: TBD | Runtime | Precision | Device | Latency (ms) | Type | |---------|-----------|--------|---------------|------| | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 147.768855 | Measured | | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 92.763967 | Measured | --- ## Model Input ### Input Tensor - Shape: TBD — not available from source data (expected `(N, 3, H, W)`, RGB) - Format: TBD - Data Type: TBD - Pixel Range: TBD ### Preprocessing TBD — not available from source data. ## Model Outputs TBD — not available from source data. DETR produces a fixed-size set of predictions (typically 100 object queries), each with a class-probability distribution (including a "no object" class) and a normalized bounding box, directly via set prediction — no anchor decoding or NMS is required by design. ### Postprocessing 1. Per-query class-probability argmax (excluding "no object") 2. Confidence threshold filtering 3. Box de-normalization to image coordinates ### Accuracy TBD — not yet measured/published for this repo. --- ## Runtime Details ### MWMX Runtime - **Engine:** Renesas MWMX (Middleware MX) native inference runtime - **Input format:** FP32 ONNX (compiled by the MWMX toolchain) - **NPU execution precision:** INT8 (auto-cast by MWMX toolchain) - **Execution target:** NPX6-48K NPU on R-Car X5H --- ## Prerequisites To run inference on Renesas R-Car X5H, you need: 1. **Renesas R-Car X5H board** with NPX6 NPU 2. **Renesas MWMX Runtime** 3. **Hugging Face CLI** to download the model ## Download ```bash hf download Renesas/DETR-R50-ONNX --repo-type=model --include "fp32/*" ``` --- ## Benchmark Methodology - **HIL runs:** Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX runtime (`metawaremx_runtime` CI pipeline, "APM50" ship-performance target) - **Precision:** FP32 ONNX input; INT8 execution (auto-cast by MWMX) - **Slices:** results reported for both 1 AI core and 12 AI cores per NPU instance