Artem Plastinkin commited on
Commit ·
72c3e54
1
Parent(s): bbe369d
Add benchmarking
Browse files
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml
CHANGED
|
@@ -29,7 +29,7 @@ benchmark:
|
|
| 29 |
|
| 30 |
performance:
|
| 31 |
fps: null
|
| 32 |
-
latency:
|
| 33 |
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
# 1-core run for this artifact failed to compile — no figure available for that slice.
|
| 35 |
|
|
|
|
| 29 |
|
| 30 |
performance:
|
| 31 |
fps: null
|
| 32 |
+
latency: 195.34 # synced to mwmx2.2 model_list.html (2026-09-22); was 38.103941
|
| 33 |
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
# 1-core run for this artifact failed to compile — no figure available for that slice.
|
| 35 |
|
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 1
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 497.55 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
|
| 33 |
+
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
+
|
| 35 |
+
metrics:
|
| 36 |
+
accuracy: null
|
| 37 |
+
top5_accuracy: null
|
| 38 |
+
|
| 39 |
+
memory:
|
| 40 |
+
peak_mb: null
|
| 41 |
+
|
| 42 |
+
power:
|
| 43 |
+
avg_w: null
|
| 44 |
+
|
| 45 |
+
# Exact commands verified against the NNAC "Getting Started" chapter. Rendered
|
| 46 |
+
# by the AI-Dashboard in place of the generic placeholder flow — see
|
| 47 |
+
# downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
|
| 48 |
+
# CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
|
| 49 |
+
# README) — these steps reproduce only this segment's latency, not an
|
| 50 |
+
# end-to-end DeepLabV3+ result. The network config below is required for this
|
| 51 |
+
# model — there is no working default.
|
| 52 |
+
reproduce:
|
| 53 |
+
steps:
|
| 54 |
+
- title: Activate the Python environment
|
| 55 |
+
command: >-
|
| 56 |
+
Activate the Python virtual environment that has the `hf` CLI
|
| 57 |
+
(huggingface_hub) and the NNAC toolchain installed, e.g.
|
| 58 |
+
`source nnac_venv/bin/activate` -- path depends on your toolchain install.
|
| 59 |
+
kind: note
|
| 60 |
+
- title: Download the ONNX model and compile config
|
| 61 |
+
command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
|
| 62 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 63 |
+
command: |
|
| 64 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 1 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
|
| 65 |
+
- title: Set up the R-Car X5H board
|
| 66 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 67 |
+
kind: note
|
| 68 |
+
- title: Copy the compiled artifact to the board
|
| 69 |
+
command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 70 |
+
kind: note
|
| 71 |
+
- title: Run on R-Car X5H (single NPU cluster)
|
| 72 |
+
command: |
|
| 73 |
+
cd binary
|
| 74 |
+
./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
|
| 75 |
+
expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
|
| 76 |
+
notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
|
int8/benchmarks/x5h_mwmx_npu_apm80_3core.yaml
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 3
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 278.55 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
|
| 33 |
+
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
+
|
| 35 |
+
metrics:
|
| 36 |
+
accuracy: null
|
| 37 |
+
top5_accuracy: null
|
| 38 |
+
|
| 39 |
+
memory:
|
| 40 |
+
peak_mb: null
|
| 41 |
+
|
| 42 |
+
power:
|
| 43 |
+
avg_w: null
|
| 44 |
+
|
| 45 |
+
# Exact commands verified against the NNAC "Getting Started" chapter. Rendered
|
| 46 |
+
# by the AI-Dashboard in place of the generic placeholder flow — see
|
| 47 |
+
# downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
|
| 48 |
+
# CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
|
| 49 |
+
# README) — these steps reproduce only this segment's latency, not an
|
| 50 |
+
# end-to-end DeepLabV3+ result. The network config below is required for this
|
| 51 |
+
# model — there is no working default.
|
| 52 |
+
reproduce:
|
| 53 |
+
steps:
|
| 54 |
+
- title: Activate the Python environment
|
| 55 |
+
command: >-
|
| 56 |
+
Activate the Python virtual environment that has the `hf` CLI
|
| 57 |
+
(huggingface_hub) and the NNAC toolchain installed, e.g.
|
| 58 |
+
`source nnac_venv/bin/activate` -- path depends on your toolchain install.
|
| 59 |
+
kind: note
|
| 60 |
+
- title: Download the ONNX model and compile config
|
| 61 |
+
command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
|
| 62 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 63 |
+
command: |
|
| 64 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 3 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
|
| 65 |
+
- title: Set up the R-Car X5H board
|
| 66 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 67 |
+
kind: note
|
| 68 |
+
- title: Copy the compiled artifact to the board
|
| 69 |
+
command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 70 |
+
kind: note
|
| 71 |
+
- title: Run on R-Car X5H (single NPU cluster, 3 AI cores)
|
| 72 |
+
command: |
|
| 73 |
+
cd binary
|
| 74 |
+
./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
|
| 75 |
+
expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
|
| 76 |
+
notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
|
int8/benchmarks/x5h_mwmx_npu_apm80_4core.yaml
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 4
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 230.27 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
|
| 33 |
+
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
+
|
| 35 |
+
metrics:
|
| 36 |
+
accuracy: null
|
| 37 |
+
top5_accuracy: null
|
| 38 |
+
|
| 39 |
+
memory:
|
| 40 |
+
peak_mb: null
|
| 41 |
+
|
| 42 |
+
power:
|
| 43 |
+
avg_w: null
|
| 44 |
+
|
| 45 |
+
# Exact commands verified against the NNAC "Getting Started" chapter. Rendered
|
| 46 |
+
# by the AI-Dashboard in place of the generic placeholder flow — see
|
| 47 |
+
# downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
|
| 48 |
+
# CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
|
| 49 |
+
# README) — these steps reproduce only this segment's latency, not an
|
| 50 |
+
# end-to-end DeepLabV3+ result. The network config below is required for this
|
| 51 |
+
# model — there is no working default.
|
| 52 |
+
reproduce:
|
| 53 |
+
steps:
|
| 54 |
+
- title: Activate the Python environment
|
| 55 |
+
command: >-
|
| 56 |
+
Activate the Python virtual environment that has the `hf` CLI
|
| 57 |
+
(huggingface_hub) and the NNAC toolchain installed, e.g.
|
| 58 |
+
`source nnac_venv/bin/activate` -- path depends on your toolchain install.
|
| 59 |
+
kind: note
|
| 60 |
+
- title: Download the ONNX model and compile config
|
| 61 |
+
command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
|
| 62 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 63 |
+
command: |
|
| 64 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 4 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
|
| 65 |
+
- title: Set up the R-Car X5H board
|
| 66 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 67 |
+
kind: note
|
| 68 |
+
- title: Copy the compiled artifact to the board
|
| 69 |
+
command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 70 |
+
kind: note
|
| 71 |
+
- title: Run on R-Car X5H (single NPU cluster, 4 AI cores)
|
| 72 |
+
command: |
|
| 73 |
+
cd binary
|
| 74 |
+
./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
|
| 75 |
+
expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
|
| 76 |
+
notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
|
int8/benchmarks/x5h_mwmx_npu_apm80_6core.yaml
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
hardware:
|
| 2 |
+
vendor: renesas
|
| 3 |
+
chip: rcar-x5h
|
| 4 |
+
cpu: arm-cortex-a720
|
| 5 |
+
npu: npx6-48k
|
| 6 |
+
npu_count: 2
|
| 7 |
+
npu_cores: 12
|
| 8 |
+
npu_default_freq_mhz: 1066
|
| 9 |
+
accelerator:
|
| 10 |
+
- npu
|
| 11 |
+
|
| 12 |
+
runtime:
|
| 13 |
+
engine: mwmx
|
| 14 |
+
toolchain_version: "MWMX SDK v4.35.0"
|
| 15 |
+
format: onnx
|
| 16 |
+
execution_provider: npu
|
| 17 |
+
execution_precision: int8
|
| 18 |
+
|
| 19 |
+
configuration:
|
| 20 |
+
npu_instances: 1
|
| 21 |
+
npu_cores_per_instance: 6
|
| 22 |
+
npu_freq_mhz: 850
|
| 23 |
+
|
| 24 |
+
benchmark:
|
| 25 |
+
type: hil
|
| 26 |
+
parameters:
|
| 27 |
+
batch_size: 1
|
| 28 |
+
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
|
| 29 |
+
|
| 30 |
+
performance:
|
| 31 |
+
fps: null
|
| 32 |
+
latency: 218.57 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
|
| 33 |
+
# of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
|
| 34 |
+
|
| 35 |
+
metrics:
|
| 36 |
+
accuracy: null
|
| 37 |
+
top5_accuracy: null
|
| 38 |
+
|
| 39 |
+
memory:
|
| 40 |
+
peak_mb: null
|
| 41 |
+
|
| 42 |
+
power:
|
| 43 |
+
avg_w: null
|
| 44 |
+
|
| 45 |
+
# Exact commands verified against the NNAC "Getting Started" chapter. Rendered
|
| 46 |
+
# by the AI-Dashboard in place of the generic placeholder flow — see
|
| 47 |
+
# downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
|
| 48 |
+
# CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
|
| 49 |
+
# README) — these steps reproduce only this segment's latency, not an
|
| 50 |
+
# end-to-end DeepLabV3+ result. The network config below is required for this
|
| 51 |
+
# model — there is no working default.
|
| 52 |
+
reproduce:
|
| 53 |
+
steps:
|
| 54 |
+
- title: Activate the Python environment
|
| 55 |
+
command: >-
|
| 56 |
+
Activate the Python virtual environment that has the `hf` CLI
|
| 57 |
+
(huggingface_hub) and the NNAC toolchain installed, e.g.
|
| 58 |
+
`source nnac_venv/bin/activate` -- path depends on your toolchain install.
|
| 59 |
+
kind: note
|
| 60 |
+
- title: Download the ONNX model and compile config
|
| 61 |
+
command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
|
| 62 |
+
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
|
| 63 |
+
command: |
|
| 64 |
+
python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 6 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
|
| 65 |
+
- title: Set up the R-Car X5H board
|
| 66 |
+
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
|
| 67 |
+
kind: note
|
| 68 |
+
- title: Copy the compiled artifact to the board
|
| 69 |
+
command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
|
| 70 |
+
kind: note
|
| 71 |
+
- title: Run on R-Car X5H (single NPU cluster, 6 AI cores)
|
| 72 |
+
command: |
|
| 73 |
+
cd binary
|
| 74 |
+
./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
|
| 75 |
+
expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
|
| 76 |
+
notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
|