Artem Plastinkin commited on
Commit
72c3e54
·
1 Parent(s): bbe369d

Add benchmarking

Browse files
int8/benchmarks/x5h_mwmx_npu_apm80_12core.yaml CHANGED
@@ -29,7 +29,7 @@ benchmark:
29
 
30
  performance:
31
  fps: null
32
- latency: 38.103941 # APM80 pipeline run (metawaremx_runtime CI); this artifact is ONE segment
33
  # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
  # 1-core run for this artifact failed to compile — no figure available for that slice.
35
 
 
29
 
30
  performance:
31
  fps: null
32
+ latency: 195.34 # synced to mwmx2.2 model_list.html (2026-09-22); was 38.103941
33
  # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
  # 1-core run for this artifact failed to compile — no figure available for that slice.
35
 
int8/benchmarks/x5h_mwmx_npu_apm80_1core.yaml ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 1
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 497.55 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
33
+ # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
+
35
+ metrics:
36
+ accuracy: null
37
+ top5_accuracy: null
38
+
39
+ memory:
40
+ peak_mb: null
41
+
42
+ power:
43
+ avg_w: null
44
+
45
+ # Exact commands verified against the NNAC "Getting Started" chapter. Rendered
46
+ # by the AI-Dashboard in place of the generic placeholder flow — see
47
+ # downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
48
+ # CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
49
+ # README) — these steps reproduce only this segment's latency, not an
50
+ # end-to-end DeepLabV3+ result. The network config below is required for this
51
+ # model — there is no working default.
52
+ reproduce:
53
+ steps:
54
+ - title: Activate the Python environment
55
+ command: >-
56
+ Activate the Python virtual environment that has the `hf` CLI
57
+ (huggingface_hub) and the NNAC toolchain installed, e.g.
58
+ `source nnac_venv/bin/activate` -- path depends on your toolchain install.
59
+ kind: note
60
+ - title: Download the ONNX model and compile config
61
+ command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
62
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
63
+ command: |
64
+ python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 1 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
65
+ - title: Set up the R-Car X5H board
66
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
67
+ kind: note
68
+ - title: Copy the compiled artifact to the board
69
+ command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
70
+ kind: note
71
+ - title: Run on R-Car X5H (single NPU cluster)
72
+ command: |
73
+ cd binary
74
+ ./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
75
+ expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
76
+ notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
int8/benchmarks/x5h_mwmx_npu_apm80_3core.yaml ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 3
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 278.55 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
33
+ # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
+
35
+ metrics:
36
+ accuracy: null
37
+ top5_accuracy: null
38
+
39
+ memory:
40
+ peak_mb: null
41
+
42
+ power:
43
+ avg_w: null
44
+
45
+ # Exact commands verified against the NNAC "Getting Started" chapter. Rendered
46
+ # by the AI-Dashboard in place of the generic placeholder flow — see
47
+ # downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
48
+ # CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
49
+ # README) — these steps reproduce only this segment's latency, not an
50
+ # end-to-end DeepLabV3+ result. The network config below is required for this
51
+ # model — there is no working default.
52
+ reproduce:
53
+ steps:
54
+ - title: Activate the Python environment
55
+ command: >-
56
+ Activate the Python virtual environment that has the `hf` CLI
57
+ (huggingface_hub) and the NNAC toolchain installed, e.g.
58
+ `source nnac_venv/bin/activate` -- path depends on your toolchain install.
59
+ kind: note
60
+ - title: Download the ONNX model and compile config
61
+ command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
62
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
63
+ command: |
64
+ python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 3 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
65
+ - title: Set up the R-Car X5H board
66
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
67
+ kind: note
68
+ - title: Copy the compiled artifact to the board
69
+ command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
70
+ kind: note
71
+ - title: Run on R-Car X5H (single NPU cluster, 3 AI cores)
72
+ command: |
73
+ cd binary
74
+ ./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
75
+ expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
76
+ notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
int8/benchmarks/x5h_mwmx_npu_apm80_4core.yaml ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 4
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 230.27 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
33
+ # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
+
35
+ metrics:
36
+ accuracy: null
37
+ top5_accuracy: null
38
+
39
+ memory:
40
+ peak_mb: null
41
+
42
+ power:
43
+ avg_w: null
44
+
45
+ # Exact commands verified against the NNAC "Getting Started" chapter. Rendered
46
+ # by the AI-Dashboard in place of the generic placeholder flow — see
47
+ # downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
48
+ # CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
49
+ # README) — these steps reproduce only this segment's latency, not an
50
+ # end-to-end DeepLabV3+ result. The network config below is required for this
51
+ # model — there is no working default.
52
+ reproduce:
53
+ steps:
54
+ - title: Activate the Python environment
55
+ command: >-
56
+ Activate the Python virtual environment that has the `hf` CLI
57
+ (huggingface_hub) and the NNAC toolchain installed, e.g.
58
+ `source nnac_venv/bin/activate` -- path depends on your toolchain install.
59
+ kind: note
60
+ - title: Download the ONNX model and compile config
61
+ command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
62
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
63
+ command: |
64
+ python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 4 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
65
+ - title: Set up the R-Car X5H board
66
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
67
+ kind: note
68
+ - title: Copy the compiled artifact to the board
69
+ command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
70
+ kind: note
71
+ - title: Run on R-Car X5H (single NPU cluster, 4 AI cores)
72
+ command: |
73
+ cd binary
74
+ ./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
75
+ expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
76
+ notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.
int8/benchmarks/x5h_mwmx_npu_apm80_6core.yaml ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ hardware:
2
+ vendor: renesas
3
+ chip: rcar-x5h
4
+ cpu: arm-cortex-a720
5
+ npu: npx6-48k
6
+ npu_count: 2
7
+ npu_cores: 12
8
+ npu_default_freq_mhz: 1066
9
+ accelerator:
10
+ - npu
11
+
12
+ runtime:
13
+ engine: mwmx
14
+ toolchain_version: "MWMX SDK v4.35.0"
15
+ format: onnx
16
+ execution_provider: npu
17
+ execution_precision: int8
18
+
19
+ configuration:
20
+ npu_instances: 1
21
+ npu_cores_per_instance: 6
22
+ npu_freq_mhz: 850
23
+
24
+ benchmark:
25
+ type: hil
26
+ parameters:
27
+ batch_size: 1
28
+ input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
29
+
30
+ performance:
31
+ fps: null
32
+ latency: 218.57 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
33
+ # of a 4-way split model ("custom_seg_split_4_split_2") — not full end-to-end latency.
34
+
35
+ metrics:
36
+ accuracy: null
37
+ top5_accuracy: null
38
+
39
+ memory:
40
+ peak_mb: null
41
+
42
+ power:
43
+ avg_w: null
44
+
45
+ # Exact commands verified against the NNAC "Getting Started" chapter. Rendered
46
+ # by the AI-Dashboard in place of the generic placeholder flow — see
47
+ # downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
48
+ # CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
49
+ # README) — these steps reproduce only this segment's latency, not an
50
+ # end-to-end DeepLabV3+ result. The network config below is required for this
51
+ # model — there is no working default.
52
+ reproduce:
53
+ steps:
54
+ - title: Activate the Python environment
55
+ command: >-
56
+ Activate the Python virtual environment that has the `hf` CLI
57
+ (huggingface_hub) and the NNAC toolchain installed, e.g.
58
+ `source nnac_venv/bin/activate` -- path depends on your toolchain install.
59
+ kind: note
60
+ - title: Download the ONNX model and compile config
61
+ command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
62
+ - title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
63
+ command: |
64
+ python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 6 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
65
+ - title: Set up the R-Car X5H board
66
+ command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
67
+ kind: note
68
+ - title: Copy the compiled artifact to the board
69
+ command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
70
+ kind: note
71
+ - title: Run on R-Car X5H (single NPU cluster, 6 AI cores)
72
+ command: |
73
+ cd binary
74
+ ./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
75
+ expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
76
+ notes: This is one segment (split_2 of 4) of the full segmentation pipeline — not whole-model latency.