msluszniak commited on
Commit
9d1d488
·
verified ·
1 Parent(s): 7245f52

Core ML: emit a (1,520,520) int32 index map instead of (1,21,520,520) logits

Browse files

The argmax now runs inside the graph, built from reduce_max and elementwise ops so it stays on the ANE. 1.05x to 1.63x faster and 7 to 48 percent lower peak memory on an iPhone 16. XNNPACK variants are unchanged.

Files changed (1) hide show
  1. coreml/config.json +3 -2
coreml/config.json CHANGED
@@ -11,6 +11,8 @@
11
  {
12
  "file": "lraspp_mobilenet_v3_large_coreml_fp16.pte",
13
  "precision": "fp16",
 
 
14
  "methods": {
15
  "forward": {
16
  "inputs": [
@@ -28,11 +30,10 @@
28
  {
29
  "shape": [
30
  1,
31
- 21,
32
  520,
33
  520
34
  ],
35
- "dtype": "float32"
36
  }
37
  ]
38
  }
 
11
  {
12
  "file": "lraspp_mobilenet_v3_large_coreml_fp16.pte",
13
  "precision": "fp16",
14
+ "quantized": false,
15
+ "default": true,
16
  "methods": {
17
  "forward": {
18
  "inputs": [
 
30
  {
31
  "shape": [
32
  1,
 
33
  520,
34
  520
35
  ],
36
+ "dtype": "int32"
37
  }
38
  ]
39
  }