Core ML: emit a (1,520,520) int32 index map instead of (1,21,520,520) logits
Browse filesThe argmax now runs inside the graph, built from reduce_max and elementwise ops so it stays on the ANE. 1.05x to 1.63x faster and 7 to 48 percent lower peak memory on an iPhone 16. XNNPACK variants are unchanged.
coreml/lraspp_mobilenet_v3_large_coreml_fp16.pte
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:365b2961edb1771bdd185d301bac1987f9c828fe3f804a72d1d66f867173cf43
|
| 3 |
+
size 6917692
|