README: ios/ is the JIT .aimodel, ios-h18p/ the h18p bundle; iPhone 18 Pro load numbers
Browse files
README.md
CHANGED
|
@@ -31,7 +31,7 @@ Zoo card, recipe and gate transcript: [coreai-model-zoo/models/granite-embedding
|
|
| 31 |
|
| 32 |
IBM's 97M-parameter **multilingual text embedder** β a ModernBERT encoder, 384-d CLS-pooled
|
| 33 |
unit vectors, Japanese and English among its languages β as a static `.aimodel` for macOS 27
|
| 34 |
-
and, ahead
|
| 35 |
[`ibm-granite/granite-embedding-97m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2)
|
| 36 |
(Apache-2.0, revision `835ad1408β¦`) is the **smallest embedder in this catalog** (390 MB fp32,
|
| 37 |
against 1.2 GB for EmbeddingGemma-300m and 1.1 GB for Qwen3-Embedding-0.6B) and its **first
|
|
@@ -107,6 +107,25 @@ The Mac h16c AOT twin also passed 70/70 (same numerics), but its timings were ta
|
|
| 107 |
lane's GPU evaluation running and are not reported. The w8 variant on Mac is gated on **CPU
|
| 108 |
only** (min cosine 0.999410, max |err| 5.67e-3, ranking exact); Mac GPU for w8 was not run.
|
| 109 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 110 |
The fixed grid computes every position, so pick the smallest grid that covers the text: S=128
|
| 111 |
for queries and short notes, S=512 for passages. **fp32 is the default.** w8 is a storage
|
| 112 |
option only β 22% smaller, not faster here β because the 180,000Γ384 fp32 vocabulary table is
|
|
@@ -126,7 +145,8 @@ matched to the wrong text) must FAIL.
|
|
| 126 |
**63 instead of 64** passes the embedding gate (cos 0.99995) and fails only the layer gate β
|
| 127 |
which is why the layer gate exists. Whole-model **fp16 fails** this layer gate on both grids.
|
| 128 |
- **Export**: the torch-exported, decomposed graph is gated before conversion, on both grids.
|
| 129 |
-
- **Runtime**: Mac CPU and GPU (JIT), Mac h16c AOT, iPhone h18p AOT β the tables above
|
|
|
|
| 130 |
- **w8**: the same gate at prepared, finalized and decomposed stages, 48 `lut_to_dense` ops
|
| 131 |
counted, palettes hashed; the iOS w8 export reuses the Mac palettes byte for byte.
|
| 132 |
|
|
@@ -145,17 +165,24 @@ platform β folder.
|
|
| 145 |
|---|---|---|---|---:|
|
| 146 |
| `macos/fp32-s512/` **(default)** | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
|
| 147 |
| `macos/fp32-s128/` | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
|
| 148 |
-
| `ios/fp32-s512/` **(default)** | iOS 27
|
| 149 |
-
| `ios/fp32-s128/` | iOS 27
|
| 150 |
| `macos/w8-fp32table-s512/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
|
| 151 |
| `macos/w8-fp32table-s128/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
|
| 152 |
-
| `ios/w8-fp32table-s512/` | iOS 27
|
| 153 |
-
| `ios/w8-fp32table-s128/` | iOS 27
|
| 154 |
-
|
| 155 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 156 |
`xcrun coreai-build compile --platform iOS --min-deployment-version 27.0 --preferred-compute gpu
|
| 157 |
-
--architecture h18p` (coreai-build 3600.83.1)
|
| 158 |
-
|
|
|
|
| 159 |
|
| 160 |
Convert yourself: [`conversion/granite_embedding/`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/README.md)
|
| 161 |
β five staged scripts; [`recipe.toml`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/recipe.toml) names the commands.
|
|
@@ -184,6 +211,7 @@ epsilon because the converter's `F.normalize` decomposition drops it.
|
|
| 184 |
|
| 185 |
Apache-2.0 at the pinned upstream revision; this repo carries IBM's unmodified card as
|
| 186 |
`UPSTREAM_README.md` and a `LICENSE-NOTE.md` listing the changes (static graph, in-graph
|
| 187 |
-
pooling, optional w8 palettes, h18p compile). Not tested:
|
| 188 |
-
|
|
|
|
| 189 |
fixtures, retrieval quality on a benchmark, sustained thermals, true cache-cold load.
|
|
|
|
| 31 |
|
| 32 |
IBM's 97M-parameter **multilingual text embedder** β a ModernBERT encoder, 384-d CLS-pooled
|
| 33 |
unit vectors, Japanese and English among its languages β as a static `.aimodel` for macOS 27
|
| 34 |
+
and iOS 27, with bundles compiled ahead of time for the iPhone 17 Pro beside it.
|
| 35 |
[`ibm-granite/granite-embedding-97m-multilingual-r2`](https://huggingface.co/ibm-granite/granite-embedding-97m-multilingual-r2)
|
| 36 |
(Apache-2.0, revision `835ad1408β¦`) is the **smallest embedder in this catalog** (390 MB fp32,
|
| 37 |
against 1.2 GB for EmbeddingGemma-300m and 1.1 GB for Qwen3-Embedding-0.6B) and its **first
|
|
|
|
| 107 |
lane's GPU evaluation running and are not reported. The w8 variant on Mac is gated on **CPU
|
| 108 |
only** (min cosine 0.999410, max |err| 5.67e-3, ranking exact); Mac GPU for w8 was not run.
|
| 109 |
|
| 110 |
+
**iPhone 18 Pro** (iPhone19,2, h19p), iOS 27.0 build **24A437**, the JIT `.aimodel` bundles in `ios/`
|
| 111 |
+
(the same files as `macos/`), loaded 2026-09-26 by the zoo's DecideGate app in its load-only mode,
|
| 112 |
+
GPU-preferred, without the increased-memory entitlement. Each first load was the first after a fresh
|
| 113 |
+
install of the app. The call is one run on all-zero inputs. Each first load wrote a specialization of
|
| 114 |
+
about the bundle's size into the app container, and the load after a relaunch reused it. One measurement
|
| 115 |
+
per graph
|
| 116 |
+
([knowledge/jit-distribution.md](https://github.com/john-rocky/coreai-model-zoo/blob/main/knowledge/jit-distribution.md)).
|
| 117 |
+
|
| 118 |
+
| JIT bundle in `ios/` | MB | first load | first call | load after relaunch |
|
| 119 |
+
|---|---:|---:|---:|---:|
|
| 120 |
+
| `fp32-s128/granite97m_fp32_s128_bound.aimodel` | 390 | 1.01 s | 619 ms | 0.26 s |
|
| 121 |
+
| `fp32-s512/granite97m_fp32_s512_bound.aimodel` | 390 | 0.40 s | 88 ms | 0.26 s |
|
| 122 |
+
| `w8-fp32table-s128/granite97m_w8_fp32table_s128.aimodel` | 305 | 0.63 s | 132 ms | 0.21 s |
|
| 123 |
+
| `w8-fp32table-s512/granite97m_w8_fp32table_s512.aimodel` | 306 | 0.33 s | 93 ms | 0.21 s |
|
| 124 |
+
|
| 125 |
+
The iPhone gate above (the iPhone 17 Pro rows) ran the `h18p` export, now in `ios-h18p/`. The JIT IR in
|
| 126 |
+
`ios/` was only loaded and called once on the iPhone 18 Pro. That call returned a 384-value float32
|
| 127 |
+
embedding with no non-finite values; its parity with the reference was not re-measured.
|
| 128 |
+
|
| 129 |
The fixed grid computes every position, so pick the smallest grid that covers the text: S=128
|
| 130 |
for queries and short notes, S=512 for passages. **fp32 is the default.** w8 is a storage
|
| 131 |
option only β 22% smaller, not faster here β because the 180,000Γ384 fp32 vocabulary table is
|
|
|
|
| 145 |
**63 instead of 64** passes the embedding gate (cos 0.99995) and fails only the layer gate β
|
| 146 |
which is why the layer gate exists. Whole-model **fp16 fails** this layer gate on both grids.
|
| 147 |
- **Export**: the torch-exported, decomposed graph is gated before conversion, on both grids.
|
| 148 |
+
- **Runtime**: Mac CPU and GPU (JIT), Mac h16c AOT, iPhone h18p AOT β the tables above; iPhone 18 Pro
|
| 149 |
+
JIT: load and one zero-input call only.
|
| 150 |
- **w8**: the same gate at prepared, finalized and decomposed stages, 48 `lut_to_dense` ops
|
| 151 |
counted, palettes hashed; the iOS w8 export reuses the Mac palettes byte for byte.
|
| 152 |
|
|
|
|
| 165 |
|---|---|---|---|---:|
|
| 166 |
| `macos/fp32-s512/` **(default)** | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
|
| 167 |
| `macos/fp32-s128/` | macOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
|
| 168 |
+
| `ios/fp32-s512/` **(default)** | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s512_bound.aimodel` | 390,431,506 |
|
| 169 |
+
| `ios/fp32-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_fp32_s128_bound.aimodel` | 389,989,146 |
|
| 170 |
| `macos/w8-fp32table-s512/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
|
| 171 |
| `macos/w8-fp32table-s128/` | macOS 27 (CPU-gated) | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
|
| 172 |
+
| `ios/w8-fp32table-s512/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s512.aimodel` | 305,569,358 |
|
| 173 |
+
| `ios/w8-fp32table-s128/` | iOS 27 | JIT `.aimodel` | `granite97m_w8_fp32table_s128.aimodel` | 305,126,985 |
|
| 174 |
+
| `ios-h18p/fp32-s512/` | iOS 27, **h18p only** | AOT `.aimodelc` | `granite97m_fp32_s512_bound.h18p.aimodelc` | 390,308,788 |
|
| 175 |
+
| `ios-h18p/fp32-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_fp32_s128_bound.h18p.aimodelc` | 390,081,410 |
|
| 176 |
+
| `ios-h18p/w8-fp32table-s512/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s512_r02.h18p.aimodelc` | 305,479,184 |
|
| 177 |
+
| `ios-h18p/w8-fp32table-s128/` | iOS 27, h18p only | AOT `.aimodelc` | `granite97m_w8_fp32table_s128_r02.h18p.aimodelc` | 305,251,774 |
|
| 178 |
+
|
| 179 |
+
`ios/` holds the same JIT bundles as `macos/`; every iPhone generation specializes them on its first
|
| 180 |
+
load. The `ios-h18p/` bundles moved there from `ios/` in revision `a27dc73e` (2026-09-26). They are
|
| 181 |
+
compiled for one device architecture (`h18p`, the iPhone 17 Pro) with
|
| 182 |
`xcrun coreai-build compile --platform iOS --min-deployment-version 27.0 --preferred-compute gpu
|
| 183 |
+
--architecture h18p` (coreai-build 3600.83.1), from a separate iOS export whose IR is reproducible
|
| 184 |
+
from the recipe but not shipped. The iPhone 18 Pro refuses an h18p bundle with
|
| 185 |
+
`incompatibleCompiledAssetArchitecture`. **Never load an `ios-h18p/` bundle on a Mac.**
|
| 186 |
|
| 187 |
Convert yourself: [`conversion/granite_embedding/`](https://github.com/john-rocky/coreai-model-zoo/blob/main/conversion/granite_embedding/README.md)
|
| 188 |
β five staged scripts; [`recipe.toml`](https://github.com/john-rocky/coreai-model-zoo/blob/main/models/granite-embedding-97m/recipe.toml) names the commands.
|
|
|
|
| 211 |
|
| 212 |
Apache-2.0 at the pinned upstream revision; this repo carries IBM's unmodified card as
|
| 213 |
`UPSTREAM_README.md` and a `LICENSE-NOTE.md` listing the changes (static graph, in-graph
|
| 214 |
+
pooling, optional w8 palettes, h18p compile). Not tested: phones other than the iPhone 17 Pro (the
|
| 215 |
+
h18p gate) and the iPhone 18 Pro (JIT load and one call), other OS builds, the JIT bundles' embedding
|
| 216 |
+
parity on an iPhone, the Mac GPU with w8, the Neural Engine, dynamic or batched shapes, S > 512, languages beyond the JA/EN
|
| 217 |
fixtures, retrieval quality on a benchmark, sustained thermals, true cache-cold load.
|