alexwengg's picture
Card: Swift wrapper is EvokeManager (FluidUse #28)
504abb1 verified
|
Raw History Blame Contribute Delete
3.81 kB
---
license: apache-2.0
library_name: coremltools
pipeline_tag: feature-extraction
base_model: ibm-granite/granite-embedding-30m-sparse
tags:
- coreml
- sparse
- splade
- sparse-encoder
- information-retrieval
- evoke
- apple-silicon
- neural-engine
---
# Granite Embedding 30M Sparse for Core ML
Core ML conversion of IBM's [`ibm-granite/granite-embedding-30m-sparse`](https://huggingface.co/ibm-granite/granite-embedding-30m-sparse)
(revision `ad82b1fd`), the learned sparse encoder behind Intelligent Internet's
[Evoke](https://github.com/Intelligent-Internet/Evoke). It turns text into a short list of weighted vocabulary terms,
including related terms the text never uses ("vaccines for seniors" β†’ vaccine, vaccination, older, elder, …), which go
into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference
converted it.
## Packages
| Package | Input | Output | Size |
| --- | --- | --- | ---: |
| `granite-embedding-30m-sparse-L64-fp16.mlpackage` | `input_ids` int32 `[1, 64]` | `max_logits` fp16 `[1, 192]`, `vocab_ids` int32 `[1, 192]` | 58 MB |
| `granite-embedding-30m-sparse-L128-fp16.mlpackage` | `[1, 128]` | same | 58 MB |
| `granite-embedding-30m-sparse-L256-fp16.mlpackage` | `[1, 256]` | same | 58 MB |
| `granite-embedding-30m-sparse-L512-fp16.mlpackage` | `[1, 512]` | same | 58 MB |
One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on
CPU). macOS 14 / iOS 17 or newer.
- **Input:** RoBERTa byte-level BPE (`tokenizer.json`), `<s>` … `</s>`, right-padded with `<pad>` (id 1) to the
package length. The attention mask is derived from `<pad>` inside the graph.
- **Output:** the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens),
sorted descending.
- **Weights (host side, fp32):** `w = log1p(max(v, 0)) ^ gamma Γ— scale` over the first `active_dims`, keep `w > 0`.
`config.json` holds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192
(gamma 0.5628, scale 1.0). Score = Ξ£ query weight Γ— document weight over shared terms.
## Accuracy
Against Evoke's shipped ONNX compilers (`Intelligent-Internet/Evoke-Model-Beta-1`), NFCorpus test (323 queries,
3,633 documents), sparse dot over the semantic terms only:
| Encoder | nDCG@10 | Recall@100 | Recall@1000 | Same top-10 as ONNX |
| --- | ---: | ---: | ---: | ---: |
| ONNX (shipped, CPU) | 0.3402 | 0.2956 | 0.5854 | β€” |
| Core ML fp16, GPU | 0.3405 | 0.2957 | 0.5856 | 99.6% |
| Core ML fp16, Neural Engine | 0.3405 | 0.2953 | 0.5847 | 98.7% |
An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences
are low-weight terms at the top-k boundary.
## Speed
Apple M5 Pro, macOS 27, one warm call (median of 100), Swift:
| Length | Neural Engine | GPU | CPU |
| --- | ---: | ---: | ---: |
| 64 | **0.74–0.87 ms** | 2.5–3.3 ms | 2.8–3.1 ms |
| 128 | **1.2–1.4 ms** | 1.5–4.3 ms | 4.1 ms |
| 256 | 2.9 ms | **1.9–3.4 ms** | 9.5–15 ms |
| 512 | 7.6 ms | **2.6 ms** | 16–27 ms |
Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask
setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding
3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU.
## Use from Swift
[FluidUse](https://github.com/FluidInference/FluidUse) wraps the packages (`EvokeManager`, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer,
host-side weighting; [PR #28](https://github.com/FluidInference/FluidUse/pull/28)) and includes `EvokeSearchDemo`, an
auto-playing search-as-you-type demo.