|
Download README.md from FluidInference/granite-embedding-30m-sparse-coreml: direct link, hf CLI and curl.
- Browser
- Download file 3.81 kB
-
https://huggingface.co/FluidInference/granite-embedding-30m-sparse-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/granite-embedding-30m-sparse-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/granite-embedding-30m-sparse-coreml/resolve/main/README.md
3.81 kB
| license: apache-2.0 | |
| library_name: coremltools | |
| pipeline_tag: feature-extraction | |
| base_model: ibm-granite/granite-embedding-30m-sparse | |
| tags: | |
| - coreml | |
| - sparse | |
| - splade | |
| - sparse-encoder | |
| - information-retrieval | |
| - evoke | |
| - apple-silicon | |
| - neural-engine | |
| # Granite Embedding 30M Sparse for Core ML | |
| Core ML conversion of IBM's [`ibm-granite/granite-embedding-30m-sparse`](https://huggingface.co/ibm-granite/granite-embedding-30m-sparse) | |
| (revision `ad82b1fd`), the learned sparse encoder behind Intelligent Internet's | |
| [Evoke](https://github.com/Intelligent-Internet/Evoke). It turns text into a short list of weighted vocabulary terms, | |
| including related terms the text never uses ("vaccines for seniors" β vaccine, vaccination, older, elder, β¦), which go | |
| into an ordinary inverted index next to BM25. 30M parameters, fp16, Apache-2.0. IBM authored the model; Fluid Inference | |
| converted it. | |
| ## Packages | |
| | Package | Input | Output | Size | | |
| | --- | --- | --- | ---: | | |
| | `granite-embedding-30m-sparse-L64-fp16.mlpackage` | `input_ids` int32 `[1, 64]` | `max_logits` fp16 `[1, 192]`, `vocab_ids` int32 `[1, 192]` | 58 MB | | |
| | `granite-embedding-30m-sparse-L128-fp16.mlpackage` | `[1, 128]` | same | 58 MB | | |
| | `granite-embedding-30m-sparse-L256-fp16.mlpackage` | `[1, 256]` | same | 58 MB | | |
| | `granite-embedding-30m-sparse-L512-fp16.mlpackage` | `[1, 512]` | same | 58 MB | | |
| One package per sequence length (fixed shapes keep the graph on the Neural Engine; enumerated shapes ran entirely on | |
| CPU). macOS 14 / iOS 17 or newer. | |
| - **Input:** RoBERTa byte-level BPE (`tokenizer.json`), `<s>` β¦ `</s>`, right-padded with `<pad>` (id 1) to the | |
| package length. The attention mask is derived from `<pad>` inside the graph. | |
| - **Output:** the 192 vocabulary terms with the highest MLM logit at any token position (max-pooled over tokens), | |
| sorted descending. | |
| - **Weights (host side, fp32):** `w = log1p(max(v, 0)) ^ gamma Γ scale` over the first `active_dims`, keep `w > 0`. | |
| `config.json` holds Evoke P2.2's constants: queries keep 50 terms (gamma 1.8519, scale 0.6964), documents keep 192 | |
| (gamma 0.5628, scale 1.0). Score = Ξ£ query weight Γ document weight over shared terms. | |
| ## Accuracy | |
| Against Evoke's shipped ONNX compilers (`Intelligent-Internet/Evoke-Model-Beta-1`), NFCorpus test (323 queries, | |
| 3,633 documents), sparse dot over the semantic terms only: | |
| | Encoder | nDCG@10 | Recall@100 | Recall@1000 | Same top-10 as ONNX | | |
| | --- | ---: | ---: | ---: | ---: | | |
| | ONNX (shipped, CPU) | 0.3402 | 0.2956 | 0.5854 | β | | |
| | Core ML fp16, GPU | 0.3405 | 0.2957 | 0.5856 | 99.6% | | |
| | Core ML fp16, Neural Engine | 0.3405 | 0.2953 | 0.5847 | 98.7% | | |
| An fp32 build (not uploaded) reproduces the ONNX term sets exactly (200/200 queries and documents); the fp16 differences | |
| are low-weight terms at the top-k boundary. | |
| ## Speed | |
| Apple M5 Pro, macOS 27, one warm call (median of 100), Swift: | |
| | Length | Neural Engine | GPU | CPU | | |
| | --- | ---: | ---: | ---: | | |
| | 64 | **0.74β0.87 ms** | 2.5β3.3 ms | 2.8β3.1 ms | | |
| | 128 | **1.2β1.4 ms** | 1.5β4.3 ms | 4.1 ms | | |
| | 256 | 2.9 ms | **1.9β3.4 ms** | 9.5β15 ms | | |
| | 512 | 7.6 ms | **2.6 ms** | 16β27 ms | | |
| Use the Neural Engine up to 128 tokens and the GPU beyond. 160 of 172 ops run on the Neural Engine; token lookup, mask | |
| setup and the final top-k stay on CPU. After several seconds idle the first Neural Engine call takes ~15 ms. Encoding | |
| 3,633 NFCorpus documents: 14.1 s Core ML GPU vs 88.2 s ONNX Runtime CPU. | |
| ## Use from Swift | |
| [FluidUse](https://github.com/FluidInference/FluidUse) wraps the packages (`EvokeManager`, Swift RoBERTa tokenizer with token ids identical to the Python tokenizer, | |
| host-side weighting; [PR #28](https://github.com/FluidInference/FluidUse/pull/28)) and includes `EvokeSearchDemo`, an | |
| auto-playing search-as-you-type demo. | |