Instructions to use smdesai/GLiNER25-Multi-Decide-FP16-CoreAI with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- GLiNER2
How to use smdesai/GLiNER25-Multi-Decide-FP16-CoreAI with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("smdesai/GLiNER25-Multi-Decide-FP16-CoreAI") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5-multi-Decide β Core AI (iOS 27+ / macOS 27+)
| File | Purpose |
|---|---|
GLiNER25-multi-Decide-FP16.aimodel/ |
Core AI source asset, FP16 (174 MB), three functions sharing one set of weights |
word_embeddings.f16 |
Word-embedding table, FP16 (384 MB), looked up on the host |
tokenizer.json, tokenizer_config.json |
Checkpoint tokenizer (mDeBERTa-v3, 250k-piece Unigram) |
model-card.md |
Upstream model card (fastino/GLiNER2.5-multi-Decide, Apache-2.0) |
Source: checkpoint revision 6bc1d43d201b0691e733626389af8c57eea3ea68, classification only
(mDeBERTa-v3-base encoder + classifier MLP). Exported with coreai-torch 0.4.3 /
coreai-core 1.0.0b3 / Torch 2.11.0. The graph uses constant relative-position buckets, a
gather-free relative-shift bias and fused SDPA. Learned weights are unchanged.
Interface
Functions context128, context256, context512:
inputs_embedsfloat16[1, L, 768]: rowinput_ids[i]ofword_embeddings.f16at positioni, including padding positions (pad id 0)attention_maskint32[1, L](1 = real token, 0 = padding)- output:
logits[1, L], one score per token position. Classification reads the positions of the schema's label markers ([L]).
Choose the smallest function that fits the request and pad to its length. The maximum is 512 tokens. Reject longer requests; do not truncate.
Word embeddings: word_embeddings.f16
The 250,112 Γ 768 word-embedding table (69% of the parameters) is kept outside the model, so the model file is small and only the rows a request uses are read. The file is little-endian FP16, row-major, no header (384,172,032 bytes). Memory-map it and copy one 1,536-byte row per token:
let table = try Data(contentsOf: tableURL, options: .alwaysMapped)
let rowBytes = 768 * MemoryLayout<Float16>.size
table.withUnsafeBytes { raw in
for (i, id) in inputIDs.enumerated() { // destination: the inputs_embeds buffer
memcpy(destination + i * 768, raw.baseAddress! + id * rowBytes, rowBytes)
}
}
The embedding LayerNorm and everything after it are in the model, so results are bit-identical to an in-model lookup. With the file in the page cache a lookup takes 40β75 Β΅s on an iPhone 17 Pro; the first requests after install, while rows are read from storage, take 1β5 ms longer. The Core ML and Core AI conversions use the same file.
Input preparation
Token IDs come from the upstream GLiNER2 processor (classify_text layout: schema
( [P] task ( [L] label ... ) ), [SEP_TEXT], then the lower-cased text words, each piece
tokenized on its own, no [CLS]/[SEP]). A Swift port that matches the Python processor
exactly (3,099 pieces and 153 requests in 26 languages) is in the conversion repository.
Required: GPU placement
let model = try await AIModel(
contentsOf: url, options: SpecializationOptions(preferredComputeUnitKind: .gpu))
The English GLiNER2.5-Decide conversion with the same graph crashed with an MPSGraph ANE-region assertion under default placement on iOS 27.2. Default placement has not been tested with this model; use GPU.
Validation (iPhone 17 Pro, iOS 27.2, GPU)
Strict gate: every decision matches the FP32 PyTorch oracle and the maximum probability error is β€ 0.005. Corpora: 43 (128), 80 (256) and 113 (512) requests in English and 25 other languages and scripts, including requests of 257β512 tokens. Three launches, identical results every launch.
| Function | Max prob. error | Median | Footprint peak |
|---|---|---|---|
| context128 | 0.0037 | 10.7β10.8 ms | 213 MiB |
| context256 | 0.0037 | 15.0β15.1 ms | 218 MiB |
| context512 | 0.0037 | 31.9 ms | 226β243 MiB |
Medians and footprint are from single-function launches; with all three functions loaded in one process the footprint peak is 223β264 MiB. The first load specializes the asset for the device (~2β3 s); later loads hit the system cache (0.01 s). Also passes on a Mac (M3 Max) GPU: max error 0.0034 / 0.0034 / 0.0045.
SHA-256
046f21c5b3a5bf0e9a8e311d6b929ca24c0fbde27085df88049c90004b22e4d4 GLiNER25-multi-Decide-FP16.aimodel/main.mlirb
2248b04176406fb2f8776b1df5849b821c0e416fdbb0c16f9772bc68eeb434f0 word_embeddings.f16
c62446df87ae18ec98b133f8f84fc449a07cc89bbf8ef192a4cb5f9c53777a7a tokenizer.json
Model tree for smdesai/GLiNER25-Multi-Decide-FP16-CoreAI
Base model
fastino/gliner2.5-multi-v1