GLiNER2.5-multi-Decide β€” Core ML (iOS 18+ / macOS 15+)

File Purpose
GLiNER25-multi-Decide-FP16.mlpackage/ Core ML package, FP16 weights and compute (208 MB), three functions sharing one set of weights
word_embeddings.f16 Word-embedding table, FP16 (384 MB), looked up on the host
tokenizer.json, tokenizer_config.json Checkpoint tokenizer (mDeBERTa-v3, 250k-piece Unigram)
model-card.md Upstream model card (fastino/GLiNER2.5-multi-Decide, Apache-2.0)

Source: checkpoint revision 6bc1d43d201b0691e733626389af8c57eea3ea68, classification only (mDeBERTa-v3-base encoder + classifier MLP). Built with coremltools 9.0 / Torch 2.7.1 from FP16 128/256/512 exports (constant relative-position buckets, finite FP16 attention mask), merged into one multifunction package with shared weights.

Interface

Functions context128, context256, context512 (select with MLModelConfiguration.functionName):

  • inputs_embeds float16 [1, L, 768]: row input_ids[i] of word_embeddings.f16 at position i, including padding positions (pad id 0)
  • attention_mask int32 [1, L] (1 = real token, 0 = padding)
  • output: logits [1, L], one score per token position. Classification reads the positions of the schema's label markers ([L]).

Choose the smallest function that fits the request and pad to its length. The maximum is 512 tokens. Reject longer requests; do not truncate.

Word embeddings: word_embeddings.f16

The 250,112 Γ— 768 word-embedding table (69% of the parameters) is kept outside the model, so the model file is small and only the rows a request uses are read. The file is little-endian FP16, row-major, no header (384,172,032 bytes). Memory-map it and copy one 1,536-byte row per token:

let table = try Data(contentsOf: tableURL, options: .alwaysMapped)
let rowBytes = 768 * MemoryLayout<Float16>.size
table.withUnsafeBytes { raw in
  for (i, id) in inputIDs.enumerated() {   // destination: the inputs_embeds buffer
    memcpy(destination + i * 768, raw.baseAddress! + id * rowBytes, rowBytes)
  }
}

The embedding LayerNorm and everything after it are in the model, so results are bit-identical to an in-model lookup. With the file in the page cache a lookup takes 40–75 Β΅s on an iPhone 17 Pro; the first requests after install, while rows are read from storage, take 1–5 ms longer. The Core ML and Core AI conversions use the same file.

Input preparation

Token IDs come from the upstream GLiNER2 processor (classify_text layout: schema ( [P] task ( [L] label ... ) ), [SEP_TEXT], then the lower-cased text words, each piece tokenized on its own, no [CLS]/[SEP]). A Swift port that matches the Python processor exactly (3,099 pieces and 153 requests in 26 languages) is in the conversion repository.

Required: GPU compute units

let configuration = MLModelConfiguration()
configuration.computeUnits = .cpuAndGPU
configuration.functionName = "context256"
let model = try MLModel(contentsOf: compiledURL, configuration: configuration)

Do not rely on .cpuOnly: FP16 on the CPU changes no decisions but exceeds the strict gate (max 0.0116 on a Mac). Xcode compiles the .mlpackage to .mlmodelc when it is bundled in an app; use MLModel.compileModel(at:) otherwise.

Validation (iPhone 17 Pro, iOS 27.2, GPU)

Strict gate: every decision matches the FP32 PyTorch oracle and the maximum probability error is ≀ 0.005. Corpora: 43 (128), 80 (256) and 113 (512) requests in English and 25 other languages and scripts, including requests of 257–512 tokens. Three launches, identical results every launch.

Function Max prob. error Median Footprint peak
context128 0.0032 8.4–9.2 ms 62 MiB
context256 0.0038 14.1–14.3 ms 67 MiB
context512 0.0038 39–46 ms 94 MiB

Footprint is from single-function launches; with all three functions loaded in one process the peak is 85–99 MiB, including the first launch after install. With the embedding table inside the model instead, that first launch peaked at 476 MiB. Also passes on a Mac (M3 Max) GPU: max error 0.0034 / 0.0034 / 0.0038. Not yet validated on an iOS 18–26 device or an older GPU.

SHA-256

ce4b3e5e06632cc00bea4419c6dc9a7e3feab2139b0c871c956dd176868341f7  GLiNER25-multi-Decide-FP16.mlpackage/Data/com.apple.CoreML/weights/weight.bin
5f908aed7b4388d6324f17023eaed44a9bcc1484ef8a86f0b180d622a0d72273  GLiNER25-multi-Decide-FP16.mlpackage/Data/com.apple.CoreML/model.mlmodel
2248b04176406fb2f8776b1df5849b821c0e416fdbb0c16f9772bc68eeb434f0  word_embeddings.f16
c62446df87ae18ec98b133f8f84fc449a07cc89bbf8ef192a4cb5f9c53777a7a  tokenizer.json
Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for smdesai/GLiNER25-Multi-Decide-FP16-CoreML

Quantized
(1)
this model