Instructions to use litert-community/GLiNER2.5-Multi-LiteRT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/GLiNER2.5-Multi-LiteRT with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- GLiNER2
How to use litert-community/GLiNER2.5-Multi-LiteRT with GLiNER2:
from gliner2 import AutoExtractor extractor = AutoExtractor.from_pretrained("litert-community/GLiNER2.5-Multi-LiteRT") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2.5 Multi for LiteRT — Android GPU FP32
GLiNER2.5 Multi extracts person, organization, location, product and date entities from English and Japanese text. These files are the official fastino/gliner2.5-multi-v1 weights converted for LiteRT with LiteRT Torch. Tested on a Samsung Galaxy S26 (SM-S942Q, SM8850, Android 16) with LiteRT 2.2.0 and explicit GPU FP32 computation: at every shipped window, the spans match the official gliner2 fp32 implementation exactly on every bilingual test input that fits (70, 75 and 80 inputs; span F1 1.000). The phone's Hexagon NPU matched 69/70, 72/75 and 77/80 span sets, so it is reported below but not offered as a supported path.
The image renders the actual output of gliner25_multi_s128_wfp16.tflite with the
fp16 embedding table (LiteRT CompiledModel, desktop CPU) on the two sentences in
host_assets/example.json. The person, organization and product names are
invented.
Files and supported configuration
Three window sizes are shipped. Use the smallest window whose N holds the encoded tokens (the five-label schema prompt plus the text) and whose T holds the processed text slots. Longer inputs are rejected, not truncated.
| File | Window N / text slots T | Packed output floats | Bytes | Role |
|---|---|---|---|---|
gliner25_multi_s128_wfp16.tflite |
128 / 48 | 78,110 | 191,876,560 | recommended |
gliner25_multi_s256_wfp16.tflite |
256 / 192 | 292,958 | 211,168,608 | recommended |
gliner25_multi_s512_wfp16.tflite |
512 / 384 | 579,422 | 250,248,384 | recommended |
fp32/gliner25_multi_s128_fp32.tflite |
128 / 48 | 78,110 | 363,985,160 | fp32 reference |
fp32/gliner25_multi_s256_fp32.tflite |
256 / 192 | 292,958 | 383,277,208 | fp32 reference |
fp32/gliner25_multi_s512_fp32.tflite |
512 / 384 | 579,422 | 422,356,980 | fp32 reference |
wfp16 files store the 96 FULLY_CONNECTED weight tensors as float16 and keep
everything else, including inputs and outputs, in float32. The fp32/ files are the
same graphs with float32 weights. Weight storage does not set the GPU compute
precision: request FP32 explicitly. At the default precision the outputs stay
finite, but one of 70 span sets changes.
Each graph is the dense part of the upstream BoundaryExtractor: the mDeBERTa-v3-base encoder, the boundary encoder, the boundary query head and the per-token projections. It takes word-embedding rows and returns one packed float32 tensor of 17 logical outputs. The host splits and tokenizes the text and looks up the rows before the graph, then runs the upstream sparse decoder after it.
The caller picks the word splitter per request: whitespace for space-delimited
text such as English, char for Japanese and other CJK text. There is no language
detector; mixed-language text needs an explicit choice too. Capacity counts
processed text slots and encoded tokens, not source characters. char gives each
non-space character its own slot and keeps runs of ASCII letters and digits
together. Text ending in 。 takes one extra slot, because the processor appends
. to text that does not end in ., ! or ?.
host_assets/ holds everything the host needs:
| File | Purpose |
|---|---|
word_embeddings_fp16.bin |
recommended: [250112,768] float16, row-major, no header, 384,172,032 B; upcast selected rows to float32 |
word_embeddings_fp32.bin |
reference: float32, 768,344,064 B, bit-identical to the checkpoint tensor |
sparse_decoder_fp32.safetensors, decoder_parameters.json |
the 16 upstream sparse-decoder tensors (860,676 B) and their list |
tokenizer.json, tokenizer_config.json, config.json, encoder_config/config.json |
exact files from the pinned checkpoint |
graph_contract_s{128,256,512}.json |
input shapes and output-slice offsets |
runtime/ |
Python host runtime: input construction, unpacking, upstream decoding |
example.json |
both example sentences with encoded inputs and official spans |
Minimal usage
Python — complete pipeline, desktop CPU
Install requirements-lock.txt and run from the repository root. --seq auto (the
default) picks the smallest fitting window; --model fp32 and --table fp32 select
the references.
python examples/run_example.py --splitter whitespace \
--text "Mira Velsan presented the Lumenquill tablet for Asterfold Labs in Bristol on October 12."
The same steps by hand:
import json, os, sys
from pathlib import Path
import numpy as np
os.environ["HF_HUB_OFFLINE"] = "1"
sys.path.insert(0, str(Path("host_assets/runtime").resolve()))
from host_runtime import HostRuntime
from ai_edge_litert.compiled_model import CompiledModel, CpuOptions, HardwareAccelerator, Options
host = HostRuntime(Path("host_assets"), table="fp16")
text = "星瀬澄香は京都で霧灯研究社の新製品『星糸端末』を紹介した。"
inputs, captured = host.prepare(text, 128, "char")
model = CompiledModel.from_file("gliner25_multi_s128_wfp16.tflite", options=Options(
hardware_accelerators=HardwareAccelerator.CPU, cpu_options=CpuOptions(num_threads=4)))
sig = "serving_default"
ins = {f"args_{i}": model.create_input_buffer_by_name(sig, f"args_{i}") for i in range(5)}
outs = {"output_0": model.create_output_buffer_by_name(sig, "output_0")}
for i, x in enumerate(inputs):
ins[f"args_{i}"].write(np.ascontiguousarray(x, dtype=np.float32))
model.run_by_name(sig, ins, outs)
packed = outs["output_0"].read(78110, np.float32).reshape(1, 1, 1, 78110)
print(json.dumps(host.decode(captured, packed, inputs)["entities"], ensure_ascii=False))
# {"person": [{"text": "星瀬澄香", "confidence": 0.9864593744277954, "start": 0, "end": 4}], ...
Kotlin — Android GPU with explicit FP32
Add com.google.ai.edge.litert:litert:2.2.0, stage the model file in app-private
storage and keep one Environment per process. For s128 the inputs are args_0
[1,128,768], args_1 [1,128], args_2 [1,48,128], args_3 [1,5,128] and
args_4 [1,48], filled as HOST_CONTRACT.md describes.
import com.google.ai.edge.litert.Accelerator
import com.google.ai.edge.litert.CompiledModel
import com.google.ai.edge.litert.Environment
import java.io.File
fun runDensePrefix(env: Environment, modelDir: File, inputs: List<FloatArray>): FloatArray {
val options = CompiledModel.Options(Accelerator.GPU).apply {
gpuOptions = CompiledModel.GpuOptions(precision = CompiledModel.GpuOptions.Precision.FP32)
}
val path = File(modelDir, "gliner25_multi_s128_wfp16.tflite").path
CompiledModel.create(path, options, env).use { model ->
val ins = (0..4).associate { "args_$it" to model.createInputBuffer("args_$it", "serving_default") }
val outs = mapOf("output_0" to model.createOutputBuffer("output_0", "serving_default"))
try {
inputs.forEachIndexed { i, x -> ins.getValue("args_$i").writeFloat(x) }
model.run(ins, outs, "serving_default")
return outs.getValue("output_0").readFloat()
} finally {
(ins.values + outs.values).forEach { it.close() }
}
}
}
// start_logits [1,5,49] begins at float 46976 (graph_contract_s128.json)
fun startLogits(packed: FloatArray): Array<FloatArray> =
Array(5) { q -> packed.copyOfRange(46976 + q * 49, 46976 + (q + 1) * 49) }
The host steps (splitter, tokenizer, embedding lookup, sparse decoder) ship in
Python only. An Android port can be checked against host_assets/example.json,
which holds the encoded ids, routing positions and official spans.
Host contract
HOST_CONTRACT.md is the complete specification. In short:
- Build the sequence the way the upstream
gliner2processor does: schema tokens for the five labels,[SEP_TEXT], then the split text. Tokenizing the joined string or adding BOS/CLS does not reproduce it. Right-pad the ids with 0 to N. args_0holds the raw table rows for every position, padding included; the graph applies the embedding LayerNorm. The other four inputs are 0/1 masks and one-hot routing rows.- The upstream sparse decoder turns the 17 output slices (offsets in
graph_contract_s*.json) into label, text, start, end and confidence per span at threshold 0.5. Offsets count Unicode code points:text[start:end]is the entity.
Measured quality and performance
The reference is the official gliner2 2.0.0 fp32 CPU implementation at revision
235cf92d, threshold 0.5, same splitter. The 80 test inputs (35 English, 45
Japanese) use the five-label schema; 70 / 75 / 80 fit the three windows. A span
counts only when label, start and end all match, so F1 1.000 means the files
reproduce the official output, not that every entity is right. Upstream misses carry
over: on ten Japanese check sentences the upstream model found 4 of 10 expected
organizations, and so do the converted files.
Desktop LiteRT CPU (ai-edge-litert 2.1.6 CompiledModel, macOS arm64, fp32 table):
| Window | Inputs | wfp16 F1 (max confidence drift) | fp32 F1 (max confidence drift) |
|---|---|---|---|
| 128 | 70 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 256 | 75 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
| 512 | 80 | 1.000 (1.3e-3) | 1.000 (1.2e-5) |
Every input returns the identical span set, also with the fp16 table (drift at most 1.4e-3).
On the Galaxy S26 (LiteRT 2.2.0 Kotlin CompiledModel API, wfp16 graphs unless noted), only the graph ran on the phone. Inputs were built with the fp16 table, and outputs were decoded on the desktop with the Python runtime. A pass needs every official span set and a confidence drift of at most 0.01. GPU and NPU each took the whole graph as one partition.
| Window | Accelerator, precision | Identical span sets | Max confidence drift | Delegated (logcat) | Load + compile | Opening ms | Sustained ms |
|---|---|---|---|---|---|---|---|
| 128 | GPU, explicit FP32 | 70/70 | 1.4e-3 | 1148/1148 nodes (LITERT_CL) | 1.5 s | 24.5 | 24.3 |
| 256 | GPU, same | 75/75 | 1.4e-3 | 1149/1149 | 1.4 s | 63.5 | 93.0 |
| 512 | GPU, same | 80/80 | 1.4e-3 | 1149/1149 | 1.7 s | 200.9 | 422.5 |
| 128 | NPU (Hexagon HTP, JIT, BURST), fp16 compute | 69/70 | 4.8e-2 | 1148/1148 ops, 1 node (DispatchDelegate) | 5.4 s | 8.3 | 8.3 |
| 256 | NPU, same | 72/75 | 4.8e-2 | 1149/1149 ops, 1 node | 45.7 s | 34.5 | 34.4 |
| 512 | NPU, same | 77/80 | 4.8e-2 | 1149/1149 ops, 1 node | 167.1 s | 186.9 | 186.5 |
| 128 | CPU (reference), XNNPACK, 4 threads | 70/70 | 1.4e-3 | 1147/1148 (XNNPACK) | 0.3 s | 32.7 | 49.5 |
| 256 | CPU, same | 75/75 | 1.4e-3 | 1148/1149 | 0.3 s | 74.6 | 128.5 |
| 512 | CPU, same | 80/80 | 1.4e-3 | 1148/1149 | 0.3 s | 210.1 | 407.3 |
| 128 | GPU, explicit FP32, fp32/ graph |
70/70 | 3.2e-4 | 1052/1052 | 1.5 s | 24.7 | 24.3 |
| 128 | GPU, default precision (fp16 compute), diagnostic only | 69/70 | 3.3e-2 | 1148/1148 | 1.2 s | 13.9 | 15.0 |
Every NPU difference sits at the 0.5 threshold: it dropped two spans the official model returns at 0.502 and 0.503 and added one at 0.519 that the official model keeps below 0.5. All NPU outputs were finite, but the drift exceeds 0.01, so the NPU is reported, not validated. The default-precision GPU row flipped one span at the threshold, which is why FP32 is required.
Each input ran 2 warm-ups and 5 timed repetitions of input write → run → output
readback. Opening is the first input's median after the phone cooled to thermal
status 0 (battery 34.1–37.6 °C). For NPU s256 and s512, the first compile had warmed
it to status 2 (39.3 °C and 41.3 °C). Sustained is the median over all inputs of the
back-to-back job, during which the phone reached at most status 2 and 43.6 °C. Each
load is a first load in a fresh app process; for the NPU it includes the on-device
compile. Latency is a single-device sample, not a benchmark.
Limits
- Schema: only the five labels above, in that order, at threshold 0.5. Other labels, relations, classifications and record schemas were not validated.
- Languages: English (
whitespace) and Japanese (char) only. - Length: one window per call. Splitting longer text was not validated.
- Devices: tested on the Galaxy S26 only.
- NPU: measured above, not validated.
Provenance, conversion and license
- Source:
fastino/gliner2.5-multi-v1, revision235cf92d6d4318da9bfca0d08975c8fa7250d13b(BoundaryExtractor, encodermicrosoft/mdeberta-v3-base), loaded withgliner22.0.0 andtransformers4.57.6. - Conversion:
litert-torch0.9.3 (torch2.12.1), fixed shapes. The dense prefix was re-expressed for the GPU delegate without changing its math (matmul routing, float masks, rank-4 attention, baked relative-position buckets). - Two fp32-exact rewrites keep the same files finite under fp16 delegates: the attention-mask fill is -10000 instead of the float32 minimum, and each of the 29 active LayerNorms is computed on x·2⁻³ with eps·2⁻⁶. On 20 inputs per window the rewritten torch graph matched the original bit for bit (60/60 comparisons, maximum difference 0.0).
- Weight storage:
ai-edge-quantizer0.8.0 FLOAT_CASTING on FULLY_CONNECTED weights only (96 tensors). - Desktop checks used
ai-edge-litert2.1.6; complete pins are inrequirements-lock.txt.
License: the checkpoint is Apache-2.0, and these converted files are released under
the same license (LICENSE). The mDeBERTa-v3 encoder is MIT
(licenses/DeBERTa-MIT.txt). The gliner2 and transformers code used by the host
runtime is Apache-2.0 (licenses/), with notices in NOTICE. Upstream: the
model, the
GLiNER2 repository and the paper
arXiv:2507.18546. Runtime:
LiteRT.
- Downloads last month
- 42
Model tree for litert-community/GLiNER2.5-Multi-LiteRT
Base model
fastino/gliner2.5-multi-v1