alexwengg's picture
Publish GLiClass Edge Apps v2 Core ML packages and config
f62ad5f verified
|
Raw History Blame Contribute Delete
4.22 kB
---
license: apache-2.0
library_name: coreml
pipeline_tag: zero-shot-classification
base_model: knowledgator/gliclass-edge-v3.0
base_model_relation: finetune
tags:
- coreml
- apple-silicon
- neural-engine
- gliclass
- decision-model
- fluiduse
---
# GLiClass Edge Apps v2 · Core ML
Application-tuned, 32,705,154-parameter GLiClass model converted to Core ML for
[FluidUse](https://github.com/FluidInference/FluidUse). It receives a state, an optional
instruction, and 2–25 candidate labels, then returns a probability for each label in one
encoder pass. It does not generate text. The source checkpoint is
[`knowledgator/gliclass-edge-v3.0`](https://huggingface.co/knowledgator/gliclass-edge-v3.0)
at revision `df03993a2ed98e5e4a0d2dd7efbbd105abe874cf`, further tuned for
FluidUse application decisions before conversion. The tokenizer is included.
## Files
| File | Use | Approximate package size |
| --- | --- | ---: |
| `gliclass_edge_apps_fp16_L128_options25.mlpackage/` | Short prompts, full precision | 65.7 MB |
| `gliclass_edge_apps_fp16_L256_options25.mlpackage/` | Medium prompts, full precision | 65.9 MB |
| `gliclass_edge_apps_fp16_L512_options25.mlpackage/` | Long prompts, full precision | 66.7 MB |
| `gliclass_edge_apps_lut8_kmeans_per_tensor_L128_options25.mlpackage/` | Compact L128 option | 33.0 MB |
| `tokenizer.json` | Matching tokenizer | 3.6 MB |
| `config.json` | Buckets, tensor shapes, source, and runtime metadata | — |
The `.mlpackage` folders are portable Core ML models targeting iOS 17 / macOS 14 or
newer. The FluidUse loader compiles them on the local machine. L128, L256, and L512
support up to 25 supplied labels; labels must fit inside the selected token limit.
The LUT8 package is available only at L128. The Swift runtime chooses the smallest
loaded bucket that fits and truncates long state text at its largest loaded bucket;
it rejects requests if their option labels no longer fit. Inspect
`GLiClassAnswer.textWasTruncated` when this matters.
## Use with FluidUse
Download this repository to a local directory, then load one or more buckets:
```bash
hf download FluidInference/gliclass-edge-apps-coreml --local-dir ./gliclass-assets
```
```swift
import FluidUse
import Foundation
let manager = try await GLiClassManager.load(
from: URL(fileURLWithPath: "./gliclass-assets", isDirectory: true),
configuration: .init(lengths: [128, 256, 512], precision: "fp16"))
let answer = try await manager.classify(
text: "The next support request asks for a refund after a duplicate charge.",
labels: ["refund", "shipping", "technical support"],
prompt: "Which team should handle this request?")
print(answer.selectedLabel, answer.probabilities)
```
For the compact bucket, set `lengths: [128]` and `precision: "lut8"`. The
[FluidUse 2048 demo](https://github.com/FluidInference/FluidUse/tree/main/Sources/GLiClass2048Demo)
and [Tetris demo](https://github.com/FluidInference/FluidUse/tree/main/Sources/LayaTetrisDemo)
use GLiClass to compare shortlisted moves. They use game heuristics to prepare
candidates; the model makes the final choice among those candidates.
## Verification and limits
The converted FP16 model changed 18 argmax choices out of 3,899 application-suite
rows relative to its PyTorch checkpoint. On an Apple M5 Pro, the L128 FP16 bucket
measured 0.843 ms median model latency; LUT8 measured 0.903 ms and agreed with
FP16 on 97.1% of L128 application-suite choices. These are local measurements
under the protocol in [RESULTS.md](RESULTS.md), with supporting JSON in
[`reports/`](reports/).
A local reproduction of the 19 static Decision Index headline benchmarks scored
19.27. It was not submitted as an official tracker result because the full public
bundle was unavailable during the run. The model was tuned on application tasks,
not games, and a game result depends on its action shortlist and harness.
Training, conversion, verification, and pinned sources are documented in the
[FluidInference conversion toolkit](https://github.com/FluidInference/mobius/pull/101).
The original GLiClass model is by Knowledgator. This application-tuned Core ML
variant is released under Apache-2.0; see [LICENSE](LICENSE).