|
Download README.md from FluidInference/gliclass-edge-apps-coreml: direct link, hf CLI and curl.
- Browser
- Download file 4.22 kB
-
https://huggingface.co/FluidInference/gliclass-edge-apps-coreml/resolve/main/README.md
- Command line
-
hf download hf://FluidInference/gliclass-edge-apps-coreml/README.md
-
curl -L -o README.md https://huggingface.co/FluidInference/gliclass-edge-apps-coreml/resolve/main/README.md
4.22 kB
| license: apache-2.0 | |
| library_name: coreml | |
| pipeline_tag: zero-shot-classification | |
| base_model: knowledgator/gliclass-edge-v3.0 | |
| base_model_relation: finetune | |
| tags: | |
| - coreml | |
| - apple-silicon | |
| - neural-engine | |
| - gliclass | |
| - decision-model | |
| - fluiduse | |
| # GLiClass Edge Apps v2 · Core ML | |
| Application-tuned, 32,705,154-parameter GLiClass model converted to Core ML for | |
| [FluidUse](https://github.com/FluidInference/FluidUse). It receives a state, an optional | |
| instruction, and 2–25 candidate labels, then returns a probability for each label in one | |
| encoder pass. It does not generate text. The source checkpoint is | |
| [`knowledgator/gliclass-edge-v3.0`](https://huggingface.co/knowledgator/gliclass-edge-v3.0) | |
| at revision `df03993a2ed98e5e4a0d2dd7efbbd105abe874cf`, further tuned for | |
| FluidUse application decisions before conversion. The tokenizer is included. | |
| ## Files | |
| | File | Use | Approximate package size | | |
| | --- | --- | ---: | | |
| | `gliclass_edge_apps_fp16_L128_options25.mlpackage/` | Short prompts, full precision | 65.7 MB | | |
| | `gliclass_edge_apps_fp16_L256_options25.mlpackage/` | Medium prompts, full precision | 65.9 MB | | |
| | `gliclass_edge_apps_fp16_L512_options25.mlpackage/` | Long prompts, full precision | 66.7 MB | | |
| | `gliclass_edge_apps_lut8_kmeans_per_tensor_L128_options25.mlpackage/` | Compact L128 option | 33.0 MB | | |
| | `tokenizer.json` | Matching tokenizer | 3.6 MB | | |
| | `config.json` | Buckets, tensor shapes, source, and runtime metadata | — | | |
| The `.mlpackage` folders are portable Core ML models targeting iOS 17 / macOS 14 or | |
| newer. The FluidUse loader compiles them on the local machine. L128, L256, and L512 | |
| support up to 25 supplied labels; labels must fit inside the selected token limit. | |
| The LUT8 package is available only at L128. The Swift runtime chooses the smallest | |
| loaded bucket that fits and truncates long state text at its largest loaded bucket; | |
| it rejects requests if their option labels no longer fit. Inspect | |
| `GLiClassAnswer.textWasTruncated` when this matters. | |
| ## Use with FluidUse | |
| Download this repository to a local directory, then load one or more buckets: | |
| ```bash | |
| hf download FluidInference/gliclass-edge-apps-coreml --local-dir ./gliclass-assets | |
| ``` | |
| ```swift | |
| import FluidUse | |
| import Foundation | |
| let manager = try await GLiClassManager.load( | |
| from: URL(fileURLWithPath: "./gliclass-assets", isDirectory: true), | |
| configuration: .init(lengths: [128, 256, 512], precision: "fp16")) | |
| let answer = try await manager.classify( | |
| text: "The next support request asks for a refund after a duplicate charge.", | |
| labels: ["refund", "shipping", "technical support"], | |
| prompt: "Which team should handle this request?") | |
| print(answer.selectedLabel, answer.probabilities) | |
| ``` | |
| For the compact bucket, set `lengths: [128]` and `precision: "lut8"`. The | |
| [FluidUse 2048 demo](https://github.com/FluidInference/FluidUse/tree/main/Sources/GLiClass2048Demo) | |
| and [Tetris demo](https://github.com/FluidInference/FluidUse/tree/main/Sources/LayaTetrisDemo) | |
| use GLiClass to compare shortlisted moves. They use game heuristics to prepare | |
| candidates; the model makes the final choice among those candidates. | |
| ## Verification and limits | |
| The converted FP16 model changed 18 argmax choices out of 3,899 application-suite | |
| rows relative to its PyTorch checkpoint. On an Apple M5 Pro, the L128 FP16 bucket | |
| measured 0.843 ms median model latency; LUT8 measured 0.903 ms and agreed with | |
| FP16 on 97.1% of L128 application-suite choices. These are local measurements | |
| under the protocol in [RESULTS.md](RESULTS.md), with supporting JSON in | |
| [`reports/`](reports/). | |
| A local reproduction of the 19 static Decision Index headline benchmarks scored | |
| 19.27. It was not submitted as an official tracker result because the full public | |
| bundle was unavailable during the run. The model was tuned on application tasks, | |
| not games, and a game result depends on its action shortlist and harness. | |
| Training, conversion, verification, and pinned sources are documented in the | |
| [FluidInference conversion toolkit](https://github.com/FluidInference/mobius/pull/101). | |
| The original GLiClass model is by Knowledgator. This application-tuned Core ML | |
| variant is released under Apache-2.0; see [LICENSE](LICENSE). | |