|
Download README.md from smdesai/MXBAIEdgeColbert: direct link, hf CLI and curl.
- Browser
- Download file 1.87 kB
-
https://huggingface.co/smdesai/MXBAIEdgeColbert/resolve/main/README.md
- Command line
-
hf download hf://smdesai/MXBAIEdgeColbert/README.md
-
curl -L -o README.md https://huggingface.co/smdesai/MXBAIEdgeColbert/resolve/main/README.md
1.87 kB
| base_model: mixedbread-ai/mxbai-edge-colbert-v0-32m | |
| license: apache-2.0 | |
| library_name: coreml | |
| pipeline_tag: sentence-similarity | |
| tags: | |
| - coreml | |
| - colbert | |
| - multi-vector | |
| - late-interaction | |
| - feature-extraction | |
| - sentence-similarity | |
| - apple | |
| - on-device | |
| # MXBAIEdgeColbert (Core ML) | |
| A [Core ML](https://developer.apple.com/documentation/coreml) export of | |
| [`mixedbread-ai/mxbai-edge-colbert-v0-32m`](https://huggingface.co/mixedbread-ai/mxbai-edge-colbert-v0-32m) | |
| for on-device ColBERT late-interaction (multi-vector) encoding on Apple platforms. | |
| ## Base model | |
| - **Base model:** [`mixedbread-ai/mxbai-edge-colbert-v0-32m`](https://huggingface.co/mixedbread-ai/mxbai-edge-colbert-v0-32m) | |
| This repository contains only the compiled Core ML encoder (`MXBAIEdgeColbert.mlmodelc`). | |
| All weights are derived from the base model above; no additional training was performed. | |
| ## What was exported | |
| The exported program wraps the base transformer plus its three ColBERT `Dense` | |
| projection layers (384 → 768 → 768 → 64, no bias / no activation), producing | |
| per-token embeddings that are L2-normalized and masked by the attention mask. | |
| | Property | Value | | |
| | --- | --- | | |
| | Inputs | `input_ids` and `attention_mask`, `int32`, static shape `(1, 256)` | | |
| | Output | `token_embeddings` (per-token, dim 64, L2-normalized, padding zeroed) | | |
| | Compute precision | `float16` | | |
| | Minimum deployment target | iOS 18 / macOS 15 | | |
| Tokenization and the ColBERT `[Q]` / `[D]` prefixing and skiplist masking are | |
| performed by the host application; this model performs encoding only. | |
| ## Usage | |
| Intended for use as the encoder in an on-device ColBERT/PLAID retrieval | |
| pipeline. Load the compiled model with Core ML, feed padded `int32` | |
| `input_ids` / `attention_mask`, and read `token_embeddings`. | |
| ## License | |
| Released under the Apache-2.0 license, following the base model. | |