MXBAIEdgeColbert / README.md
smdesai's picture
Add model card: document base model mixedbread-ai/mxbai-edge-colbert-v0-32m
da70e9a verified
|
Raw History Blame Contribute Delete
1.87 kB
---
base_model: mixedbread-ai/mxbai-edge-colbert-v0-32m
license: apache-2.0
library_name: coreml
pipeline_tag: sentence-similarity
tags:
- coreml
- colbert
- multi-vector
- late-interaction
- feature-extraction
- sentence-similarity
- apple
- on-device
---
# MXBAIEdgeColbert (Core ML)
A [Core ML](https://developer.apple.com/documentation/coreml) export of
[`mixedbread-ai/mxbai-edge-colbert-v0-32m`](https://huggingface.co/mixedbread-ai/mxbai-edge-colbert-v0-32m)
for on-device ColBERT late-interaction (multi-vector) encoding on Apple platforms.
## Base model
- **Base model:** [`mixedbread-ai/mxbai-edge-colbert-v0-32m`](https://huggingface.co/mixedbread-ai/mxbai-edge-colbert-v0-32m)
This repository contains only the compiled Core ML encoder (`MXBAIEdgeColbert.mlmodelc`).
All weights are derived from the base model above; no additional training was performed.
## What was exported
The exported program wraps the base transformer plus its three ColBERT `Dense`
projection layers (384 → 768 → 768 → 64, no bias / no activation), producing
per-token embeddings that are L2-normalized and masked by the attention mask.
| Property | Value |
| --- | --- |
| Inputs | `input_ids` and `attention_mask`, `int32`, static shape `(1, 256)` |
| Output | `token_embeddings` (per-token, dim 64, L2-normalized, padding zeroed) |
| Compute precision | `float16` |
| Minimum deployment target | iOS 18 / macOS 15 |
Tokenization and the ColBERT `[Q]` / `[D]` prefixing and skiplist masking are
performed by the host application; this model performs encoding only.
## Usage
Intended for use as the encoder in an on-device ColBERT/PLAID retrieval
pipeline. Load the compiled model with Core ML, feed padded `int32`
`input_ids` / `attention_mask`, and read `token_embeddings`.
## License
Released under the Apache-2.0 license, following the base model.