Sentence Similarity
ONNX
Safetensors
sentence-transformers
PyLate
modernbert
ColBERT
multi-vector
feature-extraction
multilingual
code search
text-embeddings-inference
🇪🇺 Region: EU
Instructions to use lightonai/mLateOn with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use lightonai/mLateOn with sentence-transformers:
from pylate import models queries = [ "Which planet is known as the Red Planet?", "What is the largest planet in our solar system?", ] documents = [ ["Mars is the Red Planet.", "Venus is Earth's twin."], ["Jupiter is the largest planet.", "Saturn has rings."], ] model = models.ColBERT(model_name_or_path="lightonai/mLateOn") queries_emb = model.encode(queries, is_query=True) docs_emb = model.encode(documents, is_query=False) - Inference
- Notebooks
- Google Colab
- Kaggle
Re-export model.onnx and model_int8.onnx from model.safetensors; fix onnx_config and tokenizer_config
#3
by ameliechatelain - opened
Fixes https://huggingface.co/lightonai/mLateOn/discussions/2 (thanks @jmzzomg ).
Root cause
The previous model.onnx was not exported from this checkpoint. Comparing its
initializers with model.safetensors (diff_onnx_weights.py): every tensor differs,
token embeddings have max |diff| 7.0e-3 and correlation 0.99999 with 16,317 of
256,002 rows identical, and the final 768->128 projection has max |diff| 5.7e-4.
The new one has initializers bit-identical to model.safetensors.
What changed
- model.onnx: re-exported with export_onnx.py. Same inputs and output as before:
input_ids / attention_mask int64 [batch, seq] with the [Q]/[D] prefix already
inserted; output float32 [batch, seq, 128], L2-normalized per token, padding
positions not zeroed. - model_int8.onnx: dynamic INT8 quantization of the new graph (QInt8 weights,
per-channel scales). - onnx_config.json: model_name
lightonai/LateOn-multilingual->lightonai/mLateOn. - tokenizer_config.json: max_length and model_max_length 299 -> 8191, matching
sentence_bert_config.json and the truncation rule in tokenizer.json. - export_onnx.py: the export, quantization and validation script, so the next
release does not repeat this.uv run export_onnx.py export|quantize|check.
Validation (ONNX Runtime 1.23.2 CPU vs PyLate 1.6.0 encode(), same tokens)
Ten multilingual documents (547 tokens) and ten matching queries (100 tokens),
MaxSim on the 10x10 query-document matrix, and single unpadded documents from
11 to 6753 tokens.
| old model.onnx | new model.onnx | old model_int8.onnx | new model_int8.onnx | |
|---|---|---|---|---|
| documents, max per-token diff | 9.5e-2 | 2.5e-5 | 1.1e-1 | 7.1e-2 |
| documents, mean token cosine | 0.99751 | 1.0000000 | 0.99377 | 0.99411 |
| queries, max per-token diff | 3.6e-2 | 4.2e-7 | 9.0e-2 | 5.2e-2 |
| queries, mean token cosine | 0.99879 | 1.0000000 | 0.99322 | 0.99600 |
| MaxSim score max diff | 7.1e-2 | 1.4e-6 | 2.1e-1 | 1.5e-1 |
| top-1 agreement | 100% | 100% | 100% | 100% |
| full ranking agreement (10 docs) | 10% | 100% | 0% | 0% |
| cosine min at 6753 tokens | n/a | 0.9999996 | n/a | 0.945 |
ameliechatelain changed pull request status to merged