multilingual-e5-small β€” Core ML

Core ML conversion of intfloat/multilingual-e5-small (revision 614241f622f53c4eeff9890bdc4f31cfecc418b3), used by the Lenz macOS video editor to search talk transcripts by meaning, on device. This repository only hosts the files the app downloads once.

Files

File Contents
SentenceEncoder.mlpackage.zip 12-layer encoder, int8 linear-symmetric weights; inputs input_ids, attention_mask (int32, batch 1–64 Γ— tokens 1–512); output hidden_states (float32, batch Γ— tokens Γ— 384)
tokenizer.zip XLM-RoBERTa tokenizer files (tokenizer.json, config) for swift-transformers
manifest.json File names, sha256s, sizes, model dims

Usage notes

  • Prefix every text with its role: passage: for indexed sentences, query: for searches. Same-prefix use measurably hurts retrieval.
  • Sentence vector = mean of hidden_states over tokens where attention_mask is 1, then L2-normalised; similarity is a dot product.
  • Pad with <pad> (id 1) and a zero mask; truncate to 512 tokens keeping the final </s>.
  • Flexible shapes run fastest on the CPU compute unit (β‰ˆ0.9 ms per sentence on an M2 Max).
  • Conversion is parity-gated: pooled vectors match PyTorch at cosine β‰₯ 0.99 on EN/DE fixtures (int8 worst 0.99983). Source: models/multilingual-e5-small/convert.py in the Lenz repository.

Versioning

Files in this repo are immutable once published. Re-conversions are published as new versions, never overwrites.

License

MIT, same as the original weights by intfloat (Liang Wang et al., "Multilingual E5 Text Embeddings", arXiv:2402.05672). This repository redistributes a converted form of those weights without modification to their values beyond int8 weight quantization.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for artin666/multilingual-e5-small-coreml

Finetuned
(202)
this model

Paper for artin666/multilingual-e5-small-coreml