Multilingual E5 Text Embeddings: A Technical Report
Paper β’ 2402.05672 β’ Published β’ 23
Core ML conversion of intfloat/multilingual-e5-small
(revision 614241f622f53c4eeff9890bdc4f31cfecc418b3), used by the Lenz macOS video editor to search
talk transcripts by meaning, on device. This repository only hosts the files the app downloads once.
| File | Contents |
|---|---|
SentenceEncoder.mlpackage.zip |
12-layer encoder, int8 linear-symmetric weights; inputs input_ids, attention_mask (int32, batch 1β64 Γ tokens 1β512); output hidden_states (float32, batch Γ tokens Γ 384) |
tokenizer.zip |
XLM-RoBERTa tokenizer files (tokenizer.json, config) for swift-transformers |
manifest.json |
File names, sha256s, sizes, model dims |
passage: for indexed sentences, query: for searches.
Same-prefix use measurably hurts retrieval.hidden_states over tokens where attention_mask is 1, then L2-normalised;
similarity is a dot product.<pad> (id 1) and a zero mask; truncate to 512 tokens keeping the final </s>.models/multilingual-e5-small/convert.py in the Lenz repository.Files in this repo are immutable once published. Re-conversions are published as new versions, never overwrites.
MIT, same as the original weights by intfloat (Liang Wang et al., "Multilingual E5 Text Embeddings", arXiv:2402.05672). This repository redistributes a converted form of those weights without modification to their values beyond int8 weight quantization.
Base model
intfloat/multilingual-e5-small