Instructions to use Horizon-Labs/multilingual-reranker-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Horizon-Labs/multilingual-reranker-base with Transformers:
# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-base") model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-base", device_map="auto") - sentence-transformers
How to use Horizon-Labs/multilingual-reranker-base with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("Horizon-Labs/multilingual-reranker-base") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Transformers.js
How to use Horizon-Labs/multilingual-reranker-base with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('text-ranking', 'Horizon-Labs/multilingual-reranker-base'); - Notebooks
- Google Colab
- Kaggle
Multilingual Reranker (base, 308M)
A cross-encoder reranker for search and RAG in 40+ languages, built on mmBERT-base. Give it a query and candidate passages, for example the top 20-100 hits of a vector or keyword search. It scores each pair, so you can re-sort the candidates and keep the best ones for your LLM. Queries and passages can be in different languages. It is distilled from Qwen3-Reranker-4B into a model about 13x smaller. Apache-2.0, trained on openly licensed web text. ONNX files for CPU and the browser (transformers.js) are included. A smaller, faster version is available as Horizon-Labs/multilingual-reranker-small. Try it in the browser.
- One output logit per (query, passage) pair: higher = more relevant.
sigmoid(logit)gives a 0-1 relevance score. - Max length: trained with 384 tokens per pair. Longer passages are truncated; split long documents into chunks.
onnx/model_quantized.onnx(int8 embeddings, 641 MB): scores differ from fp32 by 0.001 on average (sigmoid scale) over 1240 benchmark pairs.
Usage
sentence-transformers:
from sentence_transformers import CrossEncoder
model = CrossEncoder("Horizon-Labs/multilingual-reranker-base")
query = "How tall is the Eiffel Tower?"
passages = ["The Eiffel Tower is 330 metres tall.", "La tour Eiffel a été construite pour l'Exposition universelle de 1889.",
"The Statue of Liberty is 93 metres tall."]
print(model.rank(query, passages)) # [{'corpus_id': 0, 'score': ...}, ...]
transformers:
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
tok = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-base")
model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-base").eval()
enc = tok([query] * len(passages), passages, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
scores = model(**enc).logits[:, 0].sigmoid()
transformers.js:
import { AutoTokenizer, AutoModelForSequenceClassification } from "@huggingface/transformers";
const tok = await AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-base");
const model = await AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-base", { dtype: "q8" });
const enc = await tok(new Array(passages.length).fill(query), { text_pair: passages, padding: true, truncation: true });
const { logits } = await model(enc);
Evaluation
nDCG@10 for reranking the candidate lists of public reranking benchmarks (the MTEB versions). The benchmarks were used only
for evaluation, never for training or model selection. Every model sees the same queries and candidates, through the same
script (code/). MIRACL: 60 queries per language with 100 candidates each. Wikipedia (WikipediaRerankingMultilingual):
60 queries per language with 9 candidates each. "Other" = ESCI (es, jp, us), RuBQ, T2Reranking, mMARCO-ja and AskUbuntu.
| model | licence | MIRACL (18 languages) | Wikipedia (16 languages) | other (6 sets) | mean | note |
|---|---|---|---|---|---|---|
| this model (308M) | Apache-2.0 | 0.761 | 0.964 | 0.797 | 0.841 | multilingual |
| Horizon-Labs/multilingual-reranker-small (141M) | Apache-2.0 | 0.724 | 0.959 | 0.785 | 0.823 | multilingual |
| cross-encoder/ms-marco-MiniLM-L6-v2 (22M) | Apache-2.0 | 0.322 | 0.842 | 0.648 | 0.604 | English |
| BAAI/bge-reranker-base (278M) | MIT | 0.683 | 0.858 | 0.735 | 0.758 | Chinese/English |
| Alibaba-NLP/gte-reranker-modernbert-base (149M) | Apache-2.0 | 0.473 | 0.898 | 0.755 | 0.709 | English |
| BAAI/bge-reranker-v2-m3 (568M) | Apache-2.0 | 0.825 | 0.958 | 0.814 | 0.866 | multilingual; trained on MIRACL train |
| Qwen/Qwen3-Reranker-0.6B (596M) | Apache-2.0 | 0.785 | 0.958 | 0.793 | 0.845 | multilingual LLM reranker |
| Qwen/Qwen3-Reranker-4B (4B) | Apache-2.0 | 0.833 | 0.971 | 0.822 | 0.875 | multilingual LLM reranker; our teacher |
- The table shows the released checkpoint. Means over two training seeds: small MIRACL .730 / Wikipedia .959 / other .783 / mean .824; base .760 / .964 / .796 / .840.
- Models ahead of this one: MIRACL: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B; Wikipedia: Qwen3-Reranker-4B; other: bge-reranker-v2-m3, Qwen3-Reranker-4B. bge-reranker-v2-m3 was trained on MIRACL's training set; we were not.
Per set (nDCG@10):
| set | this model | bge-reranker-v2-m3 | Qwen3-Reranker-4B (teacher) |
|---|---|---|---|
| AskUbuntu duplicate questions (English) | 0.639 | 0.679 | 0.700 |
| ESCI product search, Spanish | 0.861 | 0.846 | 0.872 |
| ESCI product search, Japanese | 0.858 | 0.861 | 0.869 |
| ESCI product search, English | 0.891 | 0.895 | 0.908 |
| MIRACL ar | 0.803 | 0.845 | 0.876 |
| MIRACL bn | 0.805 | 0.893 | 0.896 |
| MIRACL de | 0.680 | 0.748 | 0.791 |
| MIRACL en | 0.696 | 0.728 | 0.790 |
| MIRACL es | 0.703 | 0.781 | 0.767 |
| MIRACL fa | 0.754 | 0.803 | 0.821 |
| MIRACL fi | 0.813 | 0.838 | 0.856 |
| MIRACL fr | 0.668 | 0.751 | 0.760 |
| MIRACL hi | 0.686 | 0.743 | 0.781 |
| MIRACL id | 0.574 | 0.722 | 0.641 |
| MIRACL ja | 0.762 | 0.840 | 0.873 |
| MIRACL ko | 0.784 | 0.841 | 0.856 |
| MIRACL ru | 0.790 | 0.838 | 0.868 |
| MIRACL sw | 0.833 | 0.904 | 0.859 |
| MIRACL te | 0.829 | 0.961 | 0.912 |
| MIRACL th | 0.802 | 0.878 | 0.887 |
| MIRACL yo | 0.886 | 0.935 | 0.903 |
| MIRACL zh | 0.830 | 0.809 | 0.851 |
| mMARCO (Japanese) | 0.760 | 0.811 | 0.776 |
| RuBQ (Russian) | 0.815 | 0.861 | 0.874 |
| T2Reranking (Chinese) | 0.755 | 0.746 | 0.756 |
| Wikipedia bg | 0.962 | 0.956 | 0.967 |
| Wikipedia bn | 0.968 | 0.964 | 0.971 |
| Wikipedia cs | 0.994 | 0.983 | 0.992 |
| Wikipedia da | 0.967 | 0.955 | 0.988 |
| Wikipedia de | 0.964 | 0.963 | 0.988 |
| Wikipedia en | 0.994 | 0.983 | 0.986 |
| Wikipedia fa | 0.950 | 0.953 | 0.956 |
| Wikipedia fi | 0.988 | 0.973 | 0.988 |
| Wikipedia hi | 0.911 | 0.917 | 0.914 |
| Wikipedia it | 0.961 | 0.969 | 0.966 |
| Wikipedia nl | 0.956 | 0.961 | 0.968 |
| Wikipedia no | 0.913 | 0.931 | 0.945 |
| Wikipedia pt | 0.975 | 0.949 | 0.978 |
| Wikipedia ro | 0.979 | 0.942 | 0.982 |
| Wikipedia sr | 0.984 | 0.973 | 0.983 |
| Wikipedia sv | 0.958 | 0.962 | 0.969 |
Training
- Passages: about 400k passages of 2-6 sentences from FineWeb-2 and FineWeb (ODC-BY) in 47 languages.
- Queries: Qwen3.8-27B (Apache-2.0) wrote a natural question and a keyword query for each passage, in the passage's language, plus English questions for 15% of the non-English passages (cross-lingual search).
- Scale: 334,400 queries with 16 candidates each (5.35M teacher-scored pairs), sampled evenly across languages. The data is published as Horizon-Labs/multilingual-rerank-distill.
- Schedule: 2 epochs, learning rate 3e-5 (2 epochs were chosen over 1 by validation loss, not by the benchmarks), 16 queries x 16 candidates per step, max 384 tokens per pair.
- Candidates: for each query, the 15 most similar passages in the same language by bge-m3 dense retrieval (hard negatives) plus the source passage.
- Labels: Qwen3-Reranker-4B (Apache-2.0) scored every (query, candidate) pair. The student learns the teacher's ranking with a listwise KL loss over each query's 16 candidates, plus a pointwise loss on the teacher's relevance probability. The teacher's scores are soft labels, so the near-duplicate passages that dense retrieval finds are not wrongly treated as negatives.
- Code:
code/in this repository.
Limitations
- The model inherits the teacher's judgement, including its mistakes. It is weaker than the teacher, especially on long or technical passages.
- Training queries are LLM-written questions and keyword queries over web text. Very domain-specific search (legal, medical, code) and conversational queries are less covered.
- Pairs are truncated at the max length. Rerank passage-sized chunks, not whole documents.
- Downloads last month
- -
Model tree for Horizon-Labs/multilingual-reranker-base
Base model
jhu-clsp/mmBERT-base