Multilingual Reranker (base, 308M)

A cross-encoder reranker for search and RAG in 40+ languages, built on mmBERT-base. Give it a query and candidate passages, for example the top 20-100 hits of a vector or keyword search. It scores each pair, so you can re-sort the candidates and keep the best ones for your LLM. Queries and passages can be in different languages. It is distilled from Qwen3-Reranker-4B into a model about 13x smaller. Apache-2.0, trained on openly licensed web text. ONNX files for CPU and the browser (transformers.js) are included. A smaller, faster version is available as Horizon-Labs/multilingual-reranker-small. Try it in the browser.

  • One output logit per (query, passage) pair: higher = more relevant. sigmoid(logit) gives a 0-1 relevance score.
  • Max length: trained with 384 tokens per pair. Longer passages are truncated; split long documents into chunks.
  • onnx/model_quantized.onnx (int8 embeddings, 641 MB): scores differ from fp32 by 0.001 on average (sigmoid scale) over 1240 benchmark pairs.

Usage

sentence-transformers:

from sentence_transformers import CrossEncoder

model = CrossEncoder("Horizon-Labs/multilingual-reranker-base")
query = "How tall is the Eiffel Tower?"
passages = ["The Eiffel Tower is 330 metres tall.", "La tour Eiffel a été construite pour l'Exposition universelle de 1889.",
            "The Statue of Liberty is 93 metres tall."]
print(model.rank(query, passages))   # [{'corpus_id': 0, 'score': ...}, ...]

transformers:

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

tok = AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-base")
model = AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-base").eval()
enc = tok([query] * len(passages), passages, padding=True, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    scores = model(**enc).logits[:, 0].sigmoid()

transformers.js:

import { AutoTokenizer, AutoModelForSequenceClassification } from "@huggingface/transformers";
const tok = await AutoTokenizer.from_pretrained("Horizon-Labs/multilingual-reranker-base");
const model = await AutoModelForSequenceClassification.from_pretrained("Horizon-Labs/multilingual-reranker-base", { dtype: "q8" });
const enc = await tok(new Array(passages.length).fill(query), { text_pair: passages, padding: true, truncation: true });
const { logits } = await model(enc);

Evaluation

nDCG@10 for reranking the candidate lists of public reranking benchmarks (the MTEB versions). The benchmarks were used only for evaluation, never for training or model selection. Every model sees the same queries and candidates, through the same script (code/). MIRACL: 60 queries per language with 100 candidates each. Wikipedia (WikipediaRerankingMultilingual): 60 queries per language with 9 candidates each. "Other" = ESCI (es, jp, us), RuBQ, T2Reranking, mMARCO-ja and AskUbuntu.

model licence MIRACL (18 languages) Wikipedia (16 languages) other (6 sets) mean note
this model (308M) Apache-2.0 0.761 0.964 0.797 0.841 multilingual
Horizon-Labs/multilingual-reranker-small (141M) Apache-2.0 0.724 0.959 0.785 0.823 multilingual
cross-encoder/ms-marco-MiniLM-L6-v2 (22M) Apache-2.0 0.322 0.842 0.648 0.604 English
BAAI/bge-reranker-base (278M) MIT 0.683 0.858 0.735 0.758 Chinese/English
Alibaba-NLP/gte-reranker-modernbert-base (149M) Apache-2.0 0.473 0.898 0.755 0.709 English
BAAI/bge-reranker-v2-m3 (568M) Apache-2.0 0.825 0.958 0.814 0.866 multilingual; trained on MIRACL train
Qwen/Qwen3-Reranker-0.6B (596M) Apache-2.0 0.785 0.958 0.793 0.845 multilingual LLM reranker
Qwen/Qwen3-Reranker-4B (4B) Apache-2.0 0.833 0.971 0.822 0.875 multilingual LLM reranker; our teacher
  • The table shows the released checkpoint. Means over two training seeds: small MIRACL .730 / Wikipedia .959 / other .783 / mean .824; base .760 / .964 / .796 / .840.
  • Models ahead of this one: MIRACL: bge-reranker-v2-m3, Qwen3-Reranker-0.6B, Qwen3-Reranker-4B; Wikipedia: Qwen3-Reranker-4B; other: bge-reranker-v2-m3, Qwen3-Reranker-4B. bge-reranker-v2-m3 was trained on MIRACL's training set; we were not.

Per set (nDCG@10):

set this model bge-reranker-v2-m3 Qwen3-Reranker-4B (teacher)
AskUbuntu duplicate questions (English) 0.639 0.679 0.700
ESCI product search, Spanish 0.861 0.846 0.872
ESCI product search, Japanese 0.858 0.861 0.869
ESCI product search, English 0.891 0.895 0.908
MIRACL ar 0.803 0.845 0.876
MIRACL bn 0.805 0.893 0.896
MIRACL de 0.680 0.748 0.791
MIRACL en 0.696 0.728 0.790
MIRACL es 0.703 0.781 0.767
MIRACL fa 0.754 0.803 0.821
MIRACL fi 0.813 0.838 0.856
MIRACL fr 0.668 0.751 0.760
MIRACL hi 0.686 0.743 0.781
MIRACL id 0.574 0.722 0.641
MIRACL ja 0.762 0.840 0.873
MIRACL ko 0.784 0.841 0.856
MIRACL ru 0.790 0.838 0.868
MIRACL sw 0.833 0.904 0.859
MIRACL te 0.829 0.961 0.912
MIRACL th 0.802 0.878 0.887
MIRACL yo 0.886 0.935 0.903
MIRACL zh 0.830 0.809 0.851
mMARCO (Japanese) 0.760 0.811 0.776
RuBQ (Russian) 0.815 0.861 0.874
T2Reranking (Chinese) 0.755 0.746 0.756
Wikipedia bg 0.962 0.956 0.967
Wikipedia bn 0.968 0.964 0.971
Wikipedia cs 0.994 0.983 0.992
Wikipedia da 0.967 0.955 0.988
Wikipedia de 0.964 0.963 0.988
Wikipedia en 0.994 0.983 0.986
Wikipedia fa 0.950 0.953 0.956
Wikipedia fi 0.988 0.973 0.988
Wikipedia hi 0.911 0.917 0.914
Wikipedia it 0.961 0.969 0.966
Wikipedia nl 0.956 0.961 0.968
Wikipedia no 0.913 0.931 0.945
Wikipedia pt 0.975 0.949 0.978
Wikipedia ro 0.979 0.942 0.982
Wikipedia sr 0.984 0.973 0.983
Wikipedia sv 0.958 0.962 0.969

Training

  • Passages: about 400k passages of 2-6 sentences from FineWeb-2 and FineWeb (ODC-BY) in 47 languages.
  • Queries: Qwen3.8-27B (Apache-2.0) wrote a natural question and a keyword query for each passage, in the passage's language, plus English questions for 15% of the non-English passages (cross-lingual search).
  • Scale: 334,400 queries with 16 candidates each (5.35M teacher-scored pairs), sampled evenly across languages. The data is published as Horizon-Labs/multilingual-rerank-distill.
  • Schedule: 2 epochs, learning rate 3e-5 (2 epochs were chosen over 1 by validation loss, not by the benchmarks), 16 queries x 16 candidates per step, max 384 tokens per pair.
  • Candidates: for each query, the 15 most similar passages in the same language by bge-m3 dense retrieval (hard negatives) plus the source passage.
  • Labels: Qwen3-Reranker-4B (Apache-2.0) scored every (query, candidate) pair. The student learns the teacher's ranking with a listwise KL loss over each query's 16 candidates, plus a pointwise loss on the teacher's relevance probability. The teacher's scores are soft labels, so the near-duplicate passages that dense retrieval finds are not wrongly treated as negatives.
  • Code: code/ in this repository.

Limitations

  • The model inherits the teacher's judgement, including its mistakes. It is weaker than the teacher, especially on long or technical passages.
  • Training queries are LLM-written questions and keyword queries over web text. Very domain-specific search (legal, medical, code) and conversational queries are less covered.
  • Pairs are truncated at the max length. Rerank passage-sized chunks, not whole documents.
Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Horizon-Labs/multilingual-reranker-base

Quantized
(277)
this model

Datasets used to train Horizon-Labs/multilingual-reranker-base

Collection including Horizon-Labs/multilingual-reranker-base