Instructions to use ALJIACHI/Mizan-Rerank-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ALJIACHI/Mizan-Rerank-v1 with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("ALJIACHI/Mizan-Rerank-v1") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Mizan-Rerank-v1
A compact open-source cross-encoder for reranking Arabic long texts, built for maximum accuracy per parameter.
Deprecated: use Mizan-Rerank-v3
This model is deprecated and no longer maintained. Use Mizan-Rerank-v3 instead.
v3 is a 306M-parameter cross-encoder fine-tuned from Alibaba-NLP/gte-multilingual-reranker-base on 71k Arabic listwise groups, with adversarial traps (negated rulings, swapped entities, shifted numbers and dates) as hard negatives. It is the most accurate Arabic reranker in the Mizan family, it runs about twice as fast as bge-reranker-v2-m3, and it is built for RAG pipelines, long-document retrieval and domain-specific Arabic search.
v3 covers the use cases v1 was built for, at 306M parameters and full speed, so there is no reason to start a new project on v1. This card is kept online for reproducibility of existing work.
v1 vs. v3 at a glance
| Mizan-Rerank-v1 | Mizan-Rerank-v3 | |
|---|---|---|
| Status | deprecated | recommended |
| Parameters | 149M | 306M |
| Base model | answerdotai/ModernBERT-base | Alibaba-NLP/gte-multilingual-reranker-base |
| Max sequence length | 8192 | 8192 (trained to 3072) |
| Training data | 741M tokens | 71,044 Arabic listwise groups with adversarial hard negatives |
| Ranking quality | baseline | best in the Mizan family |
Migrating is a one-line change — see Migrating to v3 below.
Overview
Mizan-Rerank-v1 is an open-source reranking model based on the ModernBERT architecture, designed for reordering search results over Arabic text. At only 149 million parameters it targets the efficiency end of the spectrum: it runs comfortably in about 1 GB of memory while still handling documents up to 8192 tokens.
Key Features
- Lightweight: 149M parameters, roughly half to a quarter the size of common multilingual rerankers (278M–568M)
- Long text processing: Accepts up to 8192 tokens
- Fast inference: Around 0.1 s response time on a single RTX 4090
- Arabic language optimization: Trained specifically for Arabic morphology and context
- Low memory footprint: ~1 GB VRAM, making it viable on modest GPUs and CPU-bound deployments
Performance Benchmarks
Hardware Performance (RTX 4090 24GB)
| Model | Parameters | RAM Usage | Response Time |
|---|---|---|---|
| Mizan-Rerank-v1 | 149M | 1 GB | 0.1 s |
| jina-reranker-v2-base-multilingual | 278M | 2.5 GB | 0.2 s |
| bge-reranker-v2-m3 | 568M | 4 GB | 0.3 s |
Efficiency was v1's primary advantage when it was released, and it remains the smallest model in the Mizan family. It is no longer a reason to adopt it: v3 is more accurate and faster per pair than bge-reranker-v2-m3 at 306M parameters.
Ranking Quality (ndcg@10)
The table below comes from the evaluation run used for the v2 release, so all models are measured under identical conditions and the numbers are directly comparable. It predates v3.
| Model | Parameters | Reranking | Triplet | MIRACL (Long Docs) | WikiQA | MedQA |
|---|---|---|---|---|---|---|
| ALJIACHI/Mizan-Rerank-v2 | 305M | 1.0000 | 0.9993 | 0.8091 | 0.8258 | 0.6775 |
| BAAI/bge-reranker-v2-m3 | 568M | 1.0000 | 0.9998 | 0.7231 | 0.8669 | 0.6584 |
| Alibaba-NLP/gte-multilingual-reranker-base | 305M | 1.0000 | 0.9991 | 0.7539 | 0.8275 | 0.6648 |
| ALJIACHI/Mizan-Rerank-v1 | 149M | 0.9986 | 0.9955 | 0.7370 | 0.7739 | 0.5502 |
v1 stays competitive with models 2–4× its size on the saturated reranking and triplet sets. The gap widens on long-document MIRACL and on domain-specific sets such as MedQA, which is what the later Mizan models were built to close.
Note on earlier figures: the original v1 model card reported a MIRACL ndcg@10 of 0.8865, measured on a different evaluation subset and protocol. The table above supersedes it and should be used for any comparison against later rerankers.
Training Methodology
Mizan-Rerank-v1 was trained on a corpus of 741,159,981 tokens drawn from:
- Authentic Arabic open-source datasets
- Manually curated and processed text
- Purpose-generated synthetic data
How It Works
- Query reception: The model receives a user query alongside candidate passages
- Content analysis: It jointly encodes each query–passage pair and analyzes their semantic relationship
- Relevance scoring: It assigns a relevance score to each passage
- Reranking: Passages are sorted by descending score
Usage
Using Transformers
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
model = AutoModelForSequenceClassification.from_pretrained("ALJIACHI/Mizan-Rerank-v1")
tokenizer = AutoTokenizer.from_pretrained("ALJIACHI/Mizan-Rerank-v1")
model.eval()
def get_relevance_score(query, passage):
inputs = tokenizer(
query, passage,
return_tensors="pt", padding=True, truncation=True, max_length=8192,
)
with torch.no_grad():
logits = model(**inputs).logits
# Apply sigmoid to map the raw logit into a 0–1 relevance probability
return torch.sigmoid(logits).item()
query = "ما هو تفسير الآية وجعلنا من الماء كل شيء حي"
passages = [
"تعني الآية أن الماء هو عنصر أساسي في حياة جميع الكائنات الحية، وهو ضروري لاستمرار الحياة.",
"تم اكتشاف كواكب خارج المجموعة الشمسية تحتوي على مياه متجمدة.",
"تحدث القرآن الكريم عن البرق والرعد في عدة مواضع مختلفة.",
]
scores = [(p, get_relevance_score(query, p)) for p in passages]
reranked = sorted(scores, key=lambda x: x[1], reverse=True)
for passage, score in reranked:
print(f"Score: {score:.4f} | {passage}")
Using Sentence Transformers
pip install -U sentence-transformers
from sentence_transformers import CrossEncoder
model = CrossEncoder("ALJIACHI/Mizan-Rerank-v1", max_length=8192)
ranks = model.rank(
"ما هي فوائد فيتامين د؟",
[
"يساعد فيتامين د في تعزيز صحة العظام وتقوية الجهاز المناعي، كما يلعب دوراً مهماً في امتصاص الكالسيوم.",
"يستخدم فيتامين د في بعض الصناعات الغذائية كمادة حافظة.",
"يمكن الحصول على فيتامين د من خلال التعرض لأشعة الشمس أو تناول مكملات غذائية.",
],
)
print(ranks)
Migrating to v3
The interface is identical; only the model ID changes, plus trust_remote_code=True for v3's custom architecture:
from sentence_transformers import CrossEncoder
# Before
model = CrossEncoder("ALJIACHI/Mizan-Rerank-v1", max_length=8192)
# After
model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3", max_length=3072, trust_remote_code=True)
Scores are not calibrated across versions, so if your pipeline uses an absolute score threshold to filter candidates, re-tune that threshold after switching. Full details, benchmarks and a usage guide are on the v3 model card.
Practical Examples
Example 1
السؤال: ما هو القانون الجديد بشأن الضرائب في 2024؟
| النص | الدرجة |
|---|---|
| نشرت الجريدة الرسمية قانوناً جديداً في 2024 ينص على زيادة الضرائب على الشركات الكبرى بنسبة 5% | 0.9989 |
| الضرائب تعد مصدراً مهماً للدخل القومي وتختلف نسبتها من دولة إلى أخرى. | 0.0001 |
| افتتحت الحكومة مشروعاً جديداً للطاقة المتجددة في 2024. | 0.0001 |
Example 2
السؤال: ما هو تفسير الآية وجعلنا من الماء كل شيء حي
| النص | الدرجة |
|---|---|
| تعني الآية أن الماء هو عنصر أساسي في حياة جميع الكائنات الحية، وهو ضروري لاستمرار الحياة. | 0.9996 |
| تم اكتشاف كواكب خارج المجموعة الشمسية تحتوي على مياه متجمدة. | 0.0000 |
| تحدث القرآن الكريم عن البرق والرعد في عدة مواضع مختلفة. | 0.0000 |
Example 3
السؤال: ما هي فوائد فيتامين د؟
| النص | الدرجة |
|---|---|
| يساعد فيتامين د في تعزيز صحة العظام وتقوية الجهاز المناعي، كما يلعب دوراً مهماً في امتصاص الكالسيوم. | 0.9991 |
| يستخدم فيتامين د في بعض الصناعات الغذائية كمادة حافظة. | 0.9941 |
| يمكن الحصول على فيتامين د من خلال التعرض لأشعة الشمس أو تناول مكملات غذائية. | 0.9938 |
Example 3 illustrates a known v1 weakness: it separates relevant from irrelevant text reliably, but its scores are less discriminative among passages that merely mention the query topic. Later Mizan models handle this case better.
Model Details
- Model Type: Cross Encoder
- Base Model: answerdotai/ModernBERT-base
- Parameters: 149M
- Maximum Sequence Length: 8192 tokens
- Number of Output Labels: 1
- Language: Arabic (ar), English (en)
- License: Apache 2.0
- Status: deprecated
- Replacement: ALJIACHI/Mizan-Rerank-v3
Applications
- Specialized Arabic search engines
- Archiving systems and digital libraries
- RAG (Retrieval-Augmented Generation) pipelines
- Conversational AI applications
- E-learning platforms
- Information retrieval systems
For all of the above, use v3. This model is deprecated and is documented here only for reproducibility of existing work.
Citation
If you use Mizan-Rerank-v1 in your research, please cite:
@software{Mizan_Rerank_v1_2025,
author = {Ali Aljiachi},
title = {Mizan-Rerank-v1: A Revolutionary Arabic Text Reranking Model},
year = {2025},
publisher = {Hugging Face},
url = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v1}
}
For the recommended successor model:
@software{Mizan_Rerank_v3_2026,
author = {Ali Aljiachi},
title = {Mizan-Rerank-v3: Adversarially Trained Arabic Long-Context Reranker},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v3}
}
Base model:
@misc{modernbert,
title={Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference},
author={Benjamin Warner and Antoine Chaffin and Benjamin Clavié and Orion Weller and Oskar Hallström and Said Taghadouini and Alexis Gallagher and Raja Biswas and Faisal Ladhak and Tom Aarsen and Nathan Cooper and Griffin Adams and Jeremy Howard and Iacopo Poli},
year={2024},
eprint={2412.13663},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2412.13663},
}
License
The Mizan-Rerank-v1 model weights are released under the Apache 2.0 License.
- Downloads last month
- 223
Model tree for ALJIACHI/Mizan-Rerank-v1
Base model
answerdotai/ModernBERT-base