Text Ranking
sentence-transformers
Safetensors
Korean
English
modernbert
cross-encoder
reranker
korean
english
text-embeddings-inference
Instructions to use nlpai-lab/KURE-Reranker-nano with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nlpai-lab/KURE-Reranker-nano with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("nlpai-lab/KURE-Reranker-nano") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Update model card with the current MTEB-ko-retrieval benchmark results
#1
by yjoonjang - opened
- .gitattributes +1 -0
- README.md +86 -88
- assets/pps_vs_ndcg9.png +3 -0
.gitattributes
CHANGED
|
@@ -36,3 +36,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 36 |
assets/kure_logo.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
assets/pps_performance_2col.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
assets/size_performance_2col.png filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 36 |
assets/kure_logo.png filter=lfs diff=lfs merge=lfs -text
|
| 37 |
assets/pps_performance_2col.png filter=lfs diff=lfs merge=lfs -text
|
| 38 |
assets/size_performance_2col.png filter=lfs diff=lfs merge=lfs -text
|
| 39 |
+
assets/pps_vs_ndcg9.png filter=lfs diff=lfs merge=lfs -text
|
README.md
CHANGED
|
@@ -21,35 +21,22 @@ tags:
|
|
| 21 |
|
| 22 |
# KURE-Reranker-nano
|
| 23 |
|
| 24 |
-
**KURE-Reranker-nano** is a lightweight, approximately **150M-parameter**
|
| 25 |
|
| 26 |
## Key Characteristics
|
| 27 |
|
| 28 |
-
- **Strong Korean reranking:** 88.
|
|
|
|
|
|
|
| 29 |
- **Lightweight:** Approximately 150M parameters, based on [`skt/A.X-Encoder-base`](https://huggingface.co/skt/A.X-Encoder-base).
|
| 30 |
-
- **High throughput:** 556.27 pairs per second averaged over eight benchmarks on an NVIDIA RTX A6000 48GB, the highest average in our comparison under the measurement conditions described below.
|
| 31 |
-
- **Long-input support:** Up to **8,192 tokens** for the combined query–document input.
|
| 32 |
-
- **Simple interface:** Compatible with Sentence Transformers `CrossEncoder` and Transformers `AutoModelForSequenceClassification`; no chat template or instruction prefix is required.
|
| 33 |
-
|
| 34 |
-
## Performance and Efficiency
|
| 35 |
-
|
| 36 |
-
### Reranking Performance vs. Throughput
|
| 37 |
-
|
| 38 |
-
<img src="./assets/pps_performance_2col.png" width="700" alt="Reranking Performance vs. Throughput">
|
| 39 |
-
|
| 40 |
-
### Reranking Performance vs. Model Size
|
| 41 |
-
|
| 42 |
-
<img src="./assets/size_performance_2col.png" width="700" alt="Reranking Performance vs. Model Size">
|
| 43 |
-
|
| 44 |
-
Both figures report mean nDCG@10 across nine Korean benchmarks, including MLDR. Throughput is averaged over eight benchmarks, excluding MLDR. Model-specific input limits and the best measured batch size are used for throughput; these results do not measure equal-token-length speed or end-to-end retrieval latency.
|
| 45 |
|
| 46 |
## Model Overview
|
| 47 |
|
| 48 |
| Property | Value |
|
| 49 |
|---|---|
|
| 50 |
| Model type | Cross-encoder reranker |
|
| 51 |
-
| Backbone | A.X-Encoder-base
|
| 52 |
-
| Parameters |
|
| 53 |
| Languages | Korean and English |
|
| 54 |
| Maximum inference input length | 8,192 tokens, including query, document, and special tokens |
|
| 55 |
| Output | One scalar relevance score per query–document pair |
|
|
@@ -127,77 +114,88 @@ for index in torch.argsort(scores, descending=True).tolist():
|
|
| 127 |
|
| 128 |
## Korean Reranking Evaluation
|
| 129 |
|
| 130 |
-
Scores
|
| 131 |
-
|
| 132 |
-
| Model | Params | AutoRAG | Belebele | Ko-StrategyQA | LawIRKo | MIRACL | MrTidy | PublicHealthQA | SQuADKorV1 | MLDR |
|
| 133 |
-
|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
|
| 134 |
-
| [Dongjin-kr/ko-reranker](https://huggingface.co/Dongjin-kr/ko-reranker) | 560M | 90.14 | 97.59 | 84.68 | 73.43 | 80.17 | 77.73 | 76.75 | 97.85 | 37.21 |
|
| 135 |
-
| [BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3) | 568M | 96.63 | 98.53 | 84.87 | 79.06 | 81.30 | 82.22 | 84.75 | 98.53 | 66.90 |
|
| 136 |
-
| [dragonkue/bge-reranker-v2-m3-ko](https://huggingface.co/dragonkue/bge-reranker-v2-m3-ko) | 568M | 96.84 | 97.69 | 82.32 | 67.21 | 75.73 | 67.76 | 87.08 | 98.46 | 70.61 |
|
| 137 |
-
| [nlpai-lab/LAMAR-600m](https://huggingface.co/nlpai-lab/LAMAR-600m) | 568M | 95.91 | 98.35 | 84.61 | 78.22 | 82.14 | 79.75 | 82.25 | 98.50 | 56.79 |
|
| 138 |
-
| [upskyy/ko-reranker-8k](https://huggingface.co/upskyy/ko-reranker-8k) | 568M | 92.30 | 92.91 | 81.43 | 77.70 | 72.49 | 69.98 | 83.88 | 97.18 | 59.75 |
|
| 139 |
-
| [telepix/PIXIE-Spell-Reranker-Preview-0.6B](https://huggingface.co/telepix/PIXIE-Spell-Reranker-Preview-0.6B) | 596M | 97.94 | 97.77 | 83.29 | 60.42 | 84.49 | 76.50 | 85.34 | 98.51 | 18.29 |
|
| 140 |
-
| [tomaarsen/Qwen3-Reranker-0.6B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-0.6B-seq-cls) | 596M | 93.08 | 97.79 | 83.36 | 78.67 | 85.07 | 73.59 | 84.89 | 98.08 | 78.13 |
|
| 141 |
-
| [jinaai/jina-reranker-v3](https://huggingface.co/jinaai/jina-reranker-v3) | 597M | 97.73 | 96.95 | 85.53 | 85.46 | 84.49 | 81.04 | 79.60 | 98.59 | - |
|
| 142 |
-
| [jinaai/jina-reranker-v3.5](https://huggingface.co/jinaai/jina-reranker-v3.5) | 597M | **98.38** | 97.33 | 85.39 | 85.45 | **85.65** | 81.94 | 80.94 | 98.87 | - |
|
| 143 |
-
| [cross-encoder/ettin-reranker-1b-v1](https://huggingface.co/cross-encoder/ettin-reranker-1b-v1) | 1.03B | 89.01 | 69.14 | 66.24 | 53.06 | 70.04 | 66.59 | 74.61 | 95.90 | 36.51 |
|
| 144 |
-
| [nvidia/llama-nemotron-rerank-1b-v2](https://huggingface.co/nvidia/llama-nemotron-rerank-1b-v2) | 1.24B | 94.80 | 98.83 | 85.36 | 75.27 | 81.82 | 80.26 | 84.91 | 98.58 | 67.19 |
|
| 145 |
-
| [mixedbread-ai/mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) | 1.54B | 95.31 | 97.78 | 85.63 | 81.30 | 79.39 | **87.71** | 87.72 | 96.81 | 67.87 |
|
| 146 |
-
| [BAAI/bge-reranker-v2-gemma](https://huggingface.co/BAAI/bge-reranker-v2-gemma) | 2.51B | 94.07 | 98.57 | 86.14 | 75.72 | 83.62 | 84.30 | 86.98 | 98.58 | 28.81 |
|
| 147 |
-
| [tomaarsen/Qwen3-Reranker-4B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-4B-seq-cls) | 4.02B | 97.07 | 99.06 | **87.34** | 87.52 | 85.33 | 83.21 | 86.85 | 98.61 | 81.05 |
|
| 148 |
-
| [zeroentropy/zerank-2-reranker](https://huggingface.co/zeroentropy/zerank-2-reranker) | 4.02B | 94.36 | 98.46 | 87.12 | 86.69 | 80.03 | 80.28 | 86.46 | 97.91 | 71.20 |
|
| 149 |
-
| [lightonai/LightOn-rerank-PW-4B](https://huggingface.co/lightonai/LightOn-rerank-PW-4B) | 4.54B | 93.21 | 98.82 | 85.67 | 79.38 | 80.72 | 80.91 | 86.93 | 98.03 | 76.09 |
|
| 150 |
-
| [tomaarsen/Qwen3-Reranker-8B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-8B-seq-cls) | 7.57B | 95.46 | **99.07** | 86.79 | **90.14** | 84.90 | 84.09 | **88.93** | 98.80 | **82.20** |
|
| 151 |
-
| **KURE-Reranker-nano (Ours)** | 150M | 97.41 | 98.36 | 85.83 | 83.76 | 84.03 | 79.99 | 84.80 | **98.89** | 79.78 |
|
| 152 |
-
|
| 153 |
Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
|
| 154 |
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
|
| 158 |
-
|
| 159 |
-
|
| 160 |
-
|
| 161 |
-
|
| 162 |
-
|
| 163 |
-
|
| 164 |
-
|
| 165 |
-
|
| 166 |
-
|
| 167 |
-
|
| 168 |
-
|
| 169 |
-
|
| 170 |
-
|
| 171 |
-
|
| 172 |
-
|
| 173 |
-
|
| 174 |
-
|
| 175 |
-
|
| 176 |
-
|
| 177 |
-
|
| 178 |
-
|
| 179 |
-
|
|
| 180 |
-
|
| 181 |
-
|
| 182 |
-
|
| 183 |
-
|
| 184 |
-
|
| 185 |
-
|
| 186 |
-
|
| 187 |
-
|
|
| 188 |
-
|
|
| 189 |
-
|
|
| 190 |
-
|
|
| 191 |
-
|
|
| 192 |
-
|
|
| 193 |
-
|
|
| 194 |
-
|
|
| 195 |
-
|
|
| 196 |
-
|
|
| 197 |
-
|
| 198 |
-
|
| 199 |
-
|
| 200 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 201 |
|
| 202 |
## Citation
|
| 203 |
|
|
|
|
| 21 |
|
| 22 |
# KURE-Reranker-nano
|
| 23 |
|
| 24 |
+
**KURE-Reranker-nano** is a lightweight, approximately **150M-parameter** Korean-English bilingual reranker. It jointly encodes a query and a candidate document and predicts a scalar relevance score.
|
| 25 |
|
| 26 |
## Key Characteristics
|
| 27 |
|
| 28 |
+
- **Strong Korean reranking:** 88.08 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 1B parameters and only 0.42 points below KURE-Reranker-base (1.7B).
|
| 29 |
+
- **Long-document reranking:** 79.66 nDCG@10 on MultiLongDocRetrieval (MLDR), the best among the evaluated models with fewer than 1B parameters.
|
| 30 |
+
- **High throughput:** 473.7 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 19.5× Qwen3-Reranker-4B under the measurement conditions described below.
|
| 31 |
- **Lightweight:** Approximately 150M parameters, based on [`skt/A.X-Encoder-base`](https://huggingface.co/skt/A.X-Encoder-base).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 32 |
|
| 33 |
## Model Overview
|
| 34 |
|
| 35 |
| Property | Value |
|
| 36 |
|---|---|
|
| 37 |
| Model type | Cross-encoder reranker |
|
| 38 |
+
| Backbone | A.X-Encoder-base |
|
| 39 |
+
| Parameters | 149M |
|
| 40 |
| Languages | Korean and English |
|
| 41 |
| Maximum inference input length | 8,192 tokens, including query, document, and special tokens |
|
| 42 |
| Output | One scalar relevance score per query–document pair |
|
|
|
|
| 114 |
|
| 115 |
## Korean Reranking Evaluation
|
| 116 |
|
| 117 |
+
Scores are **nDCG@10**, and throughput is pairs per second (**PPS**).
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 118 |
Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
|
| 119 |
|
| 120 |
+
<img src="./assets/pps_vs_ndcg9.png" width="700" alt="Reranking performance (mean nDCG@10) vs. throughput (mean PPS)">
|
| 121 |
+
|
| 122 |
+
The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. PPS uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency.
|
| 123 |
+
|
| 124 |
+
### Results — MTEB-ko-retrieval (9 subsets)
|
| 125 |
+
<!-- **공식 9개 subset 을 모두 평가한 모델**의 9-subset mean NDCG@10, PPS -->
|
| 126 |
+
| Model | Params | Mean NDCG@10 | Mean PPS |
|
| 127 |
+
|---|---|---|---|
|
| 128 |
+
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.9004 | 15.1 |
|
| 129 |
+
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8956 | 24.3 |
|
| 130 |
+
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8849 | 55.1 |
|
| 131 |
+
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8808 | 473.7 |
|
| 132 |
+
| zeroentropy/zerank-2-reranker | 4.0B | 0.8695 | 29.4 |
|
| 133 |
+
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8664 | 15.0 |
|
| 134 |
+
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8661 | 65.2 |
|
| 135 |
+
| BAAI/bge-reranker-v2-m3 | 568M | 0.8586 | 404.1 |
|
| 136 |
+
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8585 | 99.6 |
|
| 137 |
+
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8522 | 127.3 |
|
| 138 |
+
| nlpai-lab/LAMAR-600m | 568M | 0.8406 | 408.8 |
|
| 139 |
+
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8263 | 401.5 |
|
| 140 |
+
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8186 | 58.7 |
|
| 141 |
+
| upskyy/ko-reranker-8k | 568M | 0.8085 | 404.0 |
|
| 142 |
+
| Dongjin-kr/ko-reranker | 560M | 0.7950 | 482.8 |
|
| 143 |
+
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.7806 | 99.6 |
|
| 144 |
+
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6901 | 49.7 |
|
| 145 |
+
|
| 146 |
+
PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks.Inputs are length-sorted and dynamically padded within each batch.
|
| 147 |
+
|
| 148 |
+
`jinaai/jina-reranker-v3` and `jinaai/jina-reranker-v3.5` are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.
|
| 149 |
+
|
| 150 |
+
### Per-dataset NDCG@10
|
| 151 |
+
|
| 152 |
+
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|
| 153 |
+
|---|---|---|---|---|---|---|---|---|---|---|
|
| 154 |
+
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.8679 | 0.9546 | 0.8893 | 0.9907 | 0.8490 | 0.8409 | 0.8220 | 0.9880 | 0.9014 |
|
| 155 |
+
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8733 | 0.9707 | 0.8685 | 0.9906 | 0.8533 | 0.8321 | 0.8105 | 0.9861 | 0.8752 |
|
| 156 |
+
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8548 | 0.9762 | 0.8716 | 0.9880 | 0.8371 | 0.7970 | 0.8163 | 0.9891 | 0.8343 |
|
| 157 |
+
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8582 | 0.9741 | 0.8475 | 0.9830 | 0.8400 | 0.8011 | 0.7966 | 0.9895 | 0.8368 |
|
| 158 |
+
| jinaai/jina-reranker-v3.5 | 597M | 0.8539 | 0.9838 | 0.8094 | 0.9733 | 0.8565 | 0.8194 | — | 0.9887 | 0.8545 |
|
| 159 |
+
| jinaai/jina-reranker-v3 | 597M | 0.8553 | 0.9773 | 0.7960 | 0.9695 | 0.8449 | 0.8104 | — | 0.9859 | 0.8546 |
|
| 160 |
+
| zeroentropy/zerank-2-reranker | 4.0B | 0.8712 | 0.9436 | 0.8646 | 0.9846 | 0.8003 | 0.8027 | 0.7120 | 0.9791 | 0.8669 |
|
| 161 |
+
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8567 | 0.9321 | 0.8693 | 0.9882 | 0.8072 | 0.8091 | 0.7609 | 0.9803 | 0.7938 |
|
| 162 |
+
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8563 | 0.9531 | 0.8772 | 0.9778 | 0.7939 | 0.8771 | 0.6787 | 0.9681 | 0.8130 |
|
| 163 |
+
| BAAI/bge-reranker-v2-m3 | 568M | 0.8487 | 0.9663 | 0.8475 | 0.9853 | 0.8129 | 0.8222 | 0.6690 | 0.9853 | 0.7906 |
|
| 164 |
+
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8336 | 0.9308 | 0.8489 | 0.9779 | 0.8507 | 0.7359 | 0.7813 | 0.9808 | 0.7867 |
|
| 165 |
+
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8536 | 0.9480 | 0.8491 | 0.9883 | 0.8182 | 0.8026 | 0.6719 | 0.9858 | 0.7527 |
|
| 166 |
+
| nlpai-lab/LAMAR-600m | 568M | 0.8461 | 0.9591 | 0.8225 | 0.9835 | 0.8214 | 0.7975 | 0.5679 | 0.9850 | 0.7822 |
|
| 167 |
+
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8232 | 0.9684 | 0.8708 | 0.9769 | 0.7573 | 0.6776 | 0.7061 | 0.9846 | 0.6721 |
|
| 168 |
+
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8614 | 0.9407 | 0.8698 | 0.9857 | 0.8362 | 0.8429 | 0.2881 | 0.9858 | 0.7572 |
|
| 169 |
+
| upskyy/ko-reranker-8k | 568M | 0.8143 | 0.9230 | 0.8388 | 0.9291 | 0.7249 | 0.6998 | 0.5975 | 0.9718 | 0.7770 |
|
| 170 |
+
| Dongjin-kr/ko-reranker | 560M | 0.8468 | 0.9014 | 0.7675 | 0.9759 | 0.8017 | 0.7772 | 0.3721 | 0.9785 | 0.7343 |
|
| 171 |
+
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.8329 | 0.9794 | 0.8534 | 0.9777 | 0.8449 | 0.7650 | 0.1829 | 0.9850 | 0.6042 |
|
| 172 |
+
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6624 | 0.8901 | 0.7461 | 0.6914 | 0.7004 | 0.6659 | 0.3651 | 0.9590 | 0.5306 |
|
| 173 |
+
|
| 174 |
+
### Per-dataset PPS
|
| 175 |
+
|
| 176 |
+
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|
| 177 |
+
|---|---|---|---|---|---|---|---|---|---|---|
|
| 178 |
+
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 16.6 | 7.3 | 17.5 | 19.4 | 24.9 | 26.5 | 0.7 | 10.6 | 12.3 |
|
| 179 |
+
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 26.1 | 11.8 | 28.4 | 31.5 | 40.0 | 42.4 | 1.1 | 17.2 | 20.0 |
|
| 180 |
+
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 59.2 | 27.3 | 64.6 | 71.1 | 89.1 | 97.1 | 2.6 | 39.1 | 45.7 |
|
| 181 |
+
| **nlpai-lab/KURE-Reranker-nano** | 149M | 475.4 | 233.2 | 569.7 | 630.9 | 765.8 | 855.2 | 16.5 | 330.8 | 386.0 |
|
| 182 |
+
| jinaai/jina-reranker-v3.5 | 597M | 141.0 | 38.6 | 148.6 | 174.6 | 277.3 | 295.7 | — | 68.4 | 86.7 |
|
| 183 |
+
| jinaai/jina-reranker-v3 | 597M | 113.8 | 25.0 | 121.8 | 169.2 | 244.6 | 261.6 | — | 48.7 | 64.0 |
|
| 184 |
+
| zeroentropy/zerank-2-reranker | 4.0B | 31.0 | 12.7 | 33.8 | 37.8 | 51.0 | 56.4 | 1.1 | 18.8 | 22.3 |
|
| 185 |
+
| lightonai/LightOn-rerank-PW-4B | 4.5B | 16.2 | 7.3 | 18.6 | 19.5 | 23.4 | 26.1 | 0.7 | 10.7 | 12.4 |
|
| 186 |
+
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 70.0 | 32.9 | 76.8 | 84.1 | 104.2 | 113.6 | 3.2 | 47.0 | 54.7 |
|
| 187 |
+
| BAAI/bge-reranker-v2-m3 | 568M | 420.7 | 195.6 | 471.3 | 545.3 | 653.2 | 724.8 | 9.3 | 271.6 | 345.5 |
|
| 188 |
+
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 107.1 | 49.6 | 118.0 | 129.3 | 159.5 | 173.9 | 4.2 | 71.8 | 82.9 |
|
| 189 |
+
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 133.7 | 54.4 | 146.9 | 163.0 | 220.0 | 243.1 | 3.9 | 84.9 | 95.9 |
|
| 190 |
+
| nlpai-lab/LAMAR-600m | 568M | 418.0 | 196.5 | 480.5 | 557.2 | 656.7 | 742.5 | 9.2 | 275.0 | 343.8 |
|
| 191 |
+
| dragonkue/bge-reranker-v2-m3-ko | 568M | 416.9 | 195.1 | 467.2 | 540.7 | 653.4 | 716.7 | 9.2 | 271.9 | 342.7 |
|
| 192 |
+
| BAAI/bge-reranker-v2-gemma | 2.5B | 64.7 | 26.1 | 67.5 | 75.8 | 99.8 | 106.8 | 2.5 | 40.0 | 44.7 |
|
| 193 |
+
| upskyy/ko-reranker-8k | 568M | 411.8 | 195.6 | 474.3 | 551.4 | 649.3 | 726.0 | 9.2 | 275.1 | 342.8 |
|
| 194 |
+
| Dongjin-kr/ko-reranker | 560M | 520.8 | 258.0 | 510.1 | 554.8 | 747.6 | 765.2 | 272.7 | 334.4 | 381.2 |
|
| 195 |
+
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 104.6 | 49.4 | 117.8 | 129.8 | 160.5 | 175.6 | 4.3 | 72.1 | 82.6 |
|
| 196 |
+
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 52.4 | 19.9 | 51.4 | 65.6 | 92.7 | 97.2 | 3.8 | 31.2 | 33.4 |
|
| 197 |
+
|
| 198 |
+
Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the [Ettin reranker speed benchmark](https://huggingface.co/blog/ettin-reranker#speed).
|
| 199 |
|
| 200 |
## Citation
|
| 201 |
|
assets/pps_vs_ndcg9.png
ADDED
|
Git LFS Details
|