Update model card with the current MTEB-ko-retrieval benchmark results

#1
by yjoonjang - opened
Files changed (3) hide show
  1. .gitattributes +1 -0
  2. README.md +86 -88
  3. assets/pps_vs_ndcg9.png +3 -0
.gitattributes CHANGED
@@ -36,3 +36,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
36
  assets/kure_logo.png filter=lfs diff=lfs merge=lfs -text
37
  assets/pps_performance_2col.png filter=lfs diff=lfs merge=lfs -text
38
  assets/size_performance_2col.png filter=lfs diff=lfs merge=lfs -text
 
 
36
  assets/kure_logo.png filter=lfs diff=lfs merge=lfs -text
37
  assets/pps_performance_2col.png filter=lfs diff=lfs merge=lfs -text
38
  assets/size_performance_2col.png filter=lfs diff=lfs merge=lfs -text
39
+ assets/pps_vs_ndcg9.png filter=lfs diff=lfs merge=lfs -text
README.md CHANGED
@@ -21,35 +21,22 @@ tags:
21
 
22
  # KURE-Reranker-nano
23
 
24
- **KURE-Reranker-nano** is a lightweight, approximately **150M-parameter** cross-encoder for Korean reranking. It jointly encodes a query and a candidate document and predicts a scalar relevance score.
25
 
26
  ## Key Characteristics
27
 
28
- - **Strong Korean reranking:** 88.10 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 2B parameters.
 
 
29
  - **Lightweight:** Approximately 150M parameters, based on [`skt/A.X-Encoder-base`](https://huggingface.co/skt/A.X-Encoder-base).
30
- - **High throughput:** 556.27 pairs per second averaged over eight benchmarks on an NVIDIA RTX A6000 48GB, the highest average in our comparison under the measurement conditions described below.
31
- - **Long-input support:** Up to **8,192 tokens** for the combined query–document input.
32
- - **Simple interface:** Compatible with Sentence Transformers `CrossEncoder` and Transformers `AutoModelForSequenceClassification`; no chat template or instruction prefix is required.
33
-
34
- ## Performance and Efficiency
35
-
36
- ### Reranking Performance vs. Throughput
37
-
38
- <img src="./assets/pps_performance_2col.png" width="700" alt="Reranking Performance vs. Throughput">
39
-
40
- ### Reranking Performance vs. Model Size
41
-
42
- <img src="./assets/size_performance_2col.png" width="700" alt="Reranking Performance vs. Model Size">
43
-
44
- Both figures report mean nDCG@10 across nine Korean benchmarks, including MLDR. Throughput is averaged over eight benchmarks, excluding MLDR. Model-specific input limits and the best measured batch size are used for throughput; these results do not measure equal-token-length speed or end-to-end retrieval latency.
45
 
46
  ## Model Overview
47
 
48
  | Property | Value |
49
  |---|---|
50
  | Model type | Cross-encoder reranker |
51
- | Backbone | A.X-Encoder-base (ModernBERT) |
52
- | Parameters | Approximately 150M |
53
  | Languages | Korean and English |
54
  | Maximum inference input length | 8,192 tokens, including query, document, and special tokens |
55
  | Output | One scalar relevance score per query–document pair |
@@ -127,77 +114,88 @@ for index in torch.argsort(scores, descending=True).tolist():
127
 
128
  ## Korean Reranking Evaluation
129
 
130
- Scores below are **nDCG@10**.
131
-
132
- | Model | Params | AutoRAG | Belebele | Ko-StrategyQA | LawIRKo | MIRACL | MrTidy | PublicHealthQA | SQuADKorV1 | MLDR |
133
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
134
- | [Dongjin-kr/ko-reranker](https://huggingface.co/Dongjin-kr/ko-reranker) | 560M | 90.14 | 97.59 | 84.68 | 73.43 | 80.17 | 77.73 | 76.75 | 97.85 | 37.21 |
135
- | [BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3) | 568M | 96.63 | 98.53 | 84.87 | 79.06 | 81.30 | 82.22 | 84.75 | 98.53 | 66.90 |
136
- | [dragonkue/bge-reranker-v2-m3-ko](https://huggingface.co/dragonkue/bge-reranker-v2-m3-ko) | 568M | 96.84 | 97.69 | 82.32 | 67.21 | 75.73 | 67.76 | 87.08 | 98.46 | 70.61 |
137
- | [nlpai-lab/LAMAR-600m](https://huggingface.co/nlpai-lab/LAMAR-600m) | 568M | 95.91 | 98.35 | 84.61 | 78.22 | 82.14 | 79.75 | 82.25 | 98.50 | 56.79 |
138
- | [upskyy/ko-reranker-8k](https://huggingface.co/upskyy/ko-reranker-8k) | 568M | 92.30 | 92.91 | 81.43 | 77.70 | 72.49 | 69.98 | 83.88 | 97.18 | 59.75 |
139
- | [telepix/PIXIE-Spell-Reranker-Preview-0.6B](https://huggingface.co/telepix/PIXIE-Spell-Reranker-Preview-0.6B) | 596M | 97.94 | 97.77 | 83.29 | 60.42 | 84.49 | 76.50 | 85.34 | 98.51 | 18.29 |
140
- | [tomaarsen/Qwen3-Reranker-0.6B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-0.6B-seq-cls) | 596M | 93.08 | 97.79 | 83.36 | 78.67 | 85.07 | 73.59 | 84.89 | 98.08 | 78.13 |
141
- | [jinaai/jina-reranker-v3](https://huggingface.co/jinaai/jina-reranker-v3) | 597M | 97.73 | 96.95 | 85.53 | 85.46 | 84.49 | 81.04 | 79.60 | 98.59 | - |
142
- | [jinaai/jina-reranker-v3.5](https://huggingface.co/jinaai/jina-reranker-v3.5) | 597M | **98.38** | 97.33 | 85.39 | 85.45 | **85.65** | 81.94 | 80.94 | 98.87 | - |
143
- | [cross-encoder/ettin-reranker-1b-v1](https://huggingface.co/cross-encoder/ettin-reranker-1b-v1) | 1.03B | 89.01 | 69.14 | 66.24 | 53.06 | 70.04 | 66.59 | 74.61 | 95.90 | 36.51 |
144
- | [nvidia/llama-nemotron-rerank-1b-v2](https://huggingface.co/nvidia/llama-nemotron-rerank-1b-v2) | 1.24B | 94.80 | 98.83 | 85.36 | 75.27 | 81.82 | 80.26 | 84.91 | 98.58 | 67.19 |
145
- | [mixedbread-ai/mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) | 1.54B | 95.31 | 97.78 | 85.63 | 81.30 | 79.39 | **87.71** | 87.72 | 96.81 | 67.87 |
146
- | [BAAI/bge-reranker-v2-gemma](https://huggingface.co/BAAI/bge-reranker-v2-gemma) | 2.51B | 94.07 | 98.57 | 86.14 | 75.72 | 83.62 | 84.30 | 86.98 | 98.58 | 28.81 |
147
- | [tomaarsen/Qwen3-Reranker-4B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-4B-seq-cls) | 4.02B | 97.07 | 99.06 | **87.34** | 87.52 | 85.33 | 83.21 | 86.85 | 98.61 | 81.05 |
148
- | [zeroentropy/zerank-2-reranker](https://huggingface.co/zeroentropy/zerank-2-reranker) | 4.02B | 94.36 | 98.46 | 87.12 | 86.69 | 80.03 | 80.28 | 86.46 | 97.91 | 71.20 |
149
- | [lightonai/LightOn-rerank-PW-4B](https://huggingface.co/lightonai/LightOn-rerank-PW-4B) | 4.54B | 93.21 | 98.82 | 85.67 | 79.38 | 80.72 | 80.91 | 86.93 | 98.03 | 76.09 |
150
- | [tomaarsen/Qwen3-Reranker-8B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-8B-seq-cls) | 7.57B | 95.46 | **99.07** | 86.79 | **90.14** | 84.90 | 84.09 | **88.93** | 98.80 | **82.20** |
151
- | **KURE-Reranker-nano (Ours)** | 150M | 97.41 | 98.36 | 85.83 | 83.76 | 84.03 | 79.99 | 84.80 | **98.89** | 79.78 |
152
-
153
  Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
154
 
155
- ### Top 5: Average across 8 Benchmarks
156
-
157
- MLDR excluded; all models are compared on the same eight benchmarks.
158
-
159
- 1. [tomaarsen/Qwen3-Reranker-8B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-8B-seq-cls) — 91.02
160
- 2. [tomaarsen/Qwen3-Reranker-4B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-4B-seq-cls) — 90.62
161
- 3. [jinaai/jina-reranker-v3.5](https://huggingface.co/jinaai/jina-reranker-v3.5) — 89.24
162
- 4. KURE-Reranker-nano (Ours) — 89.13
163
- 5. [mixedbread-ai/mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) — 88.96
164
-
165
- ### Top 5: Average across 9 Benchmarks
166
-
167
- MLDR included; only models with results on all nine benchmarks are ranked.
168
-
169
- 1. [tomaarsen/Qwen3-Reranker-8B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-8B-seq-cls) — 90.04
170
- 2. [tomaarsen/Qwen3-Reranker-4B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-4B-seq-cls) — 89.56
171
- 3. KURE-Reranker-nano (Ours) — 88.10
172
- 4. [zeroentropy/zerank-2-reranker](https://huggingface.co/zeroentropy/zerank-2-reranker) — 86.95
173
- 5. [lightonai/LightOn-rerank-PW-4B](https://huggingface.co/lightonai/LightOn-rerank-PW-4B) — 86.64
174
-
175
- ### Throughput Measurement
176
-
177
- All values below are query–document pairs per second (PPS), measured on an **NVIDIA RTX A6000 48GB**. **Avg.** is the mean across the eight benchmarks excluding MLDR, consistent with the throughput figure.
178
-
179
- | Model | Params | Avg. | AutoRAG | Belebele | Ko-StrategyQA | LawIRKo | MIRACL | MrTidy | PublicHealthQA | SQuADKorV1 |
180
- |---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|
181
- | [Dongjin-kr/ko-reranker](https://huggingface.co/Dongjin-kr/ko-reranker) | 560M | 515.75 | **257.01** | 588.33 | 529.64 | 374.35 | 581.32 | 874.49 | 580.25 | **340.64** |
182
- | [BAAI/bge-reranker-v2-m3](https://huggingface.co/BAAI/bge-reranker-v2-m3) | 568M | 475.95 | 186.91 | 579.30 | 465.43 | 344.89 | 574.27 | 866.95 | 536.56 | 253.32 |
183
- | [dragonkue/bge-reranker-v2-m3-ko](https://huggingface.co/dragonkue/bge-reranker-v2-m3-ko) | 568M | 478.31 | 187.36 | 585.02 | 463.31 | 346.01 | 578.90 | 867.13 | 539.52 | 259.26 |
184
- | [nlpai-lab/LAMAR-600m](https://huggingface.co/nlpai-lab/LAMAR-600m) | 568M | 476.02 | 186.57 | 582.46 | 461.96 | 343.70 | 572.61 | 864.50 | 532.43 | 263.95 |
185
- | [upskyy/ko-reranker-8k](https://huggingface.co/upskyy/ko-reranker-8k) | 568M | 476.84 | 186.33 | 580.86 | 459.79 | 347.13 | 574.64 | 863.76 | 535.56 | 266.62 |
186
- | [telepix/PIXIE-Spell-Reranker-Preview-0.6B](https://huggingface.co/telepix/PIXIE-Spell-Reranker-Preview-0.6B) | 596M | 137.75 | 46.27 | 174.48 | 128.89 | 90.66 | 174.83 | 263.52 | 147.65 | 75.72 |
187
- | [tomaarsen/Qwen3-Reranker-0.6B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-0.6B-seq-cls) | 596M | 136.86 | 45.94 | 173.73 | 128.03 | 90.13 | 173.63 | 261.11 | 146.76 | 75.51 |
188
- | [cross-encoder/ettin-reranker-1b-v1](https://huggingface.co/cross-encoder/ettin-reranker-1b-v1) | 1.03B | 52.14 | 15.59 | 66.91 | 51.93 | 29.87 | 67.50 | 107.49 | 50.87 | 26.98 |
189
- | [nvidia/llama-nemotron-rerank-1b-v2](https://huggingface.co/nvidia/llama-nemotron-rerank-1b-v2) | 1.24B | 146.06 | 51.00 | 175.62 | 141.54 | 93.02 | 187.95 | 278.57 | 156.71 | 84.03 |
190
- | [mixedbread-ai/mxbai-rerank-large-v2](https://huggingface.co/mixedbread-ai/mxbai-rerank-large-v2) | 1.54B | 70.85 | 29.72 | 87.09 | 71.37 | 52.20 | 87.28 | 116.79 | 76.81 | 45.52 |
191
- | [BAAI/bge-reranker-v2-gemma](https://huggingface.co/BAAI/bge-reranker-v2-gemma) | 2.51B | 74.02 | 24.63 | 89.72 | 73.49 | 45.96 | 95.71 | 142.98 | 77.31 | 42.36 |
192
- | [tomaarsen/Qwen3-Reranker-4B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-4B-seq-cls) | 4.02B | 26.16 | 10.55 | 32.03 | 25.78 | 18.84 | 32.79 | 44.63 | 28.23 | 16.43 |
193
- | [zeroentropy/zerank-2-reranker](https://huggingface.co/zeroentropy/zerank-2-reranker) | 4.02B | 31.97 | 11.34 | 39.87 | 30.26 | 21.21 | 40.80 | 59.88 | 34.10 | 18.29 |
194
- | [lightonai/LightOn-rerank-PW-4B](https://huggingface.co/lightonai/LightOn-rerank-PW-4B) | 4.54B | 16.83 | 7.04 | 20.13 | 16.88 | 12.59 | 20.34 | 27.12 | 19.52 | 11.02 |
195
- | [tomaarsen/Qwen3-Reranker-8B-seq-cls](https://huggingface.co/tomaarsen/Qwen3-Reranker-8B-seq-cls) | 7.57B | 16.39 | 6.77 | 20.18 | 16.04 | 11.82 | 20.54 | 27.65 | 17.71 | 10.37 |
196
- | **KURE-Reranker-nano (Ours)** | 150M | **556.27** | 220.72 | **622.40** | **549.07** | **385.23** | **675.43** | **1044.32** | **621.59** | 331.43 |
197
-
198
- Throughput is measured on a single NVIDIA RTX A6000 48GB per model. The same samples are used across models. Inputs are length-sorted and dynamically padded within each batch.
199
-
200
- Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the [Ettin reranker speed benchmark](https://huggingface.co/blog/ettin-reranker#speed); our aggregation uses total pairs divided by total time rather than the median of three throughput measurements.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
201
 
202
  ## Citation
203
 
 
21
 
22
  # KURE-Reranker-nano
23
 
24
+ **KURE-Reranker-nano** is a lightweight, approximately **150M-parameter** Korean-English bilingual reranker. It jointly encodes a query and a candidate document and predicts a scalar relevance score.
25
 
26
  ## Key Characteristics
27
 
28
+ - **Strong Korean reranking:** 88.08 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 1B parameters and only 0.42 points below KURE-Reranker-base (1.7B).
29
+ - **Long-document reranking:** 79.66 nDCG@10 on MultiLongDocRetrieval (MLDR), the best among the evaluated models with fewer than 1B parameters.
30
+ - **High throughput:** 473.7 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 19.5× Qwen3-Reranker-4B under the measurement conditions described below.
31
  - **Lightweight:** Approximately 150M parameters, based on [`skt/A.X-Encoder-base`](https://huggingface.co/skt/A.X-Encoder-base).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
32
 
33
  ## Model Overview
34
 
35
  | Property | Value |
36
  |---|---|
37
  | Model type | Cross-encoder reranker |
38
+ | Backbone | A.X-Encoder-base |
39
+ | Parameters | 149M |
40
  | Languages | Korean and English |
41
  | Maximum inference input length | 8,192 tokens, including query, document, and special tokens |
42
  | Output | One scalar relevance score per query–document pair |
 
114
 
115
  ## Korean Reranking Evaluation
116
 
117
+ Scores are **nDCG@10**, and throughput is pairs per second (**PPS**).
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
118
  Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
119
 
120
+ <img src="./assets/pps_vs_ndcg9.png" width="700" alt="Reranking performance (mean nDCG@10) vs. throughput (mean PPS)">
121
+
122
+ The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. PPS uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency.
123
+
124
+ ### Results — MTEB-ko-retrieval (9 subsets)
125
+ <!-- **공식 9개 subset 을 모두 평가한 모델**의 9-subset mean NDCG@10, PPS -->
126
+ | Model | Params | Mean NDCG@10 | Mean PPS |
127
+ |---|---|---|---|
128
+ | tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.9004 | 15.1 |
129
+ | tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8956 | 24.3 |
130
+ | **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8849 | 55.1 |
131
+ | **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8808 | 473.7 |
132
+ | zeroentropy/zerank-2-reranker | 4.0B | 0.8695 | 29.4 |
133
+ | lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8664 | 15.0 |
134
+ | mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8661 | 65.2 |
135
+ | BAAI/bge-reranker-v2-m3 | 568M | 0.8586 | 404.1 |
136
+ | tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8585 | 99.6 |
137
+ | nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8522 | 127.3 |
138
+ | nlpai-lab/LAMAR-600m | 568M | 0.8406 | 408.8 |
139
+ | dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8263 | 401.5 |
140
+ | BAAI/bge-reranker-v2-gemma | 2.5B | 0.8186 | 58.7 |
141
+ | upskyy/ko-reranker-8k | 568M | 0.8085 | 404.0 |
142
+ | Dongjin-kr/ko-reranker | 560M | 0.7950 | 482.8 |
143
+ | telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.7806 | 99.6 |
144
+ | cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6901 | 49.7 |
145
+
146
+ PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks.Inputs are length-sorted and dynamically padded within each batch.
147
+
148
+ `jinaai/jina-reranker-v3` and `jinaai/jina-reranker-v3.5` are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.
149
+
150
+ ### Per-dataset NDCG@10
151
+
152
+ | Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
153
+ |---|---|---|---|---|---|---|---|---|---|---|
154
+ | tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.8679 | 0.9546 | 0.8893 | 0.9907 | 0.8490 | 0.8409 | 0.8220 | 0.9880 | 0.9014 |
155
+ | tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8733 | 0.9707 | 0.8685 | 0.9906 | 0.8533 | 0.8321 | 0.8105 | 0.9861 | 0.8752 |
156
+ | **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8548 | 0.9762 | 0.8716 | 0.9880 | 0.8371 | 0.7970 | 0.8163 | 0.9891 | 0.8343 |
157
+ | **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8582 | 0.9741 | 0.8475 | 0.9830 | 0.8400 | 0.8011 | 0.7966 | 0.9895 | 0.8368 |
158
+ | jinaai/jina-reranker-v3.5 | 597M | 0.8539 | 0.9838 | 0.8094 | 0.9733 | 0.8565 | 0.8194 | — | 0.9887 | 0.8545 |
159
+ | jinaai/jina-reranker-v3 | 597M | 0.8553 | 0.9773 | 0.7960 | 0.9695 | 0.8449 | 0.8104 | — | 0.9859 | 0.8546 |
160
+ | zeroentropy/zerank-2-reranker | 4.0B | 0.8712 | 0.9436 | 0.8646 | 0.9846 | 0.8003 | 0.8027 | 0.7120 | 0.9791 | 0.8669 |
161
+ | lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8567 | 0.9321 | 0.8693 | 0.9882 | 0.8072 | 0.8091 | 0.7609 | 0.9803 | 0.7938 |
162
+ | mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8563 | 0.9531 | 0.8772 | 0.9778 | 0.7939 | 0.8771 | 0.6787 | 0.9681 | 0.8130 |
163
+ | BAAI/bge-reranker-v2-m3 | 568M | 0.8487 | 0.9663 | 0.8475 | 0.9853 | 0.8129 | 0.8222 | 0.6690 | 0.9853 | 0.7906 |
164
+ | tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8336 | 0.9308 | 0.8489 | 0.9779 | 0.8507 | 0.7359 | 0.7813 | 0.9808 | 0.7867 |
165
+ | nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8536 | 0.9480 | 0.8491 | 0.9883 | 0.8182 | 0.8026 | 0.6719 | 0.9858 | 0.7527 |
166
+ | nlpai-lab/LAMAR-600m | 568M | 0.8461 | 0.9591 | 0.8225 | 0.9835 | 0.8214 | 0.7975 | 0.5679 | 0.9850 | 0.7822 |
167
+ | dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8232 | 0.9684 | 0.8708 | 0.9769 | 0.7573 | 0.6776 | 0.7061 | 0.9846 | 0.6721 |
168
+ | BAAI/bge-reranker-v2-gemma | 2.5B | 0.8614 | 0.9407 | 0.8698 | 0.9857 | 0.8362 | 0.8429 | 0.2881 | 0.9858 | 0.7572 |
169
+ | upskyy/ko-reranker-8k | 568M | 0.8143 | 0.9230 | 0.8388 | 0.9291 | 0.7249 | 0.6998 | 0.5975 | 0.9718 | 0.7770 |
170
+ | Dongjin-kr/ko-reranker | 560M | 0.8468 | 0.9014 | 0.7675 | 0.9759 | 0.8017 | 0.7772 | 0.3721 | 0.9785 | 0.7343 |
171
+ | telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.8329 | 0.9794 | 0.8534 | 0.9777 | 0.8449 | 0.7650 | 0.1829 | 0.9850 | 0.6042 |
172
+ | cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6624 | 0.8901 | 0.7461 | 0.6914 | 0.7004 | 0.6659 | 0.3651 | 0.9590 | 0.5306 |
173
+
174
+ ### Per-dataset PPS
175
+
176
+ | Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
177
+ |---|---|---|---|---|---|---|---|---|---|---|
178
+ | tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 16.6 | 7.3 | 17.5 | 19.4 | 24.9 | 26.5 | 0.7 | 10.6 | 12.3 |
179
+ | tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 26.1 | 11.8 | 28.4 | 31.5 | 40.0 | 42.4 | 1.1 | 17.2 | 20.0 |
180
+ | **nlpai-lab/KURE-Reranker-base** | 1.7B | 59.2 | 27.3 | 64.6 | 71.1 | 89.1 | 97.1 | 2.6 | 39.1 | 45.7 |
181
+ | **nlpai-lab/KURE-Reranker-nano** | 149M | 475.4 | 233.2 | 569.7 | 630.9 | 765.8 | 855.2 | 16.5 | 330.8 | 386.0 |
182
+ | jinaai/jina-reranker-v3.5 | 597M | 141.0 | 38.6 | 148.6 | 174.6 | 277.3 | 295.7 | — | 68.4 | 86.7 |
183
+ | jinaai/jina-reranker-v3 | 597M | 113.8 | 25.0 | 121.8 | 169.2 | 244.6 | 261.6 | — | 48.7 | 64.0 |
184
+ | zeroentropy/zerank-2-reranker | 4.0B | 31.0 | 12.7 | 33.8 | 37.8 | 51.0 | 56.4 | 1.1 | 18.8 | 22.3 |
185
+ | lightonai/LightOn-rerank-PW-4B | 4.5B | 16.2 | 7.3 | 18.6 | 19.5 | 23.4 | 26.1 | 0.7 | 10.7 | 12.4 |
186
+ | mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 70.0 | 32.9 | 76.8 | 84.1 | 104.2 | 113.6 | 3.2 | 47.0 | 54.7 |
187
+ | BAAI/bge-reranker-v2-m3 | 568M | 420.7 | 195.6 | 471.3 | 545.3 | 653.2 | 724.8 | 9.3 | 271.6 | 345.5 |
188
+ | tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 107.1 | 49.6 | 118.0 | 129.3 | 159.5 | 173.9 | 4.2 | 71.8 | 82.9 |
189
+ | nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 133.7 | 54.4 | 146.9 | 163.0 | 220.0 | 243.1 | 3.9 | 84.9 | 95.9 |
190
+ | nlpai-lab/LAMAR-600m | 568M | 418.0 | 196.5 | 480.5 | 557.2 | 656.7 | 742.5 | 9.2 | 275.0 | 343.8 |
191
+ | dragonkue/bge-reranker-v2-m3-ko | 568M | 416.9 | 195.1 | 467.2 | 540.7 | 653.4 | 716.7 | 9.2 | 271.9 | 342.7 |
192
+ | BAAI/bge-reranker-v2-gemma | 2.5B | 64.7 | 26.1 | 67.5 | 75.8 | 99.8 | 106.8 | 2.5 | 40.0 | 44.7 |
193
+ | upskyy/ko-reranker-8k | 568M | 411.8 | 195.6 | 474.3 | 551.4 | 649.3 | 726.0 | 9.2 | 275.1 | 342.8 |
194
+ | Dongjin-kr/ko-reranker | 560M | 520.8 | 258.0 | 510.1 | 554.8 | 747.6 | 765.2 | 272.7 | 334.4 | 381.2 |
195
+ | telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 104.6 | 49.4 | 117.8 | 129.8 | 160.5 | 175.6 | 4.3 | 72.1 | 82.6 |
196
+ | cross-encoder/ettin-reranker-1b-v1 | 1.0B | 52.4 | 19.9 | 51.4 | 65.6 | 92.7 | 97.2 | 3.8 | 31.2 | 33.4 |
197
+
198
+ Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the [Ettin reranker speed benchmark](https://huggingface.co/blog/ettin-reranker#speed).
199
 
200
  ## Citation
201
 
assets/pps_vs_ndcg9.png ADDED

Git LFS Details

  • SHA256: 99ba835554b7c39a471a033633472e0138339e45d9bfce2ef67cb3b08b4f5d51
  • Pointer size: 131 Bytes
  • Size of remote file: 110 kB