Text Ranking
sentence-transformers
Safetensors
Korean
English
qwen3
cross-encoder
reranker
korean
english
Instructions to use nlpai-lab/KURE-Reranker-base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use nlpai-lab/KURE-Reranker-base with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("nlpai-lab/KURE-Reranker-base") query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
File size: 12,929 Bytes
65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 53e42a4 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 a64a617 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 65bf9cb 5cbe128 53e42a4 5cbe128 53e42a4 a64a617 5cbe128 53e42a4 5cbe128 53e42a4 5cbe128 53e42a4 a64a617 5cbe128 53e42a4 5cbe128 53e42a4 a64a617 5cbe128 53e42a4 5cbe128 a64a617 65bf9cb 5cbe128 65bf9cb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 | ---
pipeline_tag: text-ranking
library_name: sentence-transformers
license: apache-2.0
language:
- ko
- en
base_model:
- Qwen/Qwen3-1.7B
tags:
- sentence-transformers
- cross-encoder
- reranker
- korean
- english
---
<a href="https://github.com/nlpai-lab/KURE">
<img src="./assets/kure_logo.png" width="50%" alt="KURE logo">
</a>
# KURE-Reranker-base
**KURE-Reranker-base** is a Korean-English bilingual reranker. It reads a query and a candidate document together and predicts a scalar relevance score.
## Key Characteristics
- **Strong Korean reranking:** 88.49 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 4B parameters and within 1.1 points of Qwen3-Reranker-4B.
- **Long-document reranking:** 81.63 nDCG@10 on MultiLongDocRetrieval (MLDR), ranking second only to Qwen3-Reranker-8B (82.55) among the evaluated models.
- **Efficient:** 55.1 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 2.3ร Qwen3-Reranker-4B and 3.6ร Qwen3-Reranker-8B under the measurement conditions described below.
## Model Overview
| Property | Value |
|---|---|
| Model type | Cross-encoder reranker (causal LM with yes/no logit scoring) |
| Backbone | Qwen3-1.7B |
| Parameters | 1.7B |
| Languages | Korean and English |
| Maximum inference input length | 8,192 tokens, including instruction, query, document, and prompt template |
| Output | One scalar relevance score per queryโdocument pair: logit("yes") โ logit("no") |
| License | Apache-2.0 |
## Usage
Install the required libraries:
```bash
pip install "sentence-transformers>=5.6.1" torch
```
The prompt template ships with the model. The default instruction is *"Given a web search query, retrieve relevant passages that answer the query"*, the same instruction used in training.
<details>
<summary>Sentence Transformers</summary>
```python
import torch
from sentence_transformers import CrossEncoder
model = CrossEncoder("nlpai-lab/KURE-Reranker-base", max_length=8192, model_kwargs={"dtype": torch.bfloat16})
query = "๋ํ๋ฏผ๊ตญ์ ์๋๋ ์ด๋์ธ๊ฐ์?"
documents = [
"๋ํ๋ฏผ๊ตญ์ ์๋๋ ์์ธ์
๋๋ค.",
"The capital of South Korea is Seoul.",
"๋ฐ๋๋๋ ์ด๋ ์ง์ญ์์ ์ฌ๋ฐฐ๋๋ ๊ณผ์ผ์
๋๋ค.",
]
results = model.rank(query, documents, batch_size=8, return_documents=True)
for result in results:
print(f"{result['score']:.4f}\t{result['text']}")
```
For pair scoring without sorting:
```python
pairs = [(query, document) for document in documents]
scores = model.predict(pairs, batch_size=8)
```
</details>
<details>
<summary>Transformers</summary>
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "nlpai-lab/KURE-Reranker-base"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).to(device)
model.eval()
max_length = 8192
instruction = "Given a web search query, retrieve relevant passages that answer the query"
yes_id = tokenizer.convert_tokens_to_ids("yes")
no_id = tokenizer.convert_tokens_to_ids("no")
def format_pair(query, document):
messages = [
{"role": "system", "content": instruction},
{"role": "query", "content": query},
{"role": "document", "content": document},
]
return tokenizer.apply_chat_template(messages, tokenize=False)
def truncate_document(query, document):
# Truncate the document, not the formatted text, so the closing assistant turn stays intact.
budget = max_length - len(tokenizer(format_pair(query, ""), add_special_tokens=False).input_ids)
ids = tokenizer(document, add_special_tokens=False).input_ids
return document if len(ids) <= budget else tokenizer.decode(ids[:budget])
query = "๋ํ๋ฏผ๊ตญ์ ์๋๋ ์ด๋์ธ๊ฐ์?"
documents = [
"๋ํ๋ฏผ๊ตญ์ ์๋๋ ์์ธ์
๋๋ค.",
"The capital of South Korea is Seoul.",
"๋ฐ๋๋๋ ์ด๋ ์ง์ญ์์ ์ฌ๋ฐฐ๋๋ ๊ณผ์ผ์
๋๋ค.",
]
texts = [format_pair(query, truncate_document(query, document)) for document in documents]
inputs = tokenizer(texts, padding=True, add_special_tokens=False, return_tensors="pt").to(device)
with torch.inference_mode():
logits = model(**inputs, logits_to_keep=1).logits[:, -1]
scores = (logits[:, yes_id] - logits[:, no_id]).float().cpu()
for index in torch.argsort(scores, descending=True).tolist():
print(f"{scores[index]:.4f}\t{documents[index]}")
```
</details>
## Korean Reranking Evaluation
Scores are **nDCG@10**, and throughput is pairs per second (**PPS**).
Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
<img src="./assets/pps_vs_ndcg9.png" width="700" alt="Reranking performance (mean nDCG@10) vs. throughput (mean PPS)">
The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. Throughput uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency. The x-axis is a non-uniform log scale that stretches 20โ26 and compresses 26โ50 pairs/s so that nearby points stay readable.
### Results โ MTEB-ko-retrieval (9 subsets)
<!-- **๊ณต์ 9๊ฐ subset ์ ๋ชจ๋ ํ๊ฐํ ๋ชจ๋ธ**์ 9-subset mean NDCG@10, PPS -->
| Model | Params | Mean NDCG@10 | Mean PPS |
|---|---|---|---|
| Qwen/Qwen3-Reranker-8B | 8.2B | 0.9030 | 15.2 |
| KaLM-Embedding/KaLM-Reranker-V1-Large-R2 | 7.5B | 0.8960 | 25.3 |
| Qwen/Qwen3-Reranker-4B | 4.0B | 0.8957 | 24.3 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8849 | 55.1 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8808 | 473.7 |
| zeroentropy/zerank-2-reranker | 4.0B | 0.8695 | 29.4 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8664 | 15.0 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8661 | 65.2 |
| BAAI/bge-reranker-v2-m3 | 568M | 0.8586 | 404.1 |
| Qwen/Qwen3-Reranker-0.6B | 596M | 0.8577 | 100.2 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8522 | 127.3 |
| nlpai-lab/LAMAR-600m | 568M | 0.8406 | 408.8 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8263 | 401.5 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8186 | 58.7 |
| upskyy/ko-reranker-8k | 568M | 0.8085 | 404.0 |
| Dongjin-kr/ko-reranker | 560M | 0.7950 | 482.8 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.7806 | 99.6 |
PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks. For each benchmark, every model is timed on the same ~550 sampled queryโdocument pairs (whole queries, fixed seed). Inputs are length-sorted and dynamically padded within each batch. Models run in bf16 with flash_attention_2, except `KaLM-Embedding/KaLM-Reranker-V1-Large-R2`, which uses sdpa because transformers does not support flash_attention_2 for T5Gemma2.
`jinaai/jina-reranker-v3` and `jinaai/jina-reranker-v3.5` are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.
### Per-dataset NDCG@10
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-Reranker-8B | 8.2B | 0.8720 | 0.9653 | 0.8899 | 0.9908 | 0.8513 | 0.8416 | 0.8255 | 0.9897 | 0.9013 |
| KaLM-Embedding/KaLM-Reranker-V1-Large-R2 | 7.5B | 0.8796 | 0.9415 | 0.8956 | 0.9936 | 0.8382 | 0.8500 | 0.7970 | 0.9863 | 0.8824 |
| Qwen/Qwen3-Reranker-4B | 4.0B | 0.8733 | 0.9630 | 0.8702 | 0.9904 | 0.8559 | 0.8298 | 0.8161 | 0.9861 | 0.8766 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8548 | 0.9762 | 0.8716 | 0.9880 | 0.8371 | 0.7970 | 0.8163 | 0.9891 | 0.8343 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8582 | 0.9741 | 0.8475 | 0.9830 | 0.8400 | 0.8011 | 0.7966 | 0.9895 | 0.8368 |
| jinaai/jina-reranker-v3.5 | 597M | 0.8539 | 0.9838 | 0.8094 | 0.9733 | 0.8565 | 0.8194 | โ | 0.9887 | 0.8545 |
| jinaai/jina-reranker-v3 | 597M | 0.8553 | 0.9773 | 0.7960 | 0.9695 | 0.8449 | 0.8104 | โ | 0.9859 | 0.8546 |
| zeroentropy/zerank-2-reranker | 4.0B | 0.8712 | 0.9436 | 0.8646 | 0.9846 | 0.8003 | 0.8027 | 0.7120 | 0.9791 | 0.8669 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8567 | 0.9321 | 0.8693 | 0.9882 | 0.8072 | 0.8091 | 0.7609 | 0.9803 | 0.7938 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8563 | 0.9531 | 0.8772 | 0.9778 | 0.7939 | 0.8771 | 0.6787 | 0.9681 | 0.8130 |
| BAAI/bge-reranker-v2-m3 | 568M | 0.8487 | 0.9663 | 0.8475 | 0.9853 | 0.8129 | 0.8222 | 0.6690 | 0.9853 | 0.7906 |
| Qwen/Qwen3-Reranker-0.6B | 596M | 0.8336 | 0.9321 | 0.8477 | 0.9779 | 0.8470 | 0.7325 | 0.7819 | 0.9804 | 0.7865 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8536 | 0.9480 | 0.8491 | 0.9883 | 0.8182 | 0.8026 | 0.6719 | 0.9858 | 0.7527 |
| nlpai-lab/LAMAR-600m | 568M | 0.8461 | 0.9591 | 0.8225 | 0.9835 | 0.8214 | 0.7975 | 0.5679 | 0.9850 | 0.7822 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8232 | 0.9684 | 0.8708 | 0.9769 | 0.7573 | 0.6776 | 0.7061 | 0.9846 | 0.6721 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8614 | 0.9407 | 0.8698 | 0.9857 | 0.8362 | 0.8429 | 0.2881 | 0.9858 | 0.7572 |
| upskyy/ko-reranker-8k | 568M | 0.8143 | 0.9230 | 0.8388 | 0.9291 | 0.7249 | 0.6998 | 0.5975 | 0.9718 | 0.7770 |
| Dongjin-kr/ko-reranker | 560M | 0.8468 | 0.9014 | 0.7675 | 0.9759 | 0.8017 | 0.7772 | 0.3721 | 0.9785 | 0.7343 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.8329 | 0.9794 | 0.8534 | 0.9777 | 0.8449 | 0.7650 | 0.1829 | 0.9850 | 0.6042 |
### Per-dataset PPS
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|---|---|---|---|---|---|---|---|---|---|---|
| Qwen/Qwen3-Reranker-8B | 8.2B | 16.5 | 7.4 | 17.6 | 19.4 | 25.0 | 26.6 | 0.7 | 10.8 | 12.4 |
| KaLM-Embedding/KaLM-Reranker-V1-Large-R2 | 7.5B | 28.6 | 12.4 | 28.5 | 33.7 | 39.6 | 43.8 | 0.8 | 18.7 | 21.4 |
| Qwen/Qwen3-Reranker-4B | 4.0B | 26.2 | 11.8 | 28.2 | 31.1 | 40.1 | 42.8 | 1.1 | 17.1 | 19.9 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 59.2 | 27.3 | 64.6 | 71.1 | 89.1 | 97.1 | 2.6 | 39.1 | 45.7 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 475.4 | 233.2 | 569.7 | 630.9 | 765.8 | 855.2 | 16.5 | 330.8 | 386.0 |
| jinaai/jina-reranker-v3.5 | 597M | 141.0 | 38.6 | 148.6 | 174.6 | 277.3 | 295.7 | โ | 68.4 | 86.7 |
| jinaai/jina-reranker-v3 | 597M | 113.8 | 25.0 | 121.8 | 169.2 | 244.6 | 261.6 | โ | 48.7 | 64.0 |
| zeroentropy/zerank-2-reranker | 4.0B | 31.0 | 12.7 | 33.8 | 37.8 | 51.0 | 56.4 | 1.1 | 18.8 | 22.3 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 16.2 | 7.3 | 18.6 | 19.5 | 23.4 | 26.1 | 0.7 | 10.7 | 12.4 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 70.0 | 32.9 | 76.8 | 84.1 | 104.2 | 113.6 | 3.2 | 47.0 | 54.7 |
| BAAI/bge-reranker-v2-m3 | 568M | 420.7 | 195.6 | 471.3 | 545.3 | 653.2 | 724.8 | 9.3 | 271.6 | 345.5 |
| Qwen/Qwen3-Reranker-0.6B | 596M | 105.9 | 49.4 | 117.9 | 130.3 | 162.6 | 177.4 | 4.3 | 70.8 | 83.6 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 133.7 | 54.4 | 146.9 | 163.0 | 220.0 | 243.1 | 3.9 | 84.9 | 95.9 |
| nlpai-lab/LAMAR-600m | 568M | 418.0 | 196.5 | 480.5 | 557.2 | 656.7 | 742.5 | 9.2 | 275.0 | 343.8 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 416.9 | 195.1 | 467.2 | 540.7 | 653.4 | 716.7 | 9.2 | 271.9 | 342.7 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 64.7 | 26.1 | 67.5 | 75.8 | 99.8 | 106.8 | 2.5 | 40.0 | 44.7 |
| upskyy/ko-reranker-8k | 568M | 411.8 | 195.6 | 474.3 | 551.4 | 649.3 | 726.0 | 9.2 | 275.1 | 342.8 |
| Dongjin-kr/ko-reranker | 560M | 520.8 | 258.0 | 510.1 | 554.8 | 747.6 | 765.2 | 272.7 | 334.4 | 381.2 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 104.6 | 49.4 | 117.8 | 129.8 | 160.5 | 175.6 | 4.3 | 72.1 | 82.6 |
Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the [Ettin reranker speed benchmark](https://huggingface.co/blog/ettin-reranker#speed).
## Citation
```bibtex
@misc{kure-reranker-base,
title = {KURE-Reranker-base: A KoreanโEnglish Bilingual Reranking Model},
author = {Jang, Youngjoon and Hong, Seongtae and Son, Junyoung and Lee, Taemin and Lim, Heuiseok},
year = {2026},
url = {https://huggingface.co/nlpai-lab/KURE-Reranker-base},
}
```
|