---
pipeline_tag: text-ranking
library_name: sentence-transformers
license: apache-2.0
language:
- ko
- en
base_model:
- Qwen/Qwen3-1.7B
tags:
- sentence-transformers
- cross-encoder
- reranker
- korean
- english
---
# KURE-Reranker-base
**KURE-Reranker-base** is a Korean-English bilingual reranker. It reads a query and a candidate document together and predicts a scalar relevance score.
## Key Characteristics
- **Strong Korean reranking:** 88.49 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 4B parameters and within 1.1 points of Qwen3-Reranker-4B.
- **Long-document reranking:** 81.63 nDCG@10 on MultiLongDocRetrieval (MLDR), ranking second only to Qwen3-Reranker-8B (82.20) among the evaluated models.
- **Efficient:** 55.1 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 2.3× Qwen3-Reranker-4B and 3.6× Qwen3-Reranker-8B under the measurement conditions described below.
## Model Overview
| Property | Value |
|---|---|
| Model type | Cross-encoder reranker (causal LM with yes/no logit scoring) |
| Backbone | Qwen3-1.7B |
| Parameters | 1.7B |
| Languages | Korean and English |
| Maximum inference input length | 8,192 tokens, including instruction, query, document, and prompt template |
| Output | One scalar relevance score per query–document pair: logit("yes") − logit("no") |
| License | Apache-2.0 |
## Usage
Install the required libraries:
```bash
pip install "sentence-transformers>=5.6.1" torch
```
The prompt template ships with the model. The default instruction is *"Given a web search query, retrieve relevant passages that answer the query"*, the same instruction used in training.
Sentence Transformers
```python
import torch
from sentence_transformers import CrossEncoder
model = CrossEncoder("nlpai-lab/KURE-Reranker-base", max_length=8192, model_kwargs={"dtype": torch.bfloat16})
query = "대한민국의 수도는 어디인가요?"
documents = [
"대한민국의 수도는 서울입니다.",
"The capital of South Korea is Seoul.",
"바나나는 열대 지역에서 재배되는 과일입니다.",
]
results = model.rank(query, documents, batch_size=8, return_documents=True)
for result in results:
print(f"{result['score']:.4f}\t{result['text']}")
```
For pair scoring without sorting:
```python
pairs = [(query, document) for document in documents]
scores = model.predict(pairs, batch_size=8)
```
Transformers
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "nlpai-lab/KURE-Reranker-base"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).to(device)
model.eval()
max_length = 8192
instruction = "Given a web search query, retrieve relevant passages that answer the query"
yes_id = tokenizer.convert_tokens_to_ids("yes")
no_id = tokenizer.convert_tokens_to_ids("no")
def format_pair(query, document):
messages = [
{"role": "system", "content": instruction},
{"role": "query", "content": query},
{"role": "document", "content": document},
]
return tokenizer.apply_chat_template(messages, tokenize=False)
def truncate_document(query, document):
# Truncate the document, not the formatted text, so the closing assistant turn stays intact.
budget = max_length - len(tokenizer(format_pair(query, ""), add_special_tokens=False).input_ids)
ids = tokenizer(document, add_special_tokens=False).input_ids
return document if len(ids) <= budget else tokenizer.decode(ids[:budget])
query = "대한민국의 수도는 어디인가요?"
documents = [
"대한민국의 수도는 서울입니다.",
"The capital of South Korea is Seoul.",
"바나나는 열대 지역에서 재배되는 과일입니다.",
]
texts = [format_pair(query, truncate_document(query, document)) for document in documents]
inputs = tokenizer(texts, padding=True, add_special_tokens=False, return_tensors="pt").to(device)
with torch.inference_mode():
logits = model(**inputs, logits_to_keep=1).logits[:, -1]
scores = (logits[:, yes_id] - logits[:, no_id]).float().cpu()
for index in torch.argsort(scores, descending=True).tolist():
print(f"{scores[index]:.4f}\t{documents[index]}")
```
## Korean Reranking Evaluation
Scores are **nDCG@10**, and throughput is pairs per second (**PPS**).
Evaluation was conducted using the [reranker-simple-benchmark](https://github.com/instructkr/reranker-simple-benchmark) implementation as a reference.
The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. Throughput uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency.
### Results — MTEB-ko-retrieval (9 subsets)
| Model | Params | Mean NDCG@10 | Mean PPS |
|---|---|---|---|
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.9004 | 15.1 |
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8956 | 24.3 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8849 | 55.1 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8808 | 473.7 |
| zeroentropy/zerank-2-reranker | 4.0B | 0.8695 | 29.4 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8664 | 15.0 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8661 | 65.2 |
| BAAI/bge-reranker-v2-m3 | 568M | 0.8586 | 404.1 |
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8585 | 99.6 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8522 | 127.3 |
| nlpai-lab/LAMAR-600m | 568M | 0.8406 | 408.8 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8263 | 401.5 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8186 | 58.7 |
| upskyy/ko-reranker-8k | 568M | 0.8085 | 404.0 |
| Dongjin-kr/ko-reranker | 560M | 0.7950 | 482.8 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.7806 | 99.6 |
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6901 | 49.7 |
PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks. Inputs are length-sorted and dynamically padded within each batch.
`jinaai/jina-reranker-v3` and `jinaai/jina-reranker-v3.5` are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.
### Per-dataset NDCG@10
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|---|---|---|---|---|---|---|---|---|---|---|
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 0.8679 | 0.9546 | 0.8893 | 0.9907 | 0.8490 | 0.8409 | 0.8220 | 0.9880 | 0.9014 |
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 0.8733 | 0.9707 | 0.8685 | 0.9906 | 0.8533 | 0.8321 | 0.8105 | 0.9861 | 0.8752 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 0.8548 | 0.9762 | 0.8716 | 0.9880 | 0.8371 | 0.7970 | 0.8163 | 0.9891 | 0.8343 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 0.8582 | 0.9741 | 0.8475 | 0.9830 | 0.8400 | 0.8011 | 0.7966 | 0.9895 | 0.8368 |
| jinaai/jina-reranker-v3.5 | 597M | 0.8539 | 0.9838 | 0.8094 | 0.9733 | 0.8565 | 0.8194 | — | 0.9887 | 0.8545 |
| jinaai/jina-reranker-v3 | 597M | 0.8553 | 0.9773 | 0.7960 | 0.9695 | 0.8449 | 0.8104 | — | 0.9859 | 0.8546 |
| zeroentropy/zerank-2-reranker | 4.0B | 0.8712 | 0.9436 | 0.8646 | 0.9846 | 0.8003 | 0.8027 | 0.7120 | 0.9791 | 0.8669 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 0.8567 | 0.9321 | 0.8693 | 0.9882 | 0.8072 | 0.8091 | 0.7609 | 0.9803 | 0.7938 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 0.8563 | 0.9531 | 0.8772 | 0.9778 | 0.7939 | 0.8771 | 0.6787 | 0.9681 | 0.8130 |
| BAAI/bge-reranker-v2-m3 | 568M | 0.8487 | 0.9663 | 0.8475 | 0.9853 | 0.8129 | 0.8222 | 0.6690 | 0.9853 | 0.7906 |
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 0.8336 | 0.9308 | 0.8489 | 0.9779 | 0.8507 | 0.7359 | 0.7813 | 0.9808 | 0.7867 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 0.8536 | 0.9480 | 0.8491 | 0.9883 | 0.8182 | 0.8026 | 0.6719 | 0.9858 | 0.7527 |
| nlpai-lab/LAMAR-600m | 568M | 0.8461 | 0.9591 | 0.8225 | 0.9835 | 0.8214 | 0.7975 | 0.5679 | 0.9850 | 0.7822 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 0.8232 | 0.9684 | 0.8708 | 0.9769 | 0.7573 | 0.6776 | 0.7061 | 0.9846 | 0.6721 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 0.8614 | 0.9407 | 0.8698 | 0.9857 | 0.8362 | 0.8429 | 0.2881 | 0.9858 | 0.7572 |
| upskyy/ko-reranker-8k | 568M | 0.8143 | 0.9230 | 0.8388 | 0.9291 | 0.7249 | 0.6998 | 0.5975 | 0.9718 | 0.7770 |
| Dongjin-kr/ko-reranker | 560M | 0.8468 | 0.9014 | 0.7675 | 0.9759 | 0.8017 | 0.7772 | 0.3721 | 0.9785 | 0.7343 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 0.8329 | 0.9794 | 0.8534 | 0.9777 | 0.8449 | 0.7650 | 0.1829 | 0.9850 | 0.6042 |
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 0.6624 | 0.8901 | 0.7461 | 0.6914 | 0.7004 | 0.6659 | 0.3651 | 0.9590 | 0.5306 |
### Per-dataset PPS
| Model | Params | Ko-StrategyQA | AutoRAGRetrieval | PublicHealthQA | BelebeleRetrieval | MIRACLRetrieval | MrTidyRetrieval | MultiLongDocRetrieval | SQuADKorV1Retrieval | LawIRKo |
|---|---|---|---|---|---|---|---|---|---|---|
| tomaarsen/Qwen3-Reranker-8B-seq-cls | 7.6B | 16.6 | 7.3 | 17.5 | 19.4 | 24.9 | 26.5 | 0.7 | 10.6 | 12.3 |
| tomaarsen/Qwen3-Reranker-4B-seq-cls | 4.0B | 26.1 | 11.8 | 28.4 | 31.5 | 40.0 | 42.4 | 1.1 | 17.2 | 20.0 |
| **nlpai-lab/KURE-Reranker-base** | 1.7B | 59.2 | 27.3 | 64.6 | 71.1 | 89.1 | 97.1 | 2.6 | 39.1 | 45.7 |
| **nlpai-lab/KURE-Reranker-nano** | 149M | 475.4 | 233.2 | 569.7 | 630.9 | 765.8 | 855.2 | 16.5 | 330.8 | 386.0 |
| jinaai/jina-reranker-v3.5 | 597M | 141.0 | 38.6 | 148.6 | 174.6 | 277.3 | 295.7 | — | 68.4 | 86.7 |
| jinaai/jina-reranker-v3 | 597M | 113.8 | 25.0 | 121.8 | 169.2 | 244.6 | 261.6 | — | 48.7 | 64.0 |
| zeroentropy/zerank-2-reranker | 4.0B | 31.0 | 12.7 | 33.8 | 37.8 | 51.0 | 56.4 | 1.1 | 18.8 | 22.3 |
| lightonai/LightOn-rerank-PW-4B | 4.5B | 16.2 | 7.3 | 18.6 | 19.5 | 23.4 | 26.1 | 0.7 | 10.7 | 12.4 |
| mixedbread-ai/mxbai-rerank-large-v2 | 1.5B | 70.0 | 32.9 | 76.8 | 84.1 | 104.2 | 113.6 | 3.2 | 47.0 | 54.7 |
| BAAI/bge-reranker-v2-m3 | 568M | 420.7 | 195.6 | 471.3 | 545.3 | 653.2 | 724.8 | 9.3 | 271.6 | 345.5 |
| tomaarsen/Qwen3-Reranker-0.6B-seq-cls | 596M | 107.1 | 49.6 | 118.0 | 129.3 | 159.5 | 173.9 | 4.2 | 71.8 | 82.9 |
| nvidia/llama-nemotron-rerank-1b-v2 | 1.2B | 133.7 | 54.4 | 146.9 | 163.0 | 220.0 | 243.1 | 3.9 | 84.9 | 95.9 |
| nlpai-lab/LAMAR-600m | 568M | 418.0 | 196.5 | 480.5 | 557.2 | 656.7 | 742.5 | 9.2 | 275.0 | 343.8 |
| dragonkue/bge-reranker-v2-m3-ko | 568M | 416.9 | 195.1 | 467.2 | 540.7 | 653.4 | 716.7 | 9.2 | 271.9 | 342.7 |
| BAAI/bge-reranker-v2-gemma | 2.5B | 64.7 | 26.1 | 67.5 | 75.8 | 99.8 | 106.8 | 2.5 | 40.0 | 44.7 |
| upskyy/ko-reranker-8k | 568M | 411.8 | 195.6 | 474.3 | 551.4 | 649.3 | 726.0 | 9.2 | 275.1 | 342.8 |
| Dongjin-kr/ko-reranker | 560M | 520.8 | 258.0 | 510.1 | 554.8 | 747.6 | 765.2 | 272.7 | 334.4 | 381.2 |
| telepix/PIXIE-Spell-Reranker-Preview-0.6B | 596M | 104.6 | 49.4 | 117.8 | 129.8 | 160.5 | 175.6 | 4.3 | 72.1 | 82.6 |
| cross-encoder/ettin-reranker-1b-v1 | 1.0B | 52.4 | 19.9 | 51.4 | 65.6 | 92.7 | 97.2 | 3.8 | 31.2 | 33.4 |
Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the [Ettin reranker speed benchmark](https://huggingface.co/blog/ettin-reranker#speed).
## Citation
```bibtex
@misc{kure-reranker-base,
title = {KURE-Reranker-base: A Korean–English Bilingual Reranking Model},
author = {Jang, Youngjoon and Hong, Seongtae and Son, Junyoung and Lee, Taemin and Lim, Heuiseok},
year = {2026},
url = {https://huggingface.co/nlpai-lab/KURE-Reranker-base},
}
```