KURE logo

KURE-Reranker-base

KURE-Reranker-base is a Korean-English bilingual reranker. It reads a query and a candidate document together and predicts a scalar relevance score.

Key Characteristics

  • Strong Korean reranking: 88.49 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 4B parameters and within 1.1 points of Qwen3-Reranker-4B.
  • Long-document reranking: 81.63 nDCG@10 on MultiLongDocRetrieval (MLDR), ranking second only to Qwen3-Reranker-8B (82.20) among the evaluated models.
  • Efficient: 55.1 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 2.3× Qwen3-Reranker-4B and 3.6× Qwen3-Reranker-8B under the measurement conditions described below.

Model Overview

Property Value
Model type Cross-encoder reranker (causal LM with yes/no logit scoring)
Backbone Qwen3-1.7B
Parameters 1.7B
Languages Korean and English
Maximum inference input length 8,192 tokens, including instruction, query, document, and prompt template
Output One scalar relevance score per query–document pair: logit("yes") − logit("no")
License Apache-2.0

Usage

Install the required libraries:

pip install "sentence-transformers>=5.6.1" torch

The prompt template ships with the model. The default instruction is "Given a web search query, retrieve relevant passages that answer the query", the same instruction used in training.

Sentence Transformers
import torch
from sentence_transformers import CrossEncoder

model = CrossEncoder("nlpai-lab/KURE-Reranker-base", max_length=8192, model_kwargs={"dtype": torch.bfloat16})

query = "대한민국의 수도는 어디인가요?"
documents = [
    "대한민국의 수도는 서울입니다.",
    "The capital of South Korea is Seoul.",
    "바나나는 열대 지역에서 재배되는 과일입니다.",
]

results = model.rank(query, documents, batch_size=8, return_documents=True)
for result in results:
    print(f"{result['score']:.4f}\t{result['text']}")

For pair scoring without sorting:

pairs = [(query, document) for document in documents]
scores = model.predict(pairs, batch_size=8)
Transformers
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "nlpai-lab/KURE-Reranker-base"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_id, padding_side="left")
model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16).to(device)
model.eval()

max_length = 8192
instruction = "Given a web search query, retrieve relevant passages that answer the query"
yes_id = tokenizer.convert_tokens_to_ids("yes")
no_id = tokenizer.convert_tokens_to_ids("no")

def format_pair(query, document):
    messages = [
        {"role": "system", "content": instruction},
        {"role": "query", "content": query},
        {"role": "document", "content": document},
    ]
    return tokenizer.apply_chat_template(messages, tokenize=False)

def truncate_document(query, document):
    # Truncate the document, not the formatted text, so the closing assistant turn stays intact.
    budget = max_length - len(tokenizer(format_pair(query, ""), add_special_tokens=False).input_ids)
    ids = tokenizer(document, add_special_tokens=False).input_ids
    return document if len(ids) <= budget else tokenizer.decode(ids[:budget])

query = "대한민국의 수도는 어디인가요?"
documents = [
    "대한민국의 수도는 서울입니다.",
    "The capital of South Korea is Seoul.",
    "바나나는 열대 지역에서 재배되는 과일입니다.",
]
texts = [format_pair(query, truncate_document(query, document)) for document in documents]
inputs = tokenizer(texts, padding=True, add_special_tokens=False, return_tensors="pt").to(device)

with torch.inference_mode():
    logits = model(**inputs, logits_to_keep=1).logits[:, -1]
    scores = (logits[:, yes_id] - logits[:, no_id]).float().cpu()

for index in torch.argsort(scores, descending=True).tolist():
    print(f"{scores[index]:.4f}\t{documents[index]}")

Korean Reranking Evaluation

Scores are nDCG@10, and throughput is pairs per second (PPS). Evaluation was conducted using the reranker-simple-benchmark implementation as a reference.

Reranking performance (mean nDCG@10) vs. throughput (mean PPS)

The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. Throughput uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency.

Results — MTEB-ko-retrieval (9 subsets)

Model Params Mean NDCG@10 Mean PPS
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 0.9004 15.1
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 0.8956 24.3
nlpai-lab/KURE-Reranker-base 1.7B 0.8849 55.1
nlpai-lab/KURE-Reranker-nano 149M 0.8808 473.7
zeroentropy/zerank-2-reranker 4.0B 0.8695 29.4
lightonai/LightOn-rerank-PW-4B 4.5B 0.8664 15.0
mixedbread-ai/mxbai-rerank-large-v2 1.5B 0.8661 65.2
BAAI/bge-reranker-v2-m3 568M 0.8586 404.1
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 0.8585 99.6
nvidia/llama-nemotron-rerank-1b-v2 1.2B 0.8522 127.3
nlpai-lab/LAMAR-600m 568M 0.8406 408.8
dragonkue/bge-reranker-v2-m3-ko 568M 0.8263 401.5
BAAI/bge-reranker-v2-gemma 2.5B 0.8186 58.7
upskyy/ko-reranker-8k 568M 0.8085 404.0
Dongjin-kr/ko-reranker 560M 0.7950 482.8
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 0.7806 99.6
cross-encoder/ettin-reranker-1b-v1 1.0B 0.6901 49.7

PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks. Inputs are length-sorted and dynamically padded within each batch.

jinaai/jina-reranker-v3 and jinaai/jina-reranker-v3.5 are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.

Per-dataset NDCG@10

Model Params Ko-StrategyQA AutoRAGRetrieval PublicHealthQA BelebeleRetrieval MIRACLRetrieval MrTidyRetrieval MultiLongDocRetrieval SQuADKorV1Retrieval LawIRKo
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 0.8679 0.9546 0.8893 0.9907 0.8490 0.8409 0.8220 0.9880 0.9014
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 0.8733 0.9707 0.8685 0.9906 0.8533 0.8321 0.8105 0.9861 0.8752
nlpai-lab/KURE-Reranker-base 1.7B 0.8548 0.9762 0.8716 0.9880 0.8371 0.7970 0.8163 0.9891 0.8343
nlpai-lab/KURE-Reranker-nano 149M 0.8582 0.9741 0.8475 0.9830 0.8400 0.8011 0.7966 0.9895 0.8368
jinaai/jina-reranker-v3.5 597M 0.8539 0.9838 0.8094 0.9733 0.8565 0.8194 — 0.9887 0.8545
jinaai/jina-reranker-v3 597M 0.8553 0.9773 0.7960 0.9695 0.8449 0.8104 — 0.9859 0.8546
zeroentropy/zerank-2-reranker 4.0B 0.8712 0.9436 0.8646 0.9846 0.8003 0.8027 0.7120 0.9791 0.8669
lightonai/LightOn-rerank-PW-4B 4.5B 0.8567 0.9321 0.8693 0.9882 0.8072 0.8091 0.7609 0.9803 0.7938
mixedbread-ai/mxbai-rerank-large-v2 1.5B 0.8563 0.9531 0.8772 0.9778 0.7939 0.8771 0.6787 0.9681 0.8130
BAAI/bge-reranker-v2-m3 568M 0.8487 0.9663 0.8475 0.9853 0.8129 0.8222 0.6690 0.9853 0.7906
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 0.8336 0.9308 0.8489 0.9779 0.8507 0.7359 0.7813 0.9808 0.7867
nvidia/llama-nemotron-rerank-1b-v2 1.2B 0.8536 0.9480 0.8491 0.9883 0.8182 0.8026 0.6719 0.9858 0.7527
nlpai-lab/LAMAR-600m 568M 0.8461 0.9591 0.8225 0.9835 0.8214 0.7975 0.5679 0.9850 0.7822
dragonkue/bge-reranker-v2-m3-ko 568M 0.8232 0.9684 0.8708 0.9769 0.7573 0.6776 0.7061 0.9846 0.6721
BAAI/bge-reranker-v2-gemma 2.5B 0.8614 0.9407 0.8698 0.9857 0.8362 0.8429 0.2881 0.9858 0.7572
upskyy/ko-reranker-8k 568M 0.8143 0.9230 0.8388 0.9291 0.7249 0.6998 0.5975 0.9718 0.7770
Dongjin-kr/ko-reranker 560M 0.8468 0.9014 0.7675 0.9759 0.8017 0.7772 0.3721 0.9785 0.7343
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 0.8329 0.9794 0.8534 0.9777 0.8449 0.7650 0.1829 0.9850 0.6042
cross-encoder/ettin-reranker-1b-v1 1.0B 0.6624 0.8901 0.7461 0.6914 0.7004 0.6659 0.3651 0.9590 0.5306

Per-dataset PPS

Model Params Ko-StrategyQA AutoRAGRetrieval PublicHealthQA BelebeleRetrieval MIRACLRetrieval MrTidyRetrieval MultiLongDocRetrieval SQuADKorV1Retrieval LawIRKo
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 16.6 7.3 17.5 19.4 24.9 26.5 0.7 10.6 12.3
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 26.1 11.8 28.4 31.5 40.0 42.4 1.1 17.2 20.0
nlpai-lab/KURE-Reranker-base 1.7B 59.2 27.3 64.6 71.1 89.1 97.1 2.6 39.1 45.7
nlpai-lab/KURE-Reranker-nano 149M 475.4 233.2 569.7 630.9 765.8 855.2 16.5 330.8 386.0
jinaai/jina-reranker-v3.5 597M 141.0 38.6 148.6 174.6 277.3 295.7 — 68.4 86.7
jinaai/jina-reranker-v3 597M 113.8 25.0 121.8 169.2 244.6 261.6 — 48.7 64.0
zeroentropy/zerank-2-reranker 4.0B 31.0 12.7 33.8 37.8 51.0 56.4 1.1 18.8 22.3
lightonai/LightOn-rerank-PW-4B 4.5B 16.2 7.3 18.6 19.5 23.4 26.1 0.7 10.7 12.4
mixedbread-ai/mxbai-rerank-large-v2 1.5B 70.0 32.9 76.8 84.1 104.2 113.6 3.2 47.0 54.7
BAAI/bge-reranker-v2-m3 568M 420.7 195.6 471.3 545.3 653.2 724.8 9.3 271.6 345.5
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 107.1 49.6 118.0 129.3 159.5 173.9 4.2 71.8 82.9
nvidia/llama-nemotron-rerank-1b-v2 1.2B 133.7 54.4 146.9 163.0 220.0 243.1 3.9 84.9 95.9
nlpai-lab/LAMAR-600m 568M 418.0 196.5 480.5 557.2 656.7 742.5 9.2 275.0 343.8
dragonkue/bge-reranker-v2-m3-ko 568M 416.9 195.1 467.2 540.7 653.4 716.7 9.2 271.9 342.7
BAAI/bge-reranker-v2-gemma 2.5B 64.7 26.1 67.5 75.8 99.8 106.8 2.5 40.0 44.7
upskyy/ko-reranker-8k 568M 411.8 195.6 474.3 551.4 649.3 726.0 9.2 275.1 342.8
Dongjin-kr/ko-reranker 560M 520.8 258.0 510.1 554.8 747.6 765.2 272.7 334.4 381.2
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 104.6 49.4 117.8 129.8 160.5 175.6 4.3 72.1 82.6
cross-encoder/ettin-reranker-1b-v1 1.0B 52.4 19.9 51.4 65.6 92.7 97.2 3.8 31.2 33.4

Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the Ettin reranker speed benchmark.

Citation

@misc{kure-reranker-base,
  title  = {KURE-Reranker-base: A Korean–English Bilingual Reranking Model},
  author = {Jang, Youngjoon and Hong, Seongtae and Son, Junyoung and Lee, Taemin and Lim, Heuiseok},
  year   = {2026},
  url    = {https://huggingface.co/nlpai-lab/KURE-Reranker-base},
}
Downloads last month
12
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for nlpai-lab/KURE-Reranker-base

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1230)
this model