KURE logo

KURE-Reranker-nano

KURE-Reranker-nano is a lightweight, approximately 150M-parameter Korean-English bilingual reranker. It jointly encodes a query and a candidate document and predicts a scalar relevance score.

Key Characteristics

  • Strong Korean reranking: 88.08 average nDCG@10 across nine Korean benchmarks, the highest average among the evaluated models with fewer than 1B parameters and only 0.42 points below KURE-Reranker-base (1.7B).
  • Long-document reranking: 79.66 nDCG@10 on MultiLongDocRetrieval (MLDR), the best among the evaluated models with fewer than 1B parameters.
  • High throughput: 473.7 pairs per second averaged over nine benchmarks on an NVIDIA RTX A6000 48GB, about 19.5Γ— Qwen3-Reranker-4B under the measurement conditions described below.
  • Lightweight: Approximately 150M parameters, based on skt/A.X-Encoder-base.

Model Overview

Property Value
Model type Cross-encoder reranker
Backbone A.X-Encoder-base
Parameters 149M
Languages Korean and English
Maximum inference input length 8,192 tokens, including query, document, and special tokens
Output One scalar relevance score per query–document pair
License Apache-2.0

Usage

Install the required libraries:

pip install "sentence-transformers>=5.1.0" "transformers>=4.55.4" torch
Sentence Transformers
from sentence_transformers import CrossEncoder

model = CrossEncoder("nlpai-lab/KURE-Reranker-nano", max_length=8192)

query = "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ–΄λ””μΈκ°€μš”?"
documents = [
    "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ„œμšΈμž…λ‹ˆλ‹€.",
    "The capital of South Korea is Seoul.",
    "λ°”λ‚˜λ‚˜λŠ” μ—΄λŒ€ μ§€μ—­μ—μ„œ μž¬λ°°λ˜λŠ” κ³ΌμΌμž…λ‹ˆλ‹€.",
]

results = model.rank(query, documents, batch_size=8, return_documents=True)
for result in results:
    print(f"{result['score']:.4f}\t{result['text']}")

For pair scoring without sorting:

pairs = [(query, document) for document in documents]
scores = model.predict(pairs, batch_size=8)
Transformers
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

model_id = "nlpai-lab/KURE-Reranker-nano"
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).to(device)
model.eval()

query = "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ–΄λ””μΈκ°€μš”?"
documents = [
    "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ„œμšΈμž…λ‹ˆλ‹€.",
    "The capital of South Korea is Seoul.",
    "λ°”λ‚˜λ‚˜λŠ” μ—΄λŒ€ μ§€μ—­μ—μ„œ μž¬λ°°λ˜λŠ” κ³ΌμΌμž…λ‹ˆλ‹€.",
]
inputs = tokenizer(
    [query] * len(documents), documents,
    padding=True, truncation=True, max_length=8192, return_tensors="pt",
).to(device)

with torch.inference_mode():
    scores = model(**inputs).logits.squeeze(-1).float().cpu()

for index in torch.argsort(scores, descending=True).tolist():
    print(f"{scores[index]:.4f}\t{documents[index]}")

Korean Reranking Evaluation

Scores are nDCG@10, and throughput is pairs per second (PPS). Evaluation was conducted using the reranker-simple-benchmark implementation as a reference.

Reranking performance (mean nDCG@10) vs. throughput (mean PPS)

The figure plots mean nDCG@10 against mean PPS across the nine Korean benchmarks, including MLDR. Dot color marks model size. PPS uses each model's own input limit and its best measured batch size, so it measures neither speed at equal token lengths nor end-to-end retrieval latency.

Results β€” MTEB-ko-retrieval (9 subsets)

Model Params Mean NDCG@10 Mean PPS
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 0.9004 15.1
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 0.8956 24.3
nlpai-lab/KURE-Reranker-base 1.7B 0.8849 55.1
nlpai-lab/KURE-Reranker-nano 149M 0.8808 473.7
zeroentropy/zerank-2-reranker 4.0B 0.8695 29.4
lightonai/LightOn-rerank-PW-4B 4.5B 0.8664 15.0
mixedbread-ai/mxbai-rerank-large-v2 1.5B 0.8661 65.2
BAAI/bge-reranker-v2-m3 568M 0.8586 404.1
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 0.8585 99.6
nvidia/llama-nemotron-rerank-1b-v2 1.2B 0.8522 127.3
nlpai-lab/LAMAR-600m 568M 0.8406 408.8
dragonkue/bge-reranker-v2-m3-ko 568M 0.8263 401.5
BAAI/bge-reranker-v2-gemma 2.5B 0.8186 58.7
upskyy/ko-reranker-8k 568M 0.8085 404.0
Dongjin-kr/ko-reranker 560M 0.7950 482.8
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 0.7806 99.6
cross-encoder/ettin-reranker-1b-v1 1.0B 0.6901 49.7

PPS is measured on a single NVIDIA RTX A6000 48GB per model and averaged over the nine benchmarks. Inputs are length-sorted and dynamically padded within each batch.

jinaai/jina-reranker-v3 and jinaai/jina-reranker-v3.5 are listwise rerankers that cannot be compared under the same 8,192-token condition on MLDR, so they are excluded from the 9-subset table and their MLDR cells are left blank below.

Per-dataset NDCG@10

Model Params Ko-StrategyQA AutoRAGRetrieval PublicHealthQA BelebeleRetrieval MIRACLRetrieval MrTidyRetrieval MultiLongDocRetrieval SQuADKorV1Retrieval LawIRKo
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 0.8679 0.9546 0.8893 0.9907 0.8490 0.8409 0.8220 0.9880 0.9014
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 0.8733 0.9707 0.8685 0.9906 0.8533 0.8321 0.8105 0.9861 0.8752
nlpai-lab/KURE-Reranker-base 1.7B 0.8548 0.9762 0.8716 0.9880 0.8371 0.7970 0.8163 0.9891 0.8343
nlpai-lab/KURE-Reranker-nano 149M 0.8582 0.9741 0.8475 0.9830 0.8400 0.8011 0.7966 0.9895 0.8368
jinaai/jina-reranker-v3.5 597M 0.8539 0.9838 0.8094 0.9733 0.8565 0.8194 β€” 0.9887 0.8545
jinaai/jina-reranker-v3 597M 0.8553 0.9773 0.7960 0.9695 0.8449 0.8104 β€” 0.9859 0.8546
zeroentropy/zerank-2-reranker 4.0B 0.8712 0.9436 0.8646 0.9846 0.8003 0.8027 0.7120 0.9791 0.8669
lightonai/LightOn-rerank-PW-4B 4.5B 0.8567 0.9321 0.8693 0.9882 0.8072 0.8091 0.7609 0.9803 0.7938
mixedbread-ai/mxbai-rerank-large-v2 1.5B 0.8563 0.9531 0.8772 0.9778 0.7939 0.8771 0.6787 0.9681 0.8130
BAAI/bge-reranker-v2-m3 568M 0.8487 0.9663 0.8475 0.9853 0.8129 0.8222 0.6690 0.9853 0.7906
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 0.8336 0.9308 0.8489 0.9779 0.8507 0.7359 0.7813 0.9808 0.7867
nvidia/llama-nemotron-rerank-1b-v2 1.2B 0.8536 0.9480 0.8491 0.9883 0.8182 0.8026 0.6719 0.9858 0.7527
nlpai-lab/LAMAR-600m 568M 0.8461 0.9591 0.8225 0.9835 0.8214 0.7975 0.5679 0.9850 0.7822
dragonkue/bge-reranker-v2-m3-ko 568M 0.8232 0.9684 0.8708 0.9769 0.7573 0.6776 0.7061 0.9846 0.6721
BAAI/bge-reranker-v2-gemma 2.5B 0.8614 0.9407 0.8698 0.9857 0.8362 0.8429 0.2881 0.9858 0.7572
upskyy/ko-reranker-8k 568M 0.8143 0.9230 0.8388 0.9291 0.7249 0.6998 0.5975 0.9718 0.7770
Dongjin-kr/ko-reranker 560M 0.8468 0.9014 0.7675 0.9759 0.8017 0.7772 0.3721 0.9785 0.7343
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 0.8329 0.9794 0.8534 0.9777 0.8449 0.7650 0.1829 0.9850 0.6042
cross-encoder/ettin-reranker-1b-v1 1.0B 0.6624 0.8901 0.7461 0.6914 0.7004 0.6659 0.3651 0.9590 0.5306

Per-dataset PPS

Model Params Ko-StrategyQA AutoRAGRetrieval PublicHealthQA BelebeleRetrieval MIRACLRetrieval MrTidyRetrieval MultiLongDocRetrieval SQuADKorV1Retrieval LawIRKo
tomaarsen/Qwen3-Reranker-8B-seq-cls 7.6B 16.6 7.3 17.5 19.4 24.9 26.5 0.7 10.6 12.3
tomaarsen/Qwen3-Reranker-4B-seq-cls 4.0B 26.1 11.8 28.4 31.5 40.0 42.4 1.1 17.2 20.0
nlpai-lab/KURE-Reranker-base 1.7B 59.2 27.3 64.6 71.1 89.1 97.1 2.6 39.1 45.7
nlpai-lab/KURE-Reranker-nano 149M 475.4 233.2 569.7 630.9 765.8 855.2 16.5 330.8 386.0
jinaai/jina-reranker-v3.5 597M 141.0 38.6 148.6 174.6 277.3 295.7 β€” 68.4 86.7
jinaai/jina-reranker-v3 597M 113.8 25.0 121.8 169.2 244.6 261.6 β€” 48.7 64.0
zeroentropy/zerank-2-reranker 4.0B 31.0 12.7 33.8 37.8 51.0 56.4 1.1 18.8 22.3
lightonai/LightOn-rerank-PW-4B 4.5B 16.2 7.3 18.6 19.5 23.4 26.1 0.7 10.7 12.4
mixedbread-ai/mxbai-rerank-large-v2 1.5B 70.0 32.9 76.8 84.1 104.2 113.6 3.2 47.0 54.7
BAAI/bge-reranker-v2-m3 568M 420.7 195.6 471.3 545.3 653.2 724.8 9.3 271.6 345.5
tomaarsen/Qwen3-Reranker-0.6B-seq-cls 596M 107.1 49.6 118.0 129.3 159.5 173.9 4.2 71.8 82.9
nvidia/llama-nemotron-rerank-1b-v2 1.2B 133.7 54.4 146.9 163.0 220.0 243.1 3.9 84.9 95.9
nlpai-lab/LAMAR-600m 568M 418.0 196.5 480.5 557.2 656.7 742.5 9.2 275.0 343.8
dragonkue/bge-reranker-v2-m3-ko 568M 416.9 195.1 467.2 540.7 653.4 716.7 9.2 271.9 342.7
BAAI/bge-reranker-v2-gemma 2.5B 64.7 26.1 67.5 75.8 99.8 106.8 2.5 40.0 44.7
upskyy/ko-reranker-8k 568M 411.8 195.6 474.3 551.4 649.3 726.0 9.2 275.1 342.8
Dongjin-kr/ko-reranker 560M 520.8 258.0 510.1 554.8 747.6 765.2 272.7 334.4 381.2
telepix/PIXIE-Spell-Reranker-Preview-0.6B 596M 104.6 49.4 117.8 129.8 160.5 175.6 4.3 72.1 82.6
cross-encoder/ettin-reranker-1b-v1 1.0B 52.4 19.9 51.4 65.6 92.7 97.2 3.8 31.2 33.4

Batch size starts at 8 and doubles until an out-of-memory error. Samples are repeated to fill complete batches. After warmup, three full passes are timed with CUDA events. Throughput is the total number of processed pairs divided by the accumulated GPU forward time. The highest-throughput successful batch is reported. Tokenization, data loading, and CPU preprocessing are excluded. The batch-search approach is informed by the Ettin reranker speed benchmark.

Citation

@misc{kure-reranker-nano,
  title  = {KURE-Reranker-nano: A Lightweight Korean–English Bilingual Reranking Model},
  author = {Hong, Seongtae and Jang, Youngjoon and Son, Junyoung and Lee, Taemin and Lim, Heuiseok},
  year   = {2026},
  url    = {https://huggingface.co/nlpai-lab/KURE-Reranker-nano},
}
Downloads last month
18
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for nlpai-lab/KURE-Reranker-nano

Finetuned
(15)
this model