Instructions to use nlpai-lab/RenderRank-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nlpai-lab/RenderRank-2B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("nlpai-lab/RenderRank-2B", trust_remote_code=True) model = AutoModelForMultimodalLM.from_pretrained("nlpai-lab/RenderRank-2B", trust_remote_code=True, device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - sentence-transformers
How to use nlpai-lab/RenderRank-2B with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("nlpai-lab/RenderRank-2B", trust_remote_code=True) query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
RenderRank
RenderRank: Learning to Rerank Text with Compressed Visual Tokens
Paper: RenderRank: Learning to Rerank Text with Compressed Visual Tokens
RenderRank converts text into compact visual token sequences by rendering it as images. This change in input modality reduces the number of tokens used to represent documents, allowing more content to fit within the same context budget. Queries remain text, and relevance is scored against the visual document representations.
Highlights
- Text as visual tokens: represent document text through rendered images instead of conventional text token sequences.
- Shorter input sequences: compact visual representations reduce input length and allow more document content within a fixed context budget.
- Three input modes: score document text directly with
text, render it automatically withrender, or provide existing document pages withimage.
| Property | Value |
|---|---|
| Model type | Pairwise document reranker |
| Backbone | Qwen3-VL-Reranker-2B |
| Parameters | 2B |
| Language evaluated | English |
| Maximum context length | 32,768 tokens, including query, template, document text, and visual tokens |
| Scoring | Yes/no relevance scoring |
Rendering Configuration
| Setting | Value |
|---|---|
| Font | Roboto Regular |
| Font size | 12pt at 96 DPI (16px) |
| Line spacing | 1.0 (16px line height) |
| Page width | 896px |
| Page height | 32–896px, rounded up to a multiple of 32px |
| Lines per page | Up to 56 |
| Margins | 0px on all sides |
| Colors | Black text on a white background |
| Rendering | Pillow default text antialiasing |
Whitespace, including paragraph breaks and tabs, is normalized to a single space. Lines wrap according to the font's actual pixel width; words wider than the page are split by character. Overflow continues onto subsequent pages without overlap. The final page is sized to its content, and empty documents produce one 896 × 32px page.
All generated pages are supplied in order as one document, producing one relevance score. Rendering is not limited to the first page.
BEIR Performance
All rerankers score the top 100 BM25 candidates for each query. Baselines use text queries and documents; RenderRank uses text queries and documents rendered with the default 12pt configuration above. Scores are NDCG@10; Avg. is the mean across 11 datasets. Tokens is the mean input length per query–document pair before truncation, averaged across the same datasets. For RenderRank, this includes both textual and visual tokens.
| Model | AA | CFV | DBP | FQA | FVR | HQA | NFC | SD | SF | TC | TCH | Avg. | Tokens |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| gte-reranker-modernbert-base | 66.14 | 25.80 | 42.10 | 42.54 | 89.84 | 76.15 | 34.81 | 18.90 | 75.81 | 79.66 | 33.44 | 53.20 | 359.25 |
| mxbai-rerank-large-v1 | 16.16 | 25.61 | 45.37 | 40.29 | 82.19 | 71.58 | 37.32 | 18.96 | 75.26 | 85.84 | 37.44 | 48.73 | 347.45 |
| bge-reranker-large | 30.94 | 34.11 | 44.21 | 37.47 | 89.30 | 80.09 | 33.91 | 16.68 | 73.98 | 73.21 | 34.69 | 49.87 | 409.32 |
| LAMAR-600m | 66.21 | 36.21 | 46.23 | 41.14 | 89.49 | 79.14 | 35.04 | 20.22 | 76.97 | 80.65 | 35.76 | 55.19 | 409.32 |
| Qwen3-Reranker-0.6B | 68.04 | 34.34 | 43.76 | 39.74 | 87.13 | 77.25 | 36.60 | 20.44 | 77.00 | 85.54 | 30.32 | 54.56 | 438.70 |
| llama-nemotron-rerank-1b-v2 | 54.29 | 27.37 | 45.13 | 47.11 | 87.66 | 80.34 | 38.14 | 21.53 | 79.82 | 83.14 | 32.47 | 54.27 | 363.99 |
| mxbai-rerank-large-v2 | 42.23 | 23.03 | 38.01 | 28.85 | 70.84 | 69.50 | 35.17 | 17.49 | 78.81 | 67.78 | 46.94 | 47.15 | 449.76 |
| LightOn-rerank-PW-2B | 45.52 | 23.21 | 42.60 | 35.66 | 86.51 | 73.05 | 35.61 | 16.64 | 77.70 | 81.24 | 36.91 | 50.42 | 416.22 |
| bge-reranker-v2-gemma | 74.98 | 32.03 | 45.01 | 43.18 | 87.45 | 80.47 | 36.25 | 19.58 | 77.43 | 80.08 | 37.56 | 55.82 | 399.11 |
| Qwen3-Reranker-4B | 75.40 | 37.84 | 47.02 | 44.63 | 89.14 | 79.06 | 37.48 | 23.59 | 79.47 | 85.47 | 38.77 | 57.99 | 438.70 |
| zerank-2-reranker | 44.87 | 23.65 | 44.95 | 42.75 | 83.22 | 70.97 | 38.50 | 20.23 | 79.35 | 85.24 | 38.10 | 51.98 | 380.69 |
| LightOn-rerank-PW-4B | 52.50 | 28.24 | 44.80 | 40.27 | 87.70 | 72.79 | 37.33 | 17.29 | 76.72 | 79.54 | 36.70 | 52.17 | 416.22 |
| RenderRank (Ours) | 69.11 | 35.91 | 46.19 | 42.66 | 88.92 | 78.60 | 37.31 | 20.75 | 77.00 | 84.82 | 34.27 | 55.96 | 290.07 |
AA: ArguAna; CFV: Climate-FEVER; DBP: DBPedia; FQA: FiQA; FVR: FEVER; HQA: HotpotQA; NFC: NFCorpus; SD: SCIDOCS; SF: SciFact; TC: TREC-COVID; TCH: Touché-2020.
Long-Document Performance
Results on the English subset of MLDR and three LongEmbed datasets: 2WikiMQA, QMSum, and SummScreenFD. The comparison includes models supporting at least 16K input tokens. MLDR follows the MMTEB reranking protocol; for LongEmbed, each query reranks eight candidates retrieved with Qwen3-Embedding-0.6B.
Each dataset cell shows NDCG@10 followed by the average input token count in brackets: score [tokens]. Token counts are measured per query–document pair before truncation and include both textual and visual tokens for RenderRank. Avg. reports the mean score across the four datasets.
| Model | MLDR | 2WikiMQA | QMSum | SummScreenFD | Avg. |
|---|---|---|---|---|---|
| Qwen3-Reranker-0.6B | 99.63 [8863.4] | 94.54 [9230.2] | 56.43 [13543.6] | 98.25 [8606.3] | 87.21 |
| LightOn-rerank-PW-2B | 98.91 [8878.0] | 71.27 [9230.1] | 54.30 [13876.6] | 93.05 [8977.2] | 79.38 |
| Qwen3-Reranker-4B | 99.85 [8863.4] | 94.54 [9230.2] | 59.11 [13543.6] | 99.11 [8606.3] | 88.15 |
| zerank-2-reranker | 99.57 [8805.6] | 94.42 [9173.1] | 60.10 [13485.9] | 98.97 [8548.5] | 88.26 |
| LightOn-rerank-PW-4B | 99.68 [8878.0] | 93.84 [9230.1] | 58.87 [13876.6] | 98.91 [8977.2] | 87.82 |
| RenderRank (Ours) | 99.74 [4198.0] | 94.30 [4262.2] | 59.94 [6388.8] | 99.10 [3701.9] | 88.27 |
Usage
The examples load the model from nlpai-lab/RenderRank-2B on Hugging Face.
The model applies its reranking chat template internally. Pass the instruction as shown below; do not manually prepend a chat template to the document.
Transformers
This repository provides a custom AutoModel entry point with a process() method. It supports all three document input modes.
Requirements
The Transformers interface was tested with the following versions:
transformers==5.9.0
qwen-vl-utils==0.0.14
torch==2.11.0
torchvision==0.26.0
scipy
qwen-vl-utils handles image preparation in this interface; scipy is used for score normalization. NumPy and Pillow are installed through the dependencies above. Install PyTorch and torchvision builds matching your CUDA environment.
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"nlpai-lab/RenderRank-2B",
trust_remote_code=True,
dtype=torch.bfloat16,
).to("cuda").eval()
inputs = {
"instruction": "Retrieve text relevant to the user's query.",
"query": {"text": "What is the capital of France?"},
"documents": [
# Score text directly.
{"text": "Paris is the capital of France."},
# Render text into document pages, then score the images.
{"render": "Paris is the capital of France."},
# Score an existing document image.
{"image": "/path/to/document.png"},
# Score multiple ordered pages as one document.
{"image": ["/path/to/page1.png", "/path/to/page2.png"]},
],
}
with torch.inference_mode():
scores = model.process(inputs)
ranked_indices = sorted(range(len(scores)), key=lambda i: scores[i], reverse=True)
print(scores)
print(ranked_indices)
Rendering
With Transformers, {"render": document_text} automatically renders the document before scoring. The font is included in this repository.
Alternatively, render documents in advance with the bundled rendering.py and save the page images for reuse with either Transformers or Sentence Transformers. This avoids repeated rendering; vision encoding still runs when scoring the saved images.
import sys
from pathlib import Path
from huggingface_hub import snapshot_download
repo_dir = Path(snapshot_download(
"nlpai-lab/RenderRank-2B",
allow_patterns=["rendering.py", "Roboto-Regular.ttf"],
))
sys.path.insert(0, str(repo_dir))
from rendering import DocumentImageConfig, DocumentImageRenderer
renderer = DocumentImageRenderer(
DocumentImageConfig(font_path=repo_dir / "Roboto-Regular.ttf")
)
pages = renderer.render_document("Paris is the capital of France.")
# pages is a list of PIL images, in document order.
# Save once and reuse these files for subsequent queries.
output_dir = Path("rendered_document")
output_dir.mkdir(parents=True, exist_ok=True)
for i, page in enumerate(pages, start=1):
page.save(output_dir / f"{i}.png")
The returned PIL images can be passed directly as image content or saved as numbered page files. The Transformers render mode calls this renderer internally, so no separate rendering step is needed there.
Sentence Transformers
Use a Sentence Transformers version supporting multimodal CrossEncoder models.
pip install -U sentence-transformers
The same call can score text, a single document image, and multiple page images. In the example below, 1.png, 2.png, and 3.png are ordered pages of one document.
import torch
from sentence_transformers import CrossEncoder
model = CrossEncoder(
"nlpai-lab/RenderRank-2B",
device="cuda",
max_length=32768,
model_kwargs={"dtype": torch.bfloat16},
)
query = "What is the capital of France?"
documents = [
[{"type": "text", "text": "Paris is the capital of France."}],
[{"type": "image", "image": "/path/to/document.png"}],
[
{"type": "image", "image": "/path/to/1.png"},
{"type": "image", "image": "/path/to/2.png"},
{"type": "image", "image": "/path/to/3.png"},
# Add further pages here in document order.
],
]
messages = [
[
{"role": "query", "content": [{"type": "text", "text": query}]},
{"role": "document", "content": document},
]
for document in documents
]
scores = model.predict(
messages,
prompt="Retrieve text relevant to the user's query.",
batch_size=1,
)
print(scores) # One score per document, not per page.
Larger scores indicate greater relevance. Scores are the difference between the yes and no logits, as used in our evaluation. Sentence Transformers requires one image content entry per page; render and {"image": [page1, page2]} are supported by the Transformers interface above, not this interface.
Citation
@misc{hong2026renderranklearningreranktext,
title={RenderRank: Learning to Rerank Text with Compressed Visual Tokens},
author={Seongtae Hong and Youngjoon Jang and Jungseob Lee and Hyeonseok Moon and Heuiseok Lim},
year={2026},
eprint={2609.35069},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2609.35069},
}
- Downloads last month
- -
