Instructions to use opsd-genrm/DR-GRPO-Qwen3-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use opsd-genrm/DR-GRPO-Qwen3-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="opsd-genrm/DR-GRPO-Qwen3-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("opsd-genrm/DR-GRPO-Qwen3-4B") model = AutoModelForCausalLM.from_pretrained("opsd-genrm/DR-GRPO-Qwen3-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use opsd-genrm/DR-GRPO-Qwen3-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "opsd-genrm/DR-GRPO-Qwen3-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "opsd-genrm/DR-GRPO-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/opsd-genrm/DR-GRPO-Qwen3-4B
- SGLang
How to use opsd-genrm/DR-GRPO-Qwen3-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "opsd-genrm/DR-GRPO-Qwen3-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "opsd-genrm/DR-GRPO-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "opsd-genrm/DR-GRPO-Qwen3-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "opsd-genrm/DR-GRPO-Qwen3-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use opsd-genrm/DR-GRPO-Qwen3-4B with Docker Model Runner:
docker model run hf.co/opsd-genrm/DR-GRPO-Qwen3-4B
DR-GRPO-Qwen3-4B
📄 Paper: Training LLM Judges from Language Feedback via Position-Selective Self-Distillation
Pairwise generative reward model / LLM Judge trained from
Qwen/Qwen3-4B-Instruct-2507
on opsd-genrm/dedup_filtered_HS3
— a deduplicated and filtered version of
nvidia/HelpSteer3 —
with DR. GRPO. Given a context
and two candidate responses, the model first identifies the evaluation
criteria that matter for the specific task, then compares the two
responses step by step against those criteria, and finally emits a verdict
in <verdict>A</verdict> or <verdict>B</verdict>.
Usage
Input format
Send a single user message of the form below (no system prompt). The judge
expects each turn of the dialog and each candidate response to be wrapped
with <user>...</user> and <assistant>...</assistant> tags.
{context} is the user-side input. For a single-turn query it is just
<user>\n…question…\n</user>; for a multi-turn conversation it is the full
alternating dialog (must alternate user → assistant → user → … and end on
a <user> turn). {response_a} and {response_b} are the two candidate
replies, each wrapped in a single <assistant> block.
You are an impartial judge tasked with determining which of two assistant responses is better for the given context.
Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.
[Start of Context]
{context}
[End of Context]
[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]
[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]
Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>.
Single-turn example
[Start of Context]
<user>
What is the capital of France?
</user>
[End of Context]
[Start of Assistant A's Response]
<assistant>
The capital of France is Paris.
</assistant>
[End of Assistant A's Response]
[Start of Assistant B's Response]
<assistant>
Lyon.
</assistant>
[End of Assistant B's Response]
Multi-turn example
[Start of Context]
<user>
I'm planning a 3-day trip to Tokyo next month. Any recommendations?
</user>
<assistant>
Sure — what kind of activities are you interested in (food, history, nightlife, shopping)?
</assistant>
<user>
Mostly food and history.
</user>
[End of Context]
[Start of Assistant A's Response]
<assistant>
Day 1: Tsukiji outer market for breakfast …
</assistant>
[End of Assistant A's Response]
[Start of Assistant B's Response]
<assistant>
Just go to Shibuya and figure it out when you get there.
</assistant>
[End of Assistant B's Response]
Quickstart
import re
from transformers import AutoModelForCausalLM, AutoTokenizer
REPO = "opsd-genrm/DR-GRPO-Qwen3-4B"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(
REPO, torch_dtype="bfloat16", device_map="auto"
)
PROMPT = """You are an impartial judge tasked with determining which of two assistant responses is better for the given context.
Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.
[Start of Context]
{context}
[End of Context]
[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]
[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]
Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>."""
def format_context(messages: list[dict]) -> str:
"""Wrap a (multi-turn) dialog as alternating <user>/<assistant> blocks.
`messages` is a list of {"role": "user"|"assistant", "content": str},
must alternate starting with "user", and must end on a "user" turn.
"""
parts = []
for i, m in enumerate(messages):
expected = "user" if i % 2 == 0 else "assistant"
assert m["role"] == expected, "roles must alternate user/assistant/..."
parts.append(f"<{m['role']}>\n{m['content'].strip()}\n</{m['role']}>")
return "\n\n".join(parts)
def format_response(text: str) -> str:
return f"<assistant>\n{text.strip()}\n</assistant>"
def parse_verdict(text: str) -> str | None:
"""Prefers the last <verdict>...</verdict> block: returns the last 'A' or
'B' word inside it. If the verdict tag is missing or malformed, falls back
to the last standalone 'A'/'B' anywhere in the generated text. Returns
None only when no A/B token appears at all.
"""
blocks = re.findall(r"<verdict>(.*?)</verdict>", text, re.DOTALL)
if blocks:
ab = re.findall(r"\b(A|B)\b", blocks[-1])
if ab:
return ab[-1]
ab = re.findall(r"\b(A|B)\b", text)
return ab[-1] if ab else None
def judge(context_messages: list[dict], response_a: str, response_b: str) -> str | None:
user_msg = PROMPT.format(
context=format_context(context_messages),
response_a=format_response(response_a),
response_b=format_response(response_b),
)
inputs = tok.apply_chat_template(
[{"role": "user", "content": user_msg}],
add_generation_prompt=True, return_tensors="pt",
).to(model.device)
out = model.generate(
inputs,
max_new_tokens=8192,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
)
text = tok.decode(out[0, inputs.shape[-1]:], skip_special_tokens=True)
return parse_verdict(text)
# Single-turn:
verdict = judge(
context_messages=[{"role": "user", "content": "What is the capital of France?"}],
response_a="The capital of France is Paris.",
response_b="Lyon.",
)
print(verdict) # -> "A"
Evaluation
Generation config: temperature=0.7, top_p=0.8, top_k=20,
max_tokens=8192.
Citation
If you use this model, please cite:
@article{hong2026training,
title = {Training LLM Judges from Language Feedback via Position-Selective Self-Distillation},
author = {Hong, Ilgee and Yu, Changlong and Xu, Zhenghao and Liu, Xin and Zhang, Yuwei and Lu, Qin and Yin, Bing and Zhao, Tuo},
journal = {arXiv preprint arXiv:2609.38792},
year = {2026}
}
License
Released under Apache 2.0, inheriting from the Qwen3-4B-Instruct-2507 license.
- Downloads last month
- 162