DR-GRPO-Qwen3-4B

📄 Paper: Training LLM Judges from Language Feedback via Position-Selective Self-Distillation

Pairwise generative reward model / LLM Judge trained from Qwen/Qwen3-4B-Instruct-2507 on opsd-genrm/dedup_filtered_HS3 — a deduplicated and filtered version of nvidia/HelpSteer3 — with DR. GRPO. Given a context and two candidate responses, the model first identifies the evaluation criteria that matter for the specific task, then compares the two responses step by step against those criteria, and finally emits a verdict in <verdict>A</verdict> or <verdict>B</verdict>.

Usage

Input format

Send a single user message of the form below (no system prompt). The judge expects each turn of the dialog and each candidate response to be wrapped with <user>...</user> and <assistant>...</assistant> tags.

{context} is the user-side input. For a single-turn query it is just <user>\n…question…\n</user>; for a multi-turn conversation it is the full alternating dialog (must alternate user → assistant → user → … and end on a <user> turn). {response_a} and {response_b} are the two candidate replies, each wrapped in a single <assistant> block.

You are an impartial judge tasked with determining which of two assistant responses is better for the given context.

Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.

[Start of Context]
{context}
[End of Context]

[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]

[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]

Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>.

Single-turn example

[Start of Context]
<user>
What is the capital of France?
</user>
[End of Context]

[Start of Assistant A's Response]
<assistant>
The capital of France is Paris.
</assistant>
[End of Assistant A's Response]

[Start of Assistant B's Response]
<assistant>
Lyon.
</assistant>
[End of Assistant B's Response]

Multi-turn example

[Start of Context]
<user>
I'm planning a 3-day trip to Tokyo next month. Any recommendations?
</user>

<assistant>
Sure — what kind of activities are you interested in (food, history, nightlife, shopping)?
</assistant>

<user>
Mostly food and history.
</user>
[End of Context]

[Start of Assistant A's Response]
<assistant>
Day 1: Tsukiji outer market for breakfast …
</assistant>
[End of Assistant A's Response]

[Start of Assistant B's Response]
<assistant>
Just go to Shibuya and figure it out when you get there.
</assistant>
[End of Assistant B's Response]

Quickstart

import re
from transformers import AutoModelForCausalLM, AutoTokenizer

REPO = "opsd-genrm/DR-GRPO-Qwen3-4B"
tok = AutoTokenizer.from_pretrained(REPO)
model = AutoModelForCausalLM.from_pretrained(
    REPO, torch_dtype="bfloat16", device_map="auto"
)

PROMPT = """You are an impartial judge tasked with determining which of two assistant responses is better for the given context.

Below is a context (a user query or a conversation between the user and an assistant) and two assistant responses to that context.

[Start of Context]
{context}
[End of Context]

[Start of Assistant A's Response]
{response_a}
[End of Assistant A's Response]

[Start of Assistant B's Response]
{response_b}
[End of Assistant B's Response]

Identify the quality dimensions that matter most for this specific task, then evaluate and compare the two assistant responses step by step across those dimensions. When correctness matters, solve the problem yourself and check each response for any errors. After your analysis, determine which response is better overall and provide your final verdict (A or B only) in <verdict>...</verdict>."""


def format_context(messages: list[dict]) -> str:
    """Wrap a (multi-turn) dialog as alternating <user>/<assistant> blocks.

    `messages` is a list of {"role": "user"|"assistant", "content": str},
    must alternate starting with "user", and must end on a "user" turn.
    """
    parts = []
    for i, m in enumerate(messages):
        expected = "user" if i % 2 == 0 else "assistant"
        assert m["role"] == expected, "roles must alternate user/assistant/..."
        parts.append(f"<{m['role']}>\n{m['content'].strip()}\n</{m['role']}>")
    return "\n\n".join(parts)


def format_response(text: str) -> str:
    return f"<assistant>\n{text.strip()}\n</assistant>"


def parse_verdict(text: str) -> str | None:
    """Prefers the last <verdict>...</verdict> block: returns the last 'A' or
    'B' word inside it. If the verdict tag is missing or malformed, falls back
    to the last standalone 'A'/'B' anywhere in the generated text. Returns
    None only when no A/B token appears at all.
    """
    blocks = re.findall(r"<verdict>(.*?)</verdict>", text, re.DOTALL)
    if blocks:
        ab = re.findall(r"\b(A|B)\b", blocks[-1])
        if ab:
            return ab[-1]
    ab = re.findall(r"\b(A|B)\b", text)
    return ab[-1] if ab else None


def judge(context_messages: list[dict], response_a: str, response_b: str) -> str | None:
    user_msg = PROMPT.format(
        context=format_context(context_messages),
        response_a=format_response(response_a),
        response_b=format_response(response_b),
    )
    inputs = tok.apply_chat_template(
        [{"role": "user", "content": user_msg}],
        add_generation_prompt=True, return_tensors="pt",
    ).to(model.device)
    out = model.generate(
        inputs,
        max_new_tokens=8192,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
        top_k=20,
    )
    text = tok.decode(out[0, inputs.shape[-1]:], skip_special_tokens=True)
    return parse_verdict(text)


# Single-turn:
verdict = judge(
    context_messages=[{"role": "user", "content": "What is the capital of France?"}],
    response_a="The capital of France is Paris.",
    response_b="Lyon.",
)
print(verdict)  # -> "A"

Evaluation

Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=8192.

Citation

If you use this model, please cite:

@article{hong2026training,
  title   = {Training LLM Judges from Language Feedback via Position-Selective Self-Distillation},
  author  = {Hong, Ilgee and Yu, Changlong and Xu, Zhenghao and Liu, Xin and Zhang, Yuwei and Lu, Qin and Yin, Bing and Zhao, Tuo},
  journal = {arXiv preprint arXiv:2609.38792},
  year    = {2026}
}

License

Released under Apache 2.0, inheriting from the Qwen3-4B-Instruct-2507 license.

Downloads last month
162
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for opsd-genrm/DR-GRPO-Qwen3-4B

Finetuned
(2365)
this model
Quantizations
1 model

Datasets used to train opsd-genrm/DR-GRPO-Qwen3-4B

Papers for opsd-genrm/DR-GRPO-Qwen3-4B