Instructions to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2") model = AutoModelForCausalLM.from_pretrained("JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2
- SGLang
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 with Docker Model Runner:
docker model run hf.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2
KnowWhenToHandOff-Qwen3-4B-R2 (GRPO, Responsible-AI reward)
Research artefact — not for deployment. Trained and evaluated only with simulated users on a public benchmark. It has not been tested with real customers, real data or real business systems.
Code github.com/JaspinXu/KnowWhenToHandOff · Data KnowWhenToHandOff-data · Family SFT · R1 · R2
R2 is Qwen3-4B-Instruct-2507 trained with SFT and then multi-turn GRPO on the τ²-bench airline and retail domains, with a reward that adds penalties for rule violations, unnecessary hand-offs and missed hand-offs. It has the best hand-off F1 of the study (0.812) with 10 % over-escalation, at the same task success as its SFT starting point.
Model family
| Model | Training | Hand-off F1 | Over-escalation | Use it for |
|---|---|---|---|---|
| SFT (B2) | LoRA SFT on 790 filtered teacher dialogues | 0.768 | 0.016 | starting point; precise but misses hand-offs |
| R1 | GRPO from B2, task reward only | 0.802 | 0.462 | studying the outcome-reward loophole |
| R2 | GRPO from B2, task reward + violation and hand-off penalties | 0.812 | 0.103 | the best-calibrated hand-off model |
Task success on the official test tasks is the same for all three within noise (0.33–0.35).
Quick start
A standard Qwen3 chat model with tool calling. It was trained with the
τ²-bench domain policy as the system prompt and the domain's tools, including
transfer_to_human_agents; for faithful behaviour run it inside the evaluation
harness of the code repository.
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2"
tokenizer = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(
repo, torch_dtype="auto", device_map="auto"
)
transfer = {
"type": "function",
"function": {
"name": "transfer_to_human_agents",
"description": "Transfer the user to a human agent.",
"parameters": {
"type": "object",
"properties": {"summary": {"type": "string"}},
"required": ["summary"],
},
},
}
messages = [
{"role": "system", "content": "<airline or retail policy>"},
{"role": "user", "content": "I need the unaccompanied-minor service."},
]
inputs = tokenizer.apply_chat_template(
messages, tools=[transfer], add_generation_prompt=True,
return_tensors="pt",
).to(model.device)
output = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(output[0][inputs.shape[-1]:], skip_special_tokens=True))
Serve it with vLLM (OpenAI-compatible API, tool calls parsed):
vllm serve JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2 \
--enable-auto-tool-choice --tool-call-parser hermes
Evaluation
Test split, protocol frozen in ADR-005; 95 % bootstrap intervals over tasks.
| Metric | This model | B2 (SFT start) | Source |
|---|---|---|---|
| Success (pass^1), test_main | 0.342 [0.245, 0.449] | 0.327 [0.224, 0.434] | main table |
| Paired Δ success vs B2 | +0.015 [−0.056, 0.082] | — | paired differences |
| pass^4 | 0.122 [0.041, 0.224] | 0.102 [0.041, 0.204] | main table |
| Violation rate (incl. off-reference writes) | 0.444 | 0.469 | same |
| Hand-off F1 / precision / recall, test_handoff | 0.812 / 0.886 / 0.750 | 0.768 / 0.976 / 0.633 | R2-final-test_handoff-s0-20261007 |
| Over-escalation, test_handoff | 0.103 | 0.016 | same |
| Worst-style success / style gap | 0.316 / 0.112 | 0.327 / 0.112 | styles |
Known behaviour. R2 removes R1's over-escalation but sometimes makes the opposite error on out-of-scope requests: it refuses and tells the user to contact the company directly without calling the transfer tool, or substitutes an unrelated write (under-escalation 0.25). Once it justified a hand-off with a policy that does not exist.
Training details
- Base model: Qwen/Qwen3-4B-Instruct-2507 (revision
cdbee75f17c01a7cc42f958dc650907174af0554), Apache-2.0. - Training: GRPO from the SFT checkpoint B2 with the composite reward (
configs/reward/r2.yaml: task reward − 0.25 × weighted violations − 0.5 × unneeded transfer − 1.0 × missed transfer − 0.1 × format errors); LoRA rank 64, alpha 128, lr 1e-5; 8 tasks × 8 rollouts per step, 60 steps, no KL term. - Training run
R2-s1-20261006(seed 1); exported step-60 checkpointsha256:be6198a0c51773f52ba4f20c9605a5931f23b0206e92a9fe0521a93761348c96(merged bf16, Hugging Face format). - Evaluation code at git
462b5d17b4e9add637a2ee7802974c720b625a5b; τ²-benchfc0055dc4e0a316c3f83133267fbd6faaa770992; all sources in reports/main/sources.json.
Intended use
Studying when a tool-using customer-service agent should hand off to a human, in the τ²-bench airline and retail domains. Out of scope: any production or customer-facing use, other domains, and decisions about real people.
Training data
Simulated conversations on τ²-bench v1.0.1 train-split tasks (a dev split was carved out for selection) and hand-off variants derived from them (explicit requests for a human, out-of-scope requests, hard negatives); the user simulator is Qwen3-30B-A3B-Instruct-2507. No real user data. Test tasks were never used for training or selection.
Limitations
- One user simulator (Qwen3-30B-A3B-Instruct-2507, which also generated the SFT data) and one training seed; results may not transfer to real users or other simulators, and the intervals cover task sampling, not training variance.
- Small test split (49 official tasks; retail uses the 74 tasks whose reward needs no LLM judge, so retail numbers are not comparable with leaderboards).
- Communication styles describe how users write, not who they are; no claim is made about any demographic group.
- Policy compliance is checked by five deterministic rules that cover only part of each domain policy; citation rate is a keyword rule (it counted one invented policy as a citation).
- Observed reward-hacking behaviours: docs/rl-audit.md.
Licences
The weights are a LoRA fine-tune of Qwen/Qwen3-4B-Instruct-2507 and inherit
its Apache-2.0 licence. τ²-bench code and data (pinned commit
fc0055dc4e0a316c3f83133267fbd6faaa770992) are used under the licence in
its LICENSE file;
released dialogue samples are τ²-bench-derived synthetic conversations with
no real user data.
Citation
Model DOI: 10.57967/hf/10823. Cite the model and the project:
@misc{knowwhentohandoff2026r2,
title = {KnowWhenToHandOff-Qwen3-4B-R2},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
publisher = {Hugging Face},
doi = {10.57967/hf/10823},
url = {https://huggingface.co/JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2}
}
@misc{knowwhentohandoff2026,
title = {Know When to Hand Off: Multi-turn Reinforcement Learning for
Trustworthy Customer-Service Agents},
author = {{KnowWhenToHandOff contributors}},
year = {2026},
howpublished = {\url{https://github.com/JaspinXu/KnowWhenToHandOff}}
}
- Downloads last month
- 15
Model tree for JaspinXu/KnowWhenToHandOff-Qwen3-4B-R2
Base model
Qwen/Qwen3-4B-Instruct-2507